Connect your AI agent — about a day of your team’s time, no production data, no system prompt. We run it against a fixed set of our situations and come back with a range.
You get: a range, not a number — how often your agent ended up worse than doing nothing, somewhere between X% and Y%, with every run behind it. We drew the span wide rather than narrow, and we put no confidence level on it: a small fixed set cannot do better. The paid work is what narrows it. Wide enough to tell you whether to look further; too wide to hand a customer or a board.
It comes with a short design partner agreement under mutual NDA: your keys and prompts stay with you, we never look inside your agent, and we never name you without your say-so.