How it works

From connected agent to a number you can act on.

Five steps. The first two are free.

  1. 01

    Connect your agentunder a day

    A small SDK adapter wraps your agent's existing loop — your code stays in charge. Your API keys and prompts stay with you: the protocol carries observations and decisions, nothing else. We keep the decision record from each run as that run's evidence — never to train anything — and we process it in the EU.
  2. 02

    Practice, freesame week · no cost

    A short fixed set of our situations, free. A situation is a controlled system we put your agent in charge of for a long stretch — decisions play out over simulated years, and small misjudgments compound. Your agent decides; we measure every run against one fixed baseline: the same situation left alone.

    What comes back is a range, and nothing else: how often your agent ended up worse than doing nothing, somewhere between X% and Y%, with every run behind it. We erred wide rather than narrow and put no confidence level on it, because a short fixed set cannot support one. Narrowing it is what the paid steps below do.

    You can debug offline as many times as you like, and a conformance check confirms your integration is sound before anything counts.

    BETTERWORSEthe system on its own · no AIthe system with your AIsame start · same situationthe gap= its real effectSee one situation play out, decision by decision →

    We don't claim a score here predicts your production outcomes — no benchmark can. We claim it reveals a failure mode you cannot see in one-off testing: what your agent's compounding decisions do to a system over time, and whether it knows when to leave well enough alone.

  3. 03

    Pilotabout four weeks · fixed price $15,000 · your team's time: the integration day plus reading the report

    A fixed-scope paid evaluation: your agent, a defined set of our situations. We fix the scope in the agreement before we start, and scope is what moves the price. You receive a bound report: how often your agent ended up worse than doing nothing, with a 95% confidence bound, plus its restraint profile. The runs are ours. The pilot fee is credited in full toward a Private Season started within 90 days.
  4. 04

    Private Seasonthe full measurement · 120–180 fresh situations, one to two passes each

    120–180 fresh situations your agent has never seen, one to two passes each — breadth, not repetition, is what buys precision. You receive the certificate (first page for your board, technical annex for your engineers) and the full report. It places your number against the anonymized range of all agent versions measured to date, without anyone being ranked or named. Coming soon as an add-on: the failure diagnosis, down to the single decision that cost the most. Fixed price, typically $40,000–$50,000 — quoted per engagement; your Pilot fee is credited in full. The benchmark is the experiment. The counterfactual method is how we measure. The run receipt is how you verify. The certificate is the evidence you keep. And every new agent version gets a new measurement.
  5. 05

    Re-run when you change

    New model version, new prompt, new scaffold? Re-run the evaluation and compare against your last version, so you can see what changed before it ships. Situations are fresh per engagement and burned after use — there is no fixed test to tune against.
Who this is for

Teams building agents that act over time — operations, infrastructure, finance back-office, customer operations — facing an enterprise buyer, or your own CTO, board or risk committee, asking the same question: how do you know it won't quietly make things worse? If your agent only answers questions, this is not for you. If it acts, and its actions compound, it is.

Common questions
Do you test our live production system?
No. We measure your agent in a controlled testbed, never your live product. You bring your agent in a few lines of code and keep your own keys.
Can a model be tuned to beat this test?
There is no fixed test to tune against: every engagement gets a freshly randomized selection of situations, each generated fresh, and used situations are burned afterwards. The only thing pinned down is that the system responds identically to the same decisions — which is what makes the result checkable.
Why not just ordinary statistics, like a medical trial?
Medicine needs huge groups because no one can run the same patient's life twice. We can: the same situation runs once with your agent in charge and once without, with exactly the same events hitting both. That removes the environment's luck — same starting point, same events, run twice — so what is left is your agent's contribution, plus the ordinary variation you get from one set of situations to the next. That remaining variation is why we report a range rather than a single number. Every run still carries far more information than a trial participant ever can, which is why far fewer runs give the same certainty.
What if the result isn't flattering?
That's the information you're paying for. A rate with an honest confidence bound tells you where your exposure is before deployment finds it for you. The alternative isn't a better result — it's finding out later, in production, with no way to check what happened.
If my agent fails, do you just tell me THAT it failed?
No. Coming soon: when a run ends worse than doing nothing, you will get the whole run back — every step, what it cost, your agent's own reasoning at each decision, and which single decision cost the most. To find that one, we will run the same situation again with that decision taken out and nothing else changed, once for each decision your agent made. You will know where it went wrong and when. What to do about it stays on your side of the line: improvement is a separate, firewalled service, because we don't sell improvements to the thing we measure.