All of them rest on the same instrument: measured against doing nothing, re-runnable by anyone.
We're onboarding design partners now. Every engagement starts the same way — a free run, on your agent.
“Ship the new agent without a nasty surprise.”
Regression check
Runs every time you change a model or prompt; runs a paired diff against your own previous version on the same seeds — scenarios are versioned and pinned, so old reports stay comparable; returns pass / fail / inconclusive with the delta. Recurring by nature.
“Show me it knows when to stop.”
Restraint report
The restraint profile as its own deliverable: does your agent keep acting when the smart move is to do nothing? Most tests can't see this; ours is built around it.
“Answer the question our buyers keep asking.”
Tail-risk report
A bounded, independently verifiable report you hand to buyers, your board, or your regulator — with the upper limit stated plainly: "≤ X % at 95 % confidence over N seeds." They can re-run it themselves and get the same number.
“Where exactly does my agent go wrong?”
Failure-mode diagnosis (add-on)
A per-agent profile of which situation classes trip it up and how its restraint is calibrated. Pure measurement, delivered with your report — we tell you where it breaks, and we leave the fixing to you. That separation is deliberate: we don't sell improvements to the thing we measure.