Built to be easy to try. Priced so trying is never wasted spend.
Throughout: "catastrophe" means a run that ended worse than doing nothing, measured against a stated threshold. It doesn't mean collapse.
- 1.
Free run
Connect your agent, pass conformance, get an indicative result over 7 seeds: your tier placement with stability context, wide bands, honest caveats. Seven seeds can't give you a bounded upper limit — that's what the paid tiers are for — but it tells you whether this is measuring what you thought it was. No cost, no commitment. (Only have a model? See "Connect your agent" — the reference-adapter path is the fastest start there is.)
- 2.
Pilot
A fixed-scope, high-repeat measurement on the situations that matter to you, ending in a bounded report you can put in front of a buyer. Credited in full toward a Private Season started within 90 days (terms in the pilot agreement).
- 3.
Private Season
The deep version: your own catastrophe rate with a stated 95 % upper limit over 120–180 seeds (seed-dominant by design — 120–180 different fresh situations with only 1–2 repeats each; breadth of fresh situations is what pins down a rare-failure rate), restraint profile, failure-mode diagnosis, versioned reports as your agent evolves.
- 4.
Enterprise
Ongoing measurement built into how you ship: regression checks on every release, standing scenario coverage, direct support.
Pilots are scoped to your situations, so scope and price are agreed in a 30-minute call — Talk to us →. They're built for teams putting an agent somewhere a bad month would be expensive. Your first run is free either way.