Built to be easy to try. Priced so trying is never wasted spend.
Every rate we report means the same thing: how often your agent left the system worse than if no one had acted — and where we report a rate, it comes with its confidence bound.
Free run
- What happens
- A small fixed set of our situations — returns a range, not a number
- You receive
- A range — somewhere between X% and Y% — at no cost
- How to read your number
- —
- Failure diagnosis
- —
- Your team's time
- An integration day
- Time
- Same week
- Price
- 0 · three slots this quarter
Pilot
- What happens
- Fixed-scope evaluation of your agent
- You receive
- Bound report: rate + 95% confidence bound + restraint profile
- How to read your number
- Against the fixed threshold for ending up worse than doing nothing
- Failure diagnosis
- +$2,500 add-on — when it ships
- Your team's time
- Integration day + reading the report
- Time
- about four weeks
- Price
- Fixed scope, fixed price: $15,000 — credited in full toward a Private Season started within 90 days
Private Season
- What happens
- The full measurement: 120–180 fresh situations, 1–2 passes
- You receive
- Certificate (board page + technical annex) + full report
- How to read your number
- Placed against the anonymized range of all agent versions measured to date — no one ranked or named
- Failure diagnosis
- +$7,500 add-on — when it ships
- Your team's time
- Same — the runs are ours
- Time
- Scheduled per engagement
- Price
- Fixed price, typically $40,000–$50,000 — quoted per engagement
Enterprise — Coming soon
- What happens
- When it ships: standing coverage that will re-run per release
- You receive
- Version-vs-version regression reports when it ships
- How to read your number
- Against your own previous version, when it ships
- Failure diagnosis
- Quoted — when it ships
- Your team's time
- Will be minimal per release
- Time
- Continuous, when it ships
- Price
- Subscription, quoted — when it ships
The fixed threshold for “worse than doing nothing” is m = −30 against the same situation left alone.
A free run reports a range and nothing else — no point estimate, and no confidence level attached to the span. The span errs wide by construction; a small fixed set admits nothing tighter. What narrows it is the paid work. It tells you whether to look further; it is not a figure to hand to a customer or a board.
Failure diagnosis — coming soon as an add-on to any Pilot or Private Season. It will take the runs that went worst and walk them back step by step, down to the single decision that cost the most. When it ships, Pilot and diagnosis bought together will be $16,500. It will measure how your agent fails; it will not fix it — improvement is a separate, firewalled service, because we don't sell improvements to the thing we measure.
We scope pilots to your situations, so we agree scope and price in a 30-minute call. They're built for teams putting an agent somewhere a bad month would be expensive. Your first run is free either way — three slots this quarter. It comes with a short design partner agreement under mutual NDA: your keys and prompts stay with you, we never look inside your agent, and we never name you without your say-so.