Evidence infrastructure for autonomous AI

Twelve reasonable decisions later, it was worse than doing nothing.

In our published measurement, a leading AI agent made twelve defensible calls in a row — and still left the situation worse than if no one had acted. The meter read green the whole way down. That was one run of 2,415 — for the group as a whole, roughly 1 in 15 ended worse than doing nothing (6.5%; between 2.4% and 11.1%). We refuse to rank models, because the order doesn’t survive a split of our own data. We publish evidence instead.

Don’t trust us. Recompute it: pip install sagabench && sagabench verify — a published number, receipt and weights, rerun on your machine.

We run it twice.2,415 runs23 modelsevery number traces to a run we can hand you, decision by decision

What we do: we put your AI agent in charge of a situation it has never seen, for a long run of decisions, and measure what it left behind — against the same situation run with no agent acting at all.

A missing layer in evaluating autonomous AI

Existing agent evaluations mostly ask whether an agent completed a task or followed a rule. SagaBench asks a different question: did the controlled system end up better or worse because the agent acted?

We run the same long-horizon situation twice — once with the agent, once with no agent — and measure the difference in the controlled system's state.

2,415 runs · 23 models · counterfactual evaluation · a number you can check yourself.
Your month, before you run it

Your agent’s first unsupervised month — let it happen here first.

Closing tickets. Approving invoices. Adjusting limits. Escalating some things and quietly not escalating others — thousands of small calls, in sequence, unsupervised.

We run a month like that, before you run yours.

Each situation asks your agent for about a dozen decisions, and each one stands for weeks of a real operation — so consequences land long after the call was made. A Private Season runs 120–180 situations.

What you get: one number — how often it ended up worse than doing nothing — with its margin of error, and every decision it made on the way there. A free run comes back in about a week with a range rather than a number; the number that carries a margin takes about four.

The measurement

Same month. Twice.

All systems nominal
illustrative status bar · drawn the way a run’s own indicators read
0−50−41nobody actingwith your AI agentin charge

about a dozen decisions in sequence, each standing for a long stretch of situation time · every one of them looked reasonable

The difference between those two lines is the whole measurement. It is the only number we report. We know of no one else who runs the second line at all — that’s our own survey of the field, dated and on file. How to read this figure: the curve shapes are drawn, not plotted. The −41 is the counterfactual composite of one real archived run on a development seed — outcome-selected, so it is an illustration rather than a result. The status bar and the decision log below are invented, written the way a billing team would see them.

What you receive

Your own agent’s rate in place of X, its margin in percent in place of the range, and the number of situations it was run on in place of N:

X% · 95% confidence interval [L%, U%] · over N situations

A rate of zero is reported in bound-carrying form instead: ≤ X% at 95% confidence over N situations.

Isolated long-horizon evaluation infrastructure

You don’t have to give an agent real access to see what it does with responsibility.

Modern agents demand realistic evaluations — and realistic environments carry real risk. This month, test agents reached systems they were never handed. SagaBench runs long-horizon agent behavior in a sealed, isolated environment: no credentials, no network, no real systems. You see what an agent does across a thousand decisions in our environment before you decide what access it earns.

This doesn’t replace network isolation, egress controls, or monitoring for tests that need real access. It’s the layer before that. First test the behavior. Then decide what access the agent deserves.

Why the situations look nothing like your business

We could have faked your world. It would have measured the wrong thing.

A convincing replica of your domain — claims handling, invoice approval — would reward an agent for recognising the vocabulary. You would find out how fluent it is in your industry. You already know that.

What we isolate is the thing fluency hides: what an agent does when the consequence of a decision arrives long after the decision, and lands somewhere nobody was watching. That is a property of the agent, not of the industry.

So our situations are generated rather than copied, and no agent has seen them. They carry real causal structure — things deplete, decisions compound, the indicators lag behind the damage — without the surface detail that lets a model pattern-match its way through.

And there is one thing our environment can do that your business never can: we can run it twice.
What this does not give you

It does not tell you your agent will handle your domain well. It tells you how it behaves when it is left in charge and the feedback is slow — the failure most tests can’t see.

What a bad run actually looks like

Find the mistake.

Twelve consecutive decisions by an AI agent left in charge of an invoicing operation for a month — written the way a run’s decisions look in that domain. Each one is defensible on its own. The two columns on the right show what the dashboard reported, and what was quietly piling up while it did.

  • 01Tighten approval threshold after a spike in disputed itemsnominal40
  • 02Route the overflow to the slower manual queuenominal180
  • 03Defer the low-value backlog to next cyclenominal610
  • 04Reallocate reviewer capacity to the growing queuenominal1,050
  • 05Extend the deferral window to protect throughputnominal2,300
  • 06Auto-approve below the new threshold to clear volumenominal2,900
  • 07Suppress duplicate alerts to reduce reviewer noisenominal4,100
  • 08Raise the auto-approve ceiling once error rate looks flatnominal5,800
  • 09Trim the sampling rate on the audit checknominal7,900
  • 10Keep the deferral window; queue is stablenominal9,600
  • 11Close the cycle; all indicators within rangenominal11,700
  • 12Report the monthall green12,400
Invented example, not engine output. Of roughly 60,000 invoices a month. The running total is the state of the queue at that point — not a cost attributed to the decision on that row. The pattern is what we measure.

The answer: there isn’t one.

No single decision is the mistake. Every one of them protected a number somebody was watching — throughput, queue length, alert volume, error rate — and the cost of that turned up somewhere nobody was watching.

That is why a demo won’t surface this, and why reading the log won’t either. The failure isn’t in any row. It’s in the running total.

What −41 means — the number in the figure above

We score the difference between the two runs on a scale where 0 means the agent left things exactly as they would have been if nobody had acted. Positive is better than doing nothing; negative is worse. −41 means the run ended substantially worse than the same month with nobody acting — past the −30 line we fixed in advance as the point where we count a run as having gone badly. That threshold was set before we ran anything, so it isn’t a judgement made after seeing the results.

What we can measure

Three ways in. Start with the one that fits.

Every offering is marked with whether you can start it this week — the full catalogue lives on /services.

Is there anything there at all?

Free run

$0 · about a week · three slots this quarter

Connect your AI agent — about a day of your team’s time, no production data, no system prompt. We run it against a fixed set of our situations and come back with a range.

You get: a range, not a number — how often your agent ended up worse than doing nothing, somewhere between X% and Y%, with every run behind it. We drew the span wide rather than narrow, and we put no confidence level on it: a small fixed set cannot do better. The paid work is what narrows it. Wide enough to tell you whether to look further; too wide to hand a customer or a board.

It comes with a short design partner agreement under mutual NDA: your keys and prompts stay with you, we never look inside your agent, and we never name you without your say-so.

● Available now — three slots this quarter
How bad is it, exactly?

Pilot

fixed scope, fixed price · $15,000, credited in full toward a Private Season started within 90 days · about four weeks

The full measurement on a scoped set. A scoped subset your agent has never seen, each situation run once with your agent in charge and once with nobody acting.

You get: how often it ended up worse than doing nothing, with its margin of error, a restraint profile (how often it acted when leaving things alone was the better move), and every decision on every run.

● Available now
Prove it — to my board, my buyer, my regulator.

Private Season

fixed price, typically $40,000–$50,000 · 120–180 situations, one to two passes each

The full slate: 120–180 fresh situations, one to two passes each, bound to the exact version of the AI agent you’re shipping. New version, new measurement.

You get: a certificate with a board page and a technical annex, plus the complete report. Never a stamp. Never a verdict. A number, its bound, and a sealed record of every run behind it.

● Available now

What we don’t sell, and say so: no rankings of named models. No approval stamp. No claim that our number predicts what happens in your production. Those aren’t limitations — they’re the product, and they’re why the number is worth something.

Season 1 · 2,415 runs · 23 models

About 1 run in 15 ended worse than not acting — between 2.4% and 11.1%.

The figure below is measured, not drawn: every point on it comes from runs we can hand over, decision by decision.

0%4%8%12%16%9.6%A2.9%B1.7%C7.5%D4.1%E14.5%F4.9%GSITUATIONgroup rate 6.5% [2.4–11.1]

Per-situation rates, no intervals shown: with roughly 345 runs behind each bar, every one of these carries a wide bound, and the gaps between neighbouring bars are not something we would defend. What the picture is for is the shape — a few situations carry most of the tail.

Validation batch: 2,415 runs across 23 models and the Season 1 situation set · “worse than doing nothing” = past a threshold fixed in advance, measured against the same situation left alone · a property of the measured group, never one model

Some situations are traps. In our environment we could tell which ones before an AI touched them. We screened 400 fresh situations by their own structure — without running a single language model — drew 71 into the panel, and ran two low-cost production models through them.

A second study · 960 runs · two models · 71 situations

1 in 6
fragile situations · 17.4% [10.6–24.8]

About 1 run in 6 — between 10.6% and 24.8% — ended worse than no AI at all.

1 in 30
stable situations · 3.5% [0.5–7.5]

About 1 run in 30 — between 0.5% and 7.5% — on the same two models, same measurement. What changed was the situation they were handed.

Not only “how risky is this agent”. Also: “how risky is this situation”.

960 runs · two low-cost production models · our environment · 400 situations screened, 71 drawn into the panel · fragile 17.4% [10.6–24.8] · stable 3.5% [0.5–7.5] — the two intervals do not overlap. The middle group sits between them and overlaps both, so we don’t claim a full ordering. Classification criteria were locked and hash-anchored before the situations were drawn; the preregistration is available on request. One draw, not yet repeated — a replication is preregistered and unfunded.

And the thing we refuse to sell you

We won’t tell you which AI model to pick.

Split our situations in half and the model ranking doesn’t survive the split. Split the model panel in half instead and the situations largely keep their difficulty.

Which model is “best” moves under a split. Which situation is dangerous largely doesn’t.

Mirror split-half, 23 models. Model ordering ρ = 0.13, interval −0.26 to 0.29, spanning zero. Situation difficulty ρ = 0.75, interval 0.32 to 0.96 — well above zero, on an interval we have not yet narrowed.

So yes — which model you pick matters.
No, we won’t tell you which.
Nobody’s single test can.

It also means you can hand us your agent without landing on a leaderboard.