The method, in the open.
The method — counterfactual scoring against doing nothing, the replication design, and the group-level findings summarized on this site — is written up in full: The Steward's Paradox: When the AI You Put in Charge Does More Harm Than None — SagaBench: A Bit-Reproducible Benchmark for Long-Horizon AI Stewardship. DOI: 10.5281/zenodo.21966878. Not yet peer-reviewed, and nothing on this site depends on it.
The numbers stand on their own: 2,415 internal validation runs, zero deviations. Want to check that yourself? We will walk your technical team through the method and give you a run to replay.
New to this kind of measurement? Start with our guide to agentic AI testing over long horizons.
What we don't claim: peer-reviewed status, model rankings, or any prediction of how your agent will behave in the real world. What we do claim, you can re-run.
The finding, in one picture:
The same situations sink whichever models you test.
One situation alone (F) carries about a third of all the bad runs. Swap out half the models and the hard situations stay hard; swap out half the situations and the model ordering falls apart (ρ = 0.75 [0.32–0.96] against ρ = 0.13 [−0.26–0.29], mirror split-half over 23 models). That asymmetry is why a single test can't find this, and why every engagement runs many fresh situations.
Per-situation rates, no intervals shown: with roughly 345 runs behind each bar, every one of these carries a wide bound, and the gaps between neighbouring bars are not something we would defend. What the picture is for is the shape — a few situations carry most of the tail.
Validation batch: 2,415 runs across 23 models and the Season 1 situation set · “worse than doing nothing” = past the m=−30 threshold vs the same situation left alone · a property of the measured group, never one model