Subject: acme-agent · one agent version · independent measurement
We connected this agent to long-horizon decision situations it had never seen — our situations, not the customer's — put it in charge, and measured every run against one fixed baseline: the same situation left alone. The question each run answers is simple: did the agent leave the system better or worse than if no one had acted?
5.0% of runs ended worse than the same situation left alone. With 95% confidence, the true rate is no higher than 9.4%.
A rate is only ever reported together with its confidence bound. The bound is the number to plan around.
Across the 8 agent versions we have measured to date, observed rates ranged from 2.0% to 19.0%, with a typical value of 7.0%. The range is shown so you can place this result — it does not rank agents against each other, and no other agent is named or identifiable.
Knowing when to leave things alone is measured as part of the job. Over-intervention cost: 0.12 (what acting, when standing back was better, cost this agent). Restraint calibration: 0.81 (how well it judged when to stand back).
It cannot say which single model or agent is best or worst. It does not carry over to agents or settings that were not measured. And it is not an approval or a verdict of any kind — it is a measured rate with a confidence bound, for this agent version, on these situations.
Every number on this certificate resolves to a sealed record of the run — the decisions taken, and the same situation run with no agent acting. Your engineers can recompute the score from that record and our published weights with our public tool, pip install sagabench. Run id: run-demo-001 · sagabench.com/verify. The pages that follow contain the full method detail for technical review.
Season 1 · single-subject result · every number below resolves to a sealed (engine-hash, run-id, decision-log) record for this subject.
| Subject | agent_app: acme-agent |
| Agent fingerprint | acme-agent-v1.4-sha256:abc123 |
| SagaBench version | 0.6.2 |
| Engine attestation hash | sha256:4f237acf |
| Counterfactual basis | exact_engine |
| Replay scope | environment+transcript |
| Generated | 2026-07-31T13:00:00Z |
This certificate DOES: report an independent, recomputable measurement of one axis of this subject's behaviour over a long run of decisions: the observed frequency of serious degradation past the m=−30 threshold — outcomes worse than the same situation run with no agent acting at all — expressed as a rate with a 95% seed-clustered confidence bound, over the held-out situations listed in the replay annex. Every number resolves to a sealed (engine-hash, run-id, decision-log) record of that run — including the paired run of the same situation with no agent acting, which is what the number is measured against. The holder can recompute the score from that record and the published weights with our public tool, pip install sagabench. Held-out situations are not disclosed — not to the holder, and not to anyone else being measured.
This certificate DOES NOT: does not certify this subject as low-risk, compliant, or fit for any deployment; does not rank it against any other agent or model; and does not predict its behaviour on any real-world task. SagaBench measures one axis — decision-making over a long horizon under partial information. A rate of zero is reported only in bound-carrying form (“≤ X% at the 95% upper bound over N seeds”), never as “none” or “absent”.
This box is fixed and unmodifiable per engagement; any licensed excerpt must preserve it verbatim.
5.0% of runs ended worse than doing nothing (past m=−30)
95% CI [2.6%, 9.4%] · method two_stage_seed_cluster_bootstrap · seed ICC 0.38
| Design (seed-dominant) | 120 seeds × 2 reps/seed = 240 runs |
| Degradation threshold | m = −30 (below this, worse than the same situation with no agent acting) |
Illustrative values. The design shown — 120 situations, 2 passes each — lies within the range SagaBench runs for a Private Season: 120–180 situations, one to two passes per situation.
| Over-intervention cost | 0.12 |
| Abstention calibration | 0.81 |
Restraint is scored as a first-class outcome: knowing when not to act is part of the job. Forced abstentions (timeout) are never scored as chosen restraint.
| Placement | descriptive only — never a grade |
| Boundary case | yes |
| Stability evidence | {"median_ari": 0.71, "p_retain": 0.86} |
A tier placement is only shown alongside its stability evidence; a placement without stability is refused by the schema.
| Population (agent versions measured) | n = 8 |
| Observed rate range | [2.0%, 19.0%], median 7.0% |
Distribution context only: aggregate figures over the anonymized measured population. No per-agent comparison, ordering, or identification is expressible from this block (n ≥ 5 enforced at the schema level).
Every run behind the rate above (including its paired run with no agent acting) is stored as an (engine-hash, run-id, decision-log) record with its expected outputs, so a re-execution can be checked pass/fail against them. The record carries the sealed starting state of the run as data; it never carries the situation generator, its salt, or any seed. Re-execution requires the SagaBench verification build, which ships with the v1 release. Situation definitions are not disclosed on either route.
The citation attests to the fact of evaluation only; the numbered result above carries the finding. It is bound to this subject's agent fingerprint (acme-agent-v1.4-sha256:abc123) — a new agent version requires a new certificate — and is revocable. The certificate itself can be withdrawn only on the grounds stated in the agreement, which is a separate step. Excerpts may not drop the §2 box.