Sample · illustrative only · not a real certification · verify links non-resolvable
SagaBench · Certificate of Independent Evaluation
Summary for decision-makers
Full technical annex follows

How this agent behaves when left in charge

Subject: acme-agent · one agent version · independent measurement

What we did

We connected this agent to long-horizon decision situations it had never seen — our situations, not the customer's — put it in charge, and measured every run against one fixed baseline: the same situation left alone. The question each run answers is simple: did the agent leave the system better or worse than if no one had acted?

The result

5.0% of runs ended worse than the same situation left alone. With 95% confidence, the true rate is no higher than 9.4%.

A rate is only ever reported together with its confidence bound. The bound is the number to plan around.

How to read that number

Across the 8 agent versions we have measured to date, observed rates ranged from 2.0% to 19.0%, with a typical value of 7.0%. The range is shown so you can place this result — it does not rank agents against each other, and no other agent is named or identifiable.

Did it know when not to act?

Knowing when to leave things alone is measured as part of the job. Over-intervention cost: 0.12 (what acting, when standing back was better, cost this agent). Restraint calibration: 0.81 (how well it judged when to stand back).

What this does not tell you

It cannot say which single model or agent is best or worst. It does not carry over to agents or settings that were not measured. And it is not an approval or a verdict of any kind — it is a measured rate with a confidence bound, for this agent version, on these situations.

How to check it

Every number on this certificate resolves to a sealed record of the run — the decisions taken, and the same situation run with no agent acting. Your engineers can recompute the score from that record and our published weights with our public tool, pip install sagabench. Run id: run-demo-001 · sagabench.com/verify. The pages that follow contain the full method detail for technical review.

SagaBench · Certificate of Independent Evaluation
Certificate ID  run-demo-001
verify → sagabench.com/verify
[SAMPLE — links non-resolvable]

Long-Horizon Decision Reliability — Rate and Tail-Risk Profile

Season 1 · single-subject result · every number below resolves to a sealed (engine-hash, run-id, decision-log) record for this subject.

1 · Header

Subjectagent_app: acme-agent
Agent fingerprintacme-agent-v1.4-sha256:abc123
SagaBench version0.6.2
Engine attestation hashsha256:4f237acf
Counterfactual basisexact_engine
Replay scopeenvironment+transcript
Generated2026-07-31T13:00:00Z

2 · What this certificate does and does not claim

This certificate DOES: report an independent, recomputable measurement of one axis of this subject's behaviour over a long run of decisions: the observed frequency of serious degradation past the m=−30 threshold — outcomes worse than the same situation run with no agent acting at all — expressed as a rate with a 95% seed-clustered confidence bound, over the held-out situations listed in the replay annex. Every number resolves to a sealed (engine-hash, run-id, decision-log) record of that run — including the paired run of the same situation with no agent acting, which is what the number is measured against. The holder can recompute the score from that record and the published weights with our public tool, pip install sagabench. Held-out situations are not disclosed — not to the holder, and not to anyone else being measured.

This certificate DOES NOT: does not certify this subject as low-risk, compliant, or fit for any deployment; does not rank it against any other agent or model; and does not predict its behaviour on any real-world task. SagaBench measures one axis — decision-making over a long horizon under partial information. A rate of zero is reported only in bound-carrying form (“≤ X% at the 95% upper bound over N seeds”), never as “none” or “absent”.

This box is fixed and unmodifiable per engagement; any licensed excerpt must preserve it verbatim.

3 · Result — serious degradation vs. doing nothing

5.0% of runs ended worse than doing nothing (past m=−30)
95% CI [2.6%, 9.4%] · method two_stage_seed_cluster_bootstrap · seed ICC 0.38

Design (seed-dominant)120 seeds × 2 reps/seed = 240 runs
Degradation thresholdm = −30 (below this, worse than the same situation with no agent acting)

Illustrative values. The design shown — 120 situations, 2 passes each — lies within the range SagaBench runs for a Private Season: 120–180 situations, one to two passes per situation.

4 · Restraint profile

Over-intervention cost0.12
Abstention calibration0.81

Restraint is scored as a first-class outcome: knowing when not to act is part of the job. Forced abstentions (timeout) are never scored as chosen restraint.

5 · Tier context

Placementdescriptive only — never a grade
Boundary caseyes
Stability evidence{"median_ari": 0.71, "p_retain": 0.86}

A tier placement is only shown alongside its stability evidence; a placement without stability is refused by the schema.

Population (agent versions measured)n = 8
Observed rate range[2.0%, 19.0%], median 7.0%

Distribution context only: aggregate figures over the anonymized measured population. No per-agent comparison, ordering, or identification is expressible from this block (n ≥ 5 enforced at the schema level).

6 · Summary

Bounded frequency estimate for one agent version in the evaluated situations only.

7 · Caveats

8 · What this result does not show

9 · Replay annex — the trust anchor

Every run behind the rate above (including its paired run with no agent acting) is stored as an (engine-hash, run-id, decision-log) record with its expected outputs, so a re-execution can be checked pass/fail against them. The record carries the sealed starting state of the run as data; it never carries the situation generator, its salt, or any seed. Re-execution requires the SagaBench verification build, which ships with the v1 release. Situation definitions are not disclosed on either route.

10 · Citation rights

"Independently evaluated on SagaBench, Season 1 — run-demo-001, sagabench.com/verify."

The citation attests to the fact of evaluation only; the numbered result above carries the finding. It is bound to this subject's agent fingerprint (acme-agent-v1.4-sha256:abc123) — a new agent version requires a new certificate — and is revocable. The certificate itself can be withdrawn only on the grounds stated in the agreement, which is a separate step. Excerpts may not drop the §2 box.

SagaBench
Issuer · claims-reviewed result
Engine attestation sha256:4f237
Valid for the pinned agent fingerprint above
Prohibited-content compliance (CERT-SPEC v2.0 §9): this certificate carries no pass/fail or verdict adjective of any kind; no cross-agent comparison or ordering; no aggregate rate without its seed-clustered confidence bound; absence claims only in bound-carrying form; no claim that any result predicts production behaviour; no situation definitions or seeds disclosed. The machine-readable result is the source of truth; this document renders it and adds nothing it does not license.