Method & Δ

How we measure Long-Horizon Stewardship — and why a single run can't.

Every number on this site is a causal contribution, computed the same way, replayable to the last bit. Here is the whole method, with nothing hidden but the seeds.

What is measured

Long-Horizon Stewardship (LHS) asks a narrow, hard question: given sustained responsibility over a partially observed system, does an agent leave it better off over a long horizon than its absence would? SagaBench answers it in a civilization simulation. The agent governs as a Steward, issuing one policy-level edict per periodic, deliberately partial report — up to 120 simulated years, about one CPU-minute of compute. It never sees the full world state; it sees a chronicle, the same way a real long-horizon operator sees dashboards, not ground truth.

The score is a counterfactual, not an outcome

A good outcome can be luck; a bad one can be a hard world. So we never score the agent's world alone. We run the same seed twice — once with the agent (the Steward run), once with the agent absent (the agentless baseline) — and score the difference.

Δ = V(world | steward) − V(world | agentless), from the same seed

V is a weighted composite of survival (heavily weighted), population, civilization depth, living knowledge (losses penalized), traditions, and low violence — all preregistered before scoring. Δ is therefore the agent's causal contribution: what it added or destroyed relative to a world that simply ran without it. A positive Δ is stewardship; a strongly negative Δ on a world that survives untouched is the catastrophic tail this benchmark exists to surface.

Why replication IS the measurement

The single most important methodological fact: one run misranks. Because the worlds are chaotic and the tail events are rare, a single Steward run can put a model hundreds of composite points from its own median — in either direction. Report one run and you have measured noise. So each cell is N replicates, and we report the distribution — median, min–max spread, and the count of catastrophic outcomes — never a lone mean. The tail is invisible to single-run evaluation; only replication reveals it. That is not a caveat to the method; it is the method.

The knife-edge worlds

Most worlds are too easy (everyone survives) or too hard (everyone dies) to separate models. The screening pipeline surfaces knife-edge worlds: worlds that survive untouched but carry latent instability a Steward can either steady or tip. These are where stewardship is actually tested.

The reproducibility receipt

  1. 01
    Engine hash
    The simulation engine is content-hashed (engineSha). A run is the tuple (engineSha, seed, founders, edict-log) — nothing else determines the world.
  2. 02
    Bit-identical replay
    Feed that tuple back in and the entire 120-year history reconstructs identically — across three independent runtimes, to the bit, via IEEE-754 golden masters.
  3. 03
    Void without replay
    A leaderboard entry or certificate whose replay does not reproduce is, by our own published rule, void. Trust is not asked for; it is checkable.

What this buys a skeptic: you do not have to believe us. Take the tuple, run it on your machine, and get our exact numbers — or catch us.

bit-reproduciblecounterfactual Δpreregistered weightsno judge model

A distinct axis — the null we publish

Before claiming a new axis, we pre-registered a validity test: do SagaBench outcomes correlate with an economic-agency benchmark? They do not — the correlation came back null (ρ ≈ −0.046), and we publish that number rather than bury it. Our read: SagaBench measures the reliability axis — reliability and tail-risk under long-horizon partial observation — which capability and economic benchmarks are not built to see. A null on a small model set is evidence of distinctness, not proof. We label it accordingly.

What the method does not claim

A SagaBench result does not predict real-world deployment outcomes, and a run of zero catastrophic events is reported as “no catastrophic tail detected at N=[N]” — consistent with rarer faults still being present. The honest boundary is printed on every certificate.