About

Independent research, held to its own rules — in public.

SagaBench is an independent effort to measure something the field can't yet measure well: what capable agents do when given sustained responsibility over the long haul.

Why this exists

The failures that matter for long-horizon agents are invisible to minute-scale tests. An agent can look excellent for an hour and end a civilization over a century — and a single run cannot tell you which one you have. SagaBench was built to make that tail measurable, reproducible, and impossible to wave away. We found that every frontier model we tested can, on fragile worlds, end a civilization that survives untouched — a finding only replication reveals. That is the kind of thing that ought to be checkable by anyone, not taken on trust.

What we hold ourselves to

Trustless over trusted.

Every published number ships with a replay package. If you can't re-run it and get our result, it doesn't count — including for us.

Bounded over bold.

We report what the data licenses and no more. A zero result is “no catastrophic tail detected at N=[N],” carrying its N and its confidence bound — never “safe,” “passed,” or “guaranteed.”

Sealed, not secret-for-advantage.

The holdout worlds are hidden so scores measure capability, not memorization — and retired worlds are published so our past is auditable.

Examiner first.

The evaluation is the crown jewel. When a future training product could bend evaluation integrity, we would have neither — so examination and training are separated by role, not only by data, and that wall is published.

How we handle being wrong

We publish our nulls (SagaBench does not correlate with economic-agency benchmarks — stated plainly), we bound our headline findings to exactly the worlds that support them, and we correct in the open. When an external reviewer caught us leading with a finding that was bound to a single world, we fixed the framing and wrote the check into our process so it can't recur. Getting it right matters more than looking certain.

Not affiliated with, or endorsed by, any of the labs whose models appear on the leaderboard. Every claim on this site is replayable without trusting us — that independence is the point.

We are looking for

  • an academic co-author for the methodology paper
  • a technical co-founder
  • first design-partner labs for Private Seasons

Contact: info@sagabench.com