Private Season

Know your model's tail before your customers do.

An independent, bit-reproducible measurement of your agent's long-horizon reliability. Not a safety badge. Something a hostile auditor can re-run and a procurement lawyer can approve.

Who this is for

If you deploy agents on long-horizon missions — budgets, codebases, operations, anything measured in days-to-months rather than minutes — you have no credible measure of what they do over the long haul, or of their catastrophic tail. Minute-scale evals are blind to exactly the failures that end things. A Private Season closes that gap with a measurement you can publish.

The strongest thing your certificate can say

“[Your model]: no catastrophic tail detected at N=15” — zero civilization-ending outcomes in 15 replicates on each of [M] hidden knife-edge worlds. 95% upper confidence bound on the true per-world catastrophe rate: 20.4%. Every run replayable bit-for-bit.

— and the only thing we will let it say.

What you get

A signed certificate (PDF + machine-readable JSON) served from a permanent verify URL — per-world reliability/tail profile, the counterfactual Δ distribution, and the leaderboard position.
A replay annex — the (engine-hash, seed, decision-log) triple for every replicate, so any third party can re-execute and confirm each number without trusting us and without ever seeing the sealed worlds.
Citation rights — a licensed sentence and the bounded headline claim above, yours to publish, with the honest boundaries preserved.
The honest boundary (printed on page 1, unmodifiable)

The certificate does report an independent, bit-reproducible measurement of long-horizon reliability, including the observed catastrophic-tail frequency at N replicates. It does not certify your model “safe,” guarantee the absence of faults below the detection threshold, or predict deployment outcomes. A zero result is reported as “no catastrophic tail detected at N=[N]” — never “safe,” “passed,” or “guaranteed.” This is what makes every other number on the document trustworthy.

The confidence ladder

No finite number of clean runs proves a rate of zero; higher N buys a tighter bound, not a stronger adjective.

TierReplicates N95% upper bound on true catastrophe rate (at zero events)
Private Season1520.4%
Enterprise408.8%
Enterprise+1003.7%

How it runs

High-touch and scoped to a statement of work: we agree the model, scaffolding, and N; run the replicates on the sealed holdout worlds on our hardened infrastructure; issue the certificate and replay annex. Your model never trains on the worlds; the holdout seeds are never sold, at any price, under any NDA.

Want to try before you buy?

The public leaderboard and the verify flow let you replay a published result on your own machine today — so you can see the method work before a Private Season measures your model on the sealed set.