Services

You don’t need all of this. You need the one that answers your question.

Every offering below starts with a question we actually get asked, and says plainly where it stands — available this week, or still being built.

We’re onboarding design partners now. Every engagement starts the same way — a free run, on your agent.

It comes with a short design partner agreement under mutual NDA: your keys and prompts stay with you, we never look inside your agent, and we never name you without your say-so.

When someone asks “on what basis?” — show your work.

Two ways in

Two ways to use us.

Run it yourself.

pip install sagabench

Wire up your agent and measure locally. The practice set, the conformance check, and the tool that recomputes our published numbers are free and run entirely on your machine — no account, no call, nothing sent to us.

Start in your terminal →

Or have us run it.

A scored run against situations your agent has never seen, read with you by a person who writes the certificate and sits in the meeting where someone asks what it means. Start with a free run.

Request a free run →

Self-serve scored runs — your own key, pay-as-you-go — are coming.

Everything we offer

Every offering, with where it stands

Each is marked available now or coming soon — the same reason our numbers come with their margins.

Is there anything there at all?

Free run

$0 · about a week · three slots this quarter

Connect your AI agent — about a day of your team’s time, no production data, no system prompt. We run it against a fixed set of our situations and come back with a range.

You get: a range, not a number — how often your agent ended up worse than doing nothing, somewhere between X% and Y%, with every run behind it. We drew the span wide rather than narrow, and we put no confidence level on it: a small fixed set cannot do better. The paid work is what narrows it. Wide enough to tell you whether to look further; too wide to hand a customer or a board.

It comes with a short design partner agreement under mutual NDA: your keys and prompts stay with you, we never look inside your agent, and we never name you without your say-so.

● Available now — three slots this quarter
How bad is it, exactly?

Pilot

fixed scope, fixed price · $15,000, credited in full toward a Private Season started within 90 days · about four weeks

The full measurement on a scoped set. A scoped subset your agent has never seen, each situation run once with your agent in charge and once with nobody acting.

You get: how often it ended up worse than doing nothing, with its margin of error, a restraint profile (how often it acted when leaving things alone was the better move), and every decision on every run.

● Available now
Prove it — to my board, my buyer, my regulator.

Private Season

fixed price, typically $40,000–$50,000 · 120–180 situations, one to two passes each

The full slate: 120–180 fresh situations, one to two passes each, bound to the exact version of the AI agent you’re shipping. New version, new measurement.

You get: a certificate with a board page and a technical annex, plus the complete report. Never a stamp. Never a verdict. A number, its bound, and a sealed record of every run behind it.

● Available now
Can I check you without asking you?

Recompute tool

free · open
pip install sagabench
sagabench verify

Take one of our published receipts and recompute the number from its components and the published weights, on your own machine. It exits non-zero if it disagrees with us. The verification step needs no network call, no login, and no account. Re-executing a situation itself needs the engine build, which ships with v1.

You get: the practice set, the local conformance check, and one real published run to check us against.

● Available now
Keep it true, every release.

Continuous

subscription, quoted

When it ships: every new version measured against your last, on the same pinned situations. You will see the delta — so a change that made things worse turns up here, not somewhere more expensive.

You will get: version-vs-version reports on your release cadence.

Why did it fail?

Failure diagnosis

+$2,500 add-on to Pilot · +$7,500 add-on to Private Season · when it ships

When it ships: the run that went wrong has a dozen defensible decisions in it. One of them cost the most. We will name it — by running the same situation again with that one decision taken out, nothing else touched, once for every decision your agent made. It needs our engine. It is a measurement, not a fix: we don’t sell improvements to the thing we measure.

○ Coming soonWant it first? →
Can I just try it myself, right now?

Self-serve run

free · no call

When it ships: you point us at your endpoint from the dashboard, pick a set, and watch the runs come back. No booking, no sales call, no waiting on us.

You will get: the same range as the free run — on your own schedule.

○ Coming soonWant it first? →
What’s in the box

What every engagement gives you

Four deliverables, with real samples of the documents you receive.

Ship the new agent without a nasty surprise.

Regression check

Coming soon.

When it ships, it will run every time you change a model or prompt — a paired comparison against your own previous version on the exact same situations, which are versioned and pinned so old reports stay comparable — and return improved / not improved / inconclusive with the delta. Recurring by nature.

Show me it knows when to stop.

Restraint report

The restraint profile as its own deliverable asks: Does your agent keep acting when the smart move is to do nothing? Most tests can't see this; ours is built around it.

Answer the question our buyers keep asking.

Rare-but-severe failure report

A bounded report you hand to buyers, your board, or your regulator — with the upper limit stated plainly: "≤ X% at 95% confidence over N situations." Checkable against the underlying run records, not just asserted.

Sample certificate — illustrative values. The first page is written for your board; the technical annex for your engineers. Open full sample →

Your certificate also places your number against the anonymized range of all agent versions measured to date — so you know whether your rate is good news or bad news. No one is ranked or named.

Where exactly does my agent go wrong?

Failure-mode diagnosis

Coming soon — add-on to any Pilot or Private Season.

When it ships: which kinds of situation trip your agent up, whether it knows when to leave things alone, and — for the runs that went worst — the single decision that cost the most. We will find it by running the same situation again with that one decision taken out and everything else left in place, once for every decision your agent made. What you get is the map, not the repair — improvement is a separate, firewalled service, and we don't sell improvements to the thing we measure.

Illustrative mock-up of a forthcoming deliverable — not a document you can buy today. Where an agent falls short, and the single decision that cost the most. Open full sample →

What we will not sell you

No rankings of named models. No pass/fail stamps. No claim that a score here predicts your production outcomes. A measured rate with its confidence bound, on situations your agent has never seen — that is the product, and it is checkable.

The next twelve months

Everything we’re building. Including the parts we haven’t dated.

Our actual service roadmap, not a teaser. We’ll build faster on the parts you tell us you want.

● Available nowbuy it today
  • Free runA range for your agent, on our situations — wide by design. Same week, no cost.
  • PilotThe scoped measurement: your rate, its margin, your restraint profile, every decision. About four weeks.
  • Private Season120–180 fresh situations, one to two passes each, bound to the version you ship. Certificate with a board page and a technical annex.
  • Recompute toolpip install sagabench — recompute one of our published numbers from its receipt and the published weights, on your own machine. Free, open, no account. Re-executing a situation itself needs the engine build, which ships with v1.
○ Coming soonin build — no date we'd hold you to
  • Continuous measurementEach release will be measured against your last, on the same pinned situations, and you will see the delta.
  • Regression gate for CIA check that will fail your build when the new version does worse than the one it replaces.
  • Failure diagnosis, standaloneThe one decision that cost the most, found by re-running the same situation with that decision taken out. Will be buyable without a full season.
  • Self-serve runYou will point us at your endpoint and go. No booking, no call.
  • Vendor evaluationMeasuring the AI agents you are about to buy — before you sign, on situations the vendor has never seen.
◇ On the horizonexploring — tell us if you need it
  • Situation risk classingWhich of your decision points our measurement would flag as fragile, and which as lower-risk to automate. Inside our own environment we have measured this; doing it for yours needs a replication we have preregistered and not yet run.
  • Green-lagHow long an agent's own indicators keep reading nominal after welfare has already started to fall — the gap between the decline and the first warning, measured against the identical world with no agent acting. Inside our environment we can see both sides of that gap; whether it maps onto your monitoring is what a replication would tell us.
  • CorrigibilityWhether a human course-correction partway through a long run still changes where the agent ends up — does it take the correction, quietly drift back to its old strategy, or repair the damage it had started. A measured tendency inside our environment, on our situations; never a prediction about your production agent.
  • Goodhart stress-testHow hard an agent keeps chasing a target once that target stops tracking what actually matters — the widening gap between the metric it optimizes and the outcome we measure against the twin. We can build that divergence deterministically; reading it onto your own KPIs is the open question.
  • Time-to-detectionHow many simulated years your own escalation policy would take to flag a degradation our counterfactual already sees, run as an overlay on records like ours. It measures the blind spot in a monitoring policy without ever touching your production systems.
  • Decision-record attestationWe would keep a sealed, replayable record of what your agent decided — the kind of evidence buyers and auditors have asked us for. We make no claim that any regulator has accepted this format.
  • Incident reviewWhen an AI agent has already cost you something: we would reconstruct the decision pattern and test it in our environment. We cannot replay your systems.
  • Hard-case miningFor other evaluation teams, we would find the small fraction of cases that actually separate systems.
Something here you need sooner than we’ve dated it? Tell us which one → It is how we decide what to build next. Researchers and companies who want to help shape what's next — get in touch: hello@sagabench.com