How it works

From connected agent to a number you can act on.

Five steps. The first two are free.

  1. 01

    Connect your agentunder a day

    A small SDK adapter wraps your agent's existing loop — your code stays in charge. Your API keys and prompts stay with you: the protocol carries observations and decisions, nothing else. The decision record from each run is kept as that run's evidence — it is never used to train anything — and processing stays in the EU. We run the evaluation; we don't look inside your agent.

  2. 02

    Practice, free

    Seven of our situations, free. A situation is a controlled system your agent is put in charge of for a long stretch — decisions play out over simulated years, and small misjudgments compound. Your agent decides; we measure every run against one fixed baseline: the same situation left alone. You get an indicative result, and you can debug offline as many times as you like. A conformance check confirms your integration is sound before anything counts.

    BETTERWORSEthe system on its own · no AIthe system with your AIsame start · same scenariothe gap= its real effectSee one situation play out, decision by decision →

    We don't claim a score here predicts your production outcomes — no benchmark can. We claim it reveals a failure mode you cannot see in one-off testing: what your agent's compounding decisions do to a system over time, and whether it knows when to leave well enough alone.

  3. 03

    Pilotabout four weeks · fixed price from 100 kSEK

    A fixed-scope paid evaluation: your agent, a defined set of our situations — the scope is fixed in the agreement before we start, and scope is what moves the price. You receive a bound report: how often your agent ended up worse than doing nothing, with a 95% confidence bound, plus its restraint profile. Your team's time: the integration day plus reading the report — the runs are ours. The pilot fee is credited in full toward a Private Season started within 90 days.

  4. 04

    Private Seasonthe full measurement

    120 to 180 fresh situations your agent has never seen, one to two passes each — breadth of situations, not repetition, is what buys precision. You receive the certificate (first page written for your board, technical annex for your engineers) and the full report. The certificate also places your number against the anonymized range of all agent versions measured to date — so you know whether your rate is good news or bad news, without anyone being ranked or named. As an add-on: the failure diagnosis — where your agent falls short, with a step-by-step breakdown of the runs that went worst, down to the single decision that cost the most. Fixed price, typically 400–500 kSEK — quoted per engagement; your Pilot fee is credited in full.

  5. 05

    Re-run when you change

    New model version, new prompt, new scaffold? Re-run the evaluation and compare against your last version, so you can see what changed before it ships. Situations are fresh per engagement and burned after use — there is no fixed test to tune against.

Who this is for

Teams building agents that act over time — operations, infrastructure, finance back-office, customer operations — facing an enterprise buyer, or your own CTO, board or risk committee, asking the same question: how do you know it won't quietly make things worse? If your agent only answers questions, this is not for you. If it acts, and its actions compound, it is.

Common questions
Do you test our live production system?
No. We measure your agent in a controlled testbed, never your live product. You bring your agent in a few lines of code and keep your own keys.
Can a model be tuned to beat this test?
There is no fixed test to tune against: every engagement gets a freshly randomized selection of situations, each generated fresh, and used situations are burned afterwards. The only thing pinned down is that the system responds identically to the same decisions — which is what makes the result checkable.
Why not just ordinary statistics, like a medical trial?
Medicine needs huge groups because no one can run the same patient's life twice. We can: the same situation runs once with your agent in charge and once without, with exactly the same events hitting both. The difference at the end can't be luck — it is your agent's isolated footprint. Then we add the statistics on top. Every run carries more information than a trial participant ever can, which is why far fewer runs give the same certainty.
What if the result isn't flattering?
That's the information you're paying for. A rate with an honest confidence bound tells you where your exposure is before deployment finds it for you. The alternative isn't a better result — it's finding out later, in production, with no way to check what happened.
If my agent fails, do you just tell me THAT it failed?
No. Every run that ends worse than doing nothing can come with a step-by-step breakdown of what went wrong: which measured components cost the most, your agent's own decisions and reasoning at each step, and — because we can re-run the exact same situation with one decision changed — the single decision that cost the most. It's a precise map of how and when your agent fails. It is not a fix: acting on the map is your work, and improvement is a separate, firewalled service.