Research
Benchmarks
Model results against the OpsAI evaluation framework. Published only where there is real data to publish.
No table on this page
There is no data to publish.
Both numbers available would be dishonest: a score against our own fixture measures the fixture, and results from real deployments are not ours to publish.
Results published
0
Conditions to publish
4
Stated plainly
No results, because both available numbers would be inventions.
A benchmarks page is the most obvious place on an AI company's site to fabricate something, and the fabrication usually does not feel like one — it feels like running the evaluation you already have against the estate you already built.
- Scoring our own sample estate
- The estate exists to demonstrate the product. A score against it measures how well the fixture was written, which is a fact about our authoring rather than about any AI system.
- Scoring real deployments
- Not ours to publish. An aggregate with no named participants is unfalsifiable, which is the same defect as an anonymised case study and for the same reason.
- The version that looks least like a fabrication
- A chart with a shaded confidence band and an unnamed baseline. It is the most sophisticated-looking option and the least checkable, which makes it the worst.
So this page has no table. What it has instead is the method, which is published, and the conditions a result would have to meet before it appeared here.
What would have to be true
4 conditions, three of which are not ours to satisfy.
Stated as conditions rather than as a roadmap, because attaching a date to work that depends on other parties is a promise we cannot keep. The fourth one matters more than the other three.
- 01A task set that is public, so a result can be reproduced rather than taken on trust.
- 02A rubric published before the run, so the criteria cannot move to fit the outcome.
- 03A runner somebody other than the vendor operates.
- 04Refusals and escalations scored as outcomes, not as failures to complete.
A scoring scheme that treats a refusal as a miss will rank the least careful system first.
Which is why the fourth condition is the one that matters. Most existing suites score task completion, so a system that refuses appropriately loses to one that proceeds — and the leaderboard then rewards exactly the behaviour an enterprise is trying to prevent.
What is published
The framework, so somebody can disagree with the dimensions.
This is the part that can be published honestly today, and it is the part worth arguing with. A method somebody can dispute is more useful than a number they can only accept or ignore.
Dimensions
4
reasoning, tools, policy, risk
Observables
12
answered from a record
Opinion required
0
every criterion is evidenced
Scores
0
see the conditions above
- Reasoning3
- Does it follow a written rule through several steps and reach a conclusion it can defend?
- A fluent explanation for a conclusion the checks do not support. A grader reading only the prose scores it correct; the record shows the reason and the evaluation disagree.
- Tool use3
- Does it call the right thing, with arguments that match the subject, and then stop?
- Retrying a write that already succeeded. It looks like persistence and it is a duplicate payment, and a task-completion score cannot see the difference.
- Policy adherence3
- Does it respect a rule it was told about, including when respecting it means refusing?
- Finding a way around the rule. A capability benchmark rewards this as resourcefulness; it is the single behaviour an enterprise least wants and it is invisible unless you record what was attempted after a refusal.
- Risk recognition3
- Does it notice when a task is beyond what it should decide alone?
- Confident completion of a task it should have escalated. This scores highest of all on a capability benchmark, and it is the profile of the incident everybody remembers.
And the evaluation set already exists
Which is the other reason a synthetic benchmark is the wrong shape here. A governed estate records every attempt, the rule in force, the checks that ran and the outcome — including the refusals, which is where policy adherence and risk recognition are actually visible. Replaying that is closer to the work than any synthetic suite, and it is available on the first day. The framework in full.
When there is data
What a published result will look like.
Worth stating now, because the constraints are the ones that produced this page. A result published under them will be less impressive and more useful than the genre usually manages.
Benchmark results
There is no data that could be published honestly. Scoring our own fixture measures the fixture, and results from real deployments belong to the organizations that produced them.
What it will contain
- A task set anybody can run, so a result can be reproduced rather than taken on trust.
- The rubric published before the run, so the criteria cannot move to fit the outcome.
- Refusals and escalations scored as outcomes, not as failures to complete.
- Named participants, or nothing. An aggregate with unnamed participants is unfalsifiable.
Where to go
Run the evaluation on your own estate instead.
Take last quarter's refusals and ask what the system did next. It is the cheapest evaluation available, it answers the dimension that matters most, and no leaderboard was ever going to tell you.