The ticket is closed. The customer's $745 appliance is still stuck in Nashville.
An AI agent handles a customer service inquiry. Nine tool calls, all well-formed. It reads the refund policy, opens a ticket, documents the timeline. Then it marks the ticket resolved and asks if there's anything else it can help with.
Two problems: the carrier exception is still open, so the correct status was on hold. And the customer never got a real answer. The agent's trajectory looked clean. The database state tells a different story.
That's the gap Microsoft's ThinkingBox benchmark measures. Across 507 stateful business workflows—retail, auto insurance, travel, neobank, consulting—it runs each task 20 times against various models and grades on terminal backend state and side effects, not whether the agent invoked tools or generated plausible text.
Tool calls are not outcomes
Most agent benchmarks check whether the model called the right functions or whether its final response sounds correct. ThinkingBox checks the records it left behind.
The gap is substantial. Across 121,680 valid trials covering 12 models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. The executable checks found:
- Wrong field values in 77.61% of failures
- Unintended extra effects in 43.30%
- Missing required effects in 25.36%
An agent can sound correct while writing the wrong value, changing the wrong record, or creating side effects it shouldn't. Only the database settles the question.
One success is not reliability
If an agent processes a refund correctly once and mishandles it the next four times, you don't have a working refund agent. So ThinkingBox runs every task 20 independent times from identical clean backends and reports three numbers:
- pass@1: share of all attempts that succeeded (how does it usually do?)
- pass@20: share of tasks solved at least once in 20 tries (can it ever do this?)
- Observed 20/20: tasks that passed all 20 recorded attempts (can it always be correct?)
The industry mostly reports pass@1. It reads like a normal capability ranking. Claude Opus 5.5 leads at 67.16% overall. Kimi-K3 is the strongest open-weights model at 57.37%. Domain matters as much as model: Claude Opus 4.6 scores 68.62% on retail but 8.30% on auto insurance.
Then you run it 20 times and ask how much of that score survives.
The consistency cliff
Only three models hold on to most of their single-attempt scores when you demand perfect consistency:
- GPT-6 Astra retains 78% of its pass@1 rate
- Claude Opus 5.5 retains 71%
- Claude Opus 5 retains 71%
At the other end, GLM-5.1, Kimi-K2.6, and DeepSeek-V4-Pro each keep about 8% of their single-attempt performance.
The gap between what a model can do once and what it does every time is the whole story.
Breadth versus dependability
Kimi-K3 has the broadest coverage. It solves 93.89% of the benchmark at least once—476 of 507 tasks. Only 31 tasks defeat it entirely, the lowest count in the field. On retail workflows it leads at 82.24% pass@1, ahead of every proprietary model.
Kimi-K3 is also among the least consistent. Just 68 of 507 tasks succeed in all 20 attempts. That's 13.41%.
Claude Opus 5 inverts this. It solves fewer tasks at least once (79.09%; 106 defeat it entirely) but completes 47.53% of the benchmark on every single attempt.
A newer model does not fix this. Claude Opus 5.5 scores higher than Claude Opus 5 on single-attempt average—67.16% against 66.50%—and solves more tasks at least once. It passes exactly the same number of tasks on all 20 attempts: 241. Half a point of headline accuracy bought zero additional dependability.
Kimi-K3 solves 75 more tasks at least once than Opus 5. Opus 5 solves 173 more tasks consistently than Kimi-K3.
If you're choosing a model for work that touches real records, pass@20 is the wrong column.
What consistency costs
Microsoft measured cost per successful task attempt using undiscounted list rates from OpenRouter, reversing promotional discounts and excluding quantized endpoints. They divided one run's cost by the number of attempts that succeeded:
Cost per successful task attempt = estimated cost for 507 attempts ÷ (507 × pass@1)
This is a comparative efficiency index, not an invoice. But it surfaces the tradeoff between raw capability and deployment economics.
The numbers matter because production systems don't get one shot. Every failure is a retry, a human escalation, or a broken workflow. The model that succeeds 80% of the time at $X per attempt and 20% of the time never costs less than the model that succeeds 60% of the time always.
Failure signatures
ThinkingBox's executable checks capture three classes of failure:
- Wrong field values: the agent wrote to the right record but with incorrect data
- Extra effects: side effects that shouldn't exist (duplicate tickets, unnecessary notifications)
- Missing effects: required state changes that never happened
These overlap. A single failed attempt can exhibit all three. The kitchen appliance ticket from the opening is a missing-effect failure: the status should have been on hold, not resolved.
The benchmark is available as executable test cases. The specific example above is sandbox_external_retail_group1.py:test_case_ST003_006. You can run it, watch it fail, and inspect exactly which field was wrong.
How it works
ThinkingBox runs agents against isolated MCP tool sessions, then grades the terminal backend state. Each task specifies:
- An initial backend state (database fixtures, existing records)
- A natural-language request
- Executable assertions about required end state
The agent gets tool access through the Model Context Protocol. It can read, write, update, and delete records. When it signals completion, ThinkingBox runs the assertions against the actual database state.
No LLM grader. No vibes. Just: did the record end up in the required state?
Run it yourself
The benchmark is available through Hugging Face's OpenEnv. The tasks, fixtures, and assertions are open. You can run any model you want, inspect failures, and contribute additional domains.
This is not a closed leaderboard. It's infrastructure for testing whether an agent does what you need it to do, repeatedly, against state.
What this means for agent deployment
Pass@1 measures whether a model has the capability. Observed 20/20 measures whether you can depend on it.
The industry has spent two years optimizing for the first number. The gap between the two is where production agents actually fail. An 80% pass@1 agent that only achieves 10% perfect consistency costs more to run than a 60% agent with 50% consistency, because every failure is a retry or an escalation.
Breadth and dependability are different dimensions. You can have a model that solves almost anything once (Kimi-K3) or a model that solves fewer things but does them reliably (Claude Opus 5). Picking the wrong one for your use case is the difference between a working system and an expensive retry loop.
The source is here. The paper is here. The data is here. Go run it.