The agent said the ticket was resolved. The database said otherwise. Microsoft's ThinkingBox grades agents on what they leave behind
A tool call is not an outcome. Microsoft just put numbers on how often that gap bites in real business workflows.
What happened
Microsoft and Hugging Face released ThinkingBox, a sandbox and benchmark that grades AI agents on the records they leave in the backend, not on what they say. It covers 507 stateful workflows across retail, auto insurance, travel, neobank IT and consulting, and every task runs 20 times from a clean database.
The opening example is painfully familiar: a support agent makes nine sensible tool calls, then closes the ticket as resolved when the correct end state was on hold, because the courier exception was still open.
The numbers that matter:
- In one ablation, 67.24% of failed attempts still ended cleanly, called a state-changing tool and reported no tool error. Checks found wrong field values in 77.61% of those failures.
- Claude Opus 5 drops from 66.50% on a single attempt to 47.53% of tasks passed on all 20 attempts. Kimi-K3 solves 93.89% of tasks at least once but only 13.41% every time.
- Roughly four in five failures were classified as tool usage problems: failed preconditions, empty lookups and tool errors the agent did not recover from.
- Priced for consistency, the cheapest model per single success was not the cheapest per dependable task.
My take
This is the clearest evidence yet for something I see in client builds. The agent's summary is a claim. The CRM record is the evidence.
If you run agents that touch tickets, deals or bookings, three changes are worth making this week:
- Add a check step after the agent finishes that reads the actual record and compares it to the expected state. Status, owner, amount. If it does not match, route it to a person.
- Classify tool errors. An empty search result and a timeout need different handling, and most agent failures here were exactly that kind of unhandled case.
- Test one workflow 20 times, not once. A demo that passes once tells you the agent can do it, not that it will.
The authors also recommend cutting the tool list to what the workflow needs and requiring approval for changes you cannot cheaply reverse. That matches what works in production. Pick models by how often they get it right every time, not by the best single run.
More posts
- GitHub's ReviewBench shows how to test an AI step before you ship it. Build a golden set for your automationsOct 6, 2026
- Wikimedia found OpenAI agents editing wikis and making millions of API calls. Protect your client's public endpointsOct 6, 2026
- Cohere North 2 gives agents memory, shared skills and token caps. The checklist every agent rollout needsOct 6, 2026
