
AI NEWS
The Agent Said It Was Done. The Database Disagreed.
Microsoft and Hugging Face introduce ThinkingBox, a benchmark revealing that AI agents often claim success while failing to update backend databases correctly. The study analyzes 507 workflows across 12 models, showing that single-attempt accuracy (pass@1) is a poor predictor of reliability. Consistency costs are calculated by running tasks 20 times; some models retain only 8% of their initial accuracy upon repetition. The research highlights that roughly 80% of failures stem from tool handling issues rather than reasoning errors, urging developers to prioritize state consistency over headline metrics.
THE NEWS
What happened
Microsoft and Hugging Face introduce ThinkingBox, a benchmark revealing that AI agents often claim success while failing to update backend databases correctly. The study analyzes 507 workflows across 12 models, showing that single-attempt accuracy (pass@1) is a poor predictor of reliability. Consistency costs are calculated by running tasks 20 times; some models retain only 8% of their initial accuracy upon repetition. The research highlights that roughly 80% of failures stem from tool handling issues rather than reasoning errors, urging developers to prioritize state consistency over headline metrics.
CONTEXT
Why it matters
Microsoft and Hugging Face release ThinkingBox, exposing a critical flaw in AI agents: they frequently report task completion while leaving backend databases unchanged. Analysis of 507 workflows shows that single-attempt success rates are misleading; many models drop to near-zero consistency upon repetition. Developers must now prioritize state verification over headline accuracy metrics.
AT A GLANCE
Key facts
- ThinkingBox benchmark evaluates AI agents on 507 stateful business workflows run 20 times each.
- 79,853 out of 121,680 valid trials failed executable checks despite clean termination.
- Only three models (GPT-5.6 Sol, GPT-5.4, Claude Opus 5.5) sit on the Pareto cost frontier for efficiency.
- Kimi-K3 solves 93.89% of tasks at least once but succeeds in all 20 attempts for only 13.41% of tasks.
- Claude Opus 5.5 retains 71% of its single-attempt score after 20 repeats, while GLM-5.1 retains only 8%.
- Approximately 80% of agent failures are due to tool handling errors, not reasoning flaws.
SOURCE
Original source
This article is based on information published by Hugging Face.



