AI & agents
Microsoft and Hugging Face release the ThinkingBox benchmark
- Published
- Source
- Microsoft and Hugging Face joint tech blog
Summary
Microsoft and Hugging Face published ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their sentences or tool calls. It runs 507 stateful business workflows 20 times each and is reproducible through OpenEnv. In the authors' measurements, even the most dependable models passed only 241 of 507 tasks (about 47%) on every attempt, showing that a single success is not the same as reliability.
Why it matters for our work
It offers a concrete standard for measuring trustworthiness before deploying agents that change real business records. Teams can use it directly to decide where human review must stay in the loop.
Translated from the Korean original. Summaries may be translated and edited. Commentary reflects our perspective; forecasts remain the source’s views.