Microsoft Brings Its Database-Verified AI Agent Benchmark to Hugging Face
The 507-task benchmark checks the records agents leave behind. Its latest results separate finding a solution from repeating it reliably.
Loading page…
The 507-task benchmark checks the records agents leave behind. Its latest results separate finding a solution from repeating it reliably.
Listen to this story
On October 3, Microsoft and Hugging Face made Microsoft’s ThinkingBox evaluation environment and 507-task ThinkingBox-Bench available on Hugging Face. The release tests whether agents leave business systems in a rule-compliant final state, not merely whether their tool calls run; in the authors’ 12-model ablation, 67.24% of failed trials ended without a final tool error. Results across 18 models, with each task run 20 times, show that broad one-off coverage and repeatable success can rank models differently—a distinction for teams evaluating agents for real workflows.
Tasks specify an initial state, user goal, tools and business rules; executable checks catch incorrect records, missing changes and unwanted side effects.
Kimi-K3 solved 476 of 507 tasks at least once but passed all 20 attempts on only 68; Claude Opus 5.5 reached the 20/20 bar on 241 tasks.
GPT-5.4 had the lowest reported cost per dependable task at $6.80; GPT-5.6 Sol had the lowest estimated cost per successful attempt at $0.127.
An AI agent can finish without a tool error and still leave a business record wrong. Microsoft’s ThinkingBox checks that gap against the database itself. In an October 3 joint post, Microsoft and Hugging Face said the evaluation environment and its 507-task benchmark are now available through Hugging Face.
One benchmark example follows a customer whose appliance is stuck at a courier distribution center. The agent makes nine tool calls, checks the refund policy and creates a support ticket. It then marks the ticket resolved, although the unresolved carrier exception requires an on-hold status.
ThinkingBox checks the outcome, not just whether the agent called a valid tool. Each task specifies a starting state, a user goal, available tools and business rules. Executable checks examine the final records for wrong values, missing changes and unwanted extra effects.
The accompanying research paper describes isolated, MCP-compatible sessions connecting agents to software tools. A simulated user supplies private details only when asked. Some tasks also check required properties of the final response, so database correctness is not the only requirement.
In the authors’ 12-model ablation, 79,853 of 121,680 valid trials failed executable checks. Of those failures, 67.24% ended cleanly after a state-changing tool call, without a final tool error.
The authors also classified failed traces by their observable failure pattern. Tool handling accounted for 79.9% of failures, calculated as an unweighted average of each model’s share. Agents commonly attempted the workflow but failed to recover from tool errors, unmet preconditions or empty lookups.
Wrong state updates accounted for 10.3%, incomplete user resolutions for 7.0%, and no state-changing action for 2.9%. These labels describe what happened in a trace; they do not establish a unique underlying cause.
ThinkingBox-Bench spans retail, auto insurance, travel, neobank and consulting workflows. The October post presents results for 18 models, running every task 20 times from an identical clean backend. It reports three measures that answer different questions:
The authors’ results show why those distinctions matter. Kimi-K3 solved 476 of 507 tasks at least once, the broadest coverage tested. But only 68 tasks passed every attempt. Claude Opus 5 solved fewer tasks at least once, yet passed all 20 attempts on 241 tasks.
Claude Opus 5.5 led the single-attempt scores at 67.16%, above Opus 5’s 66.50%. Both nevertheless passed exactly 241 tasks on all 20 attempts. The newer model’s higher average did not increase the number of tasks meeting that consistency bar.
The post also divides the estimated cost of a full 20-run campaign by the number of tasks passing every attempt. GPT-5.4 had the lowest reported cost per dependable task, $6.80, with 128 tasks meeting the bar. Opus 5.5 reached 241 at $7.80.
Pricing individual successes gives a different winner. GPT-5.6 Sol had the lowest estimated cost per successful attempt, $0.127. But its cost per dependable task was $9.76, above GPT-5.4’s. The cheaper individual success did not translate into cheaper repeated success.
Those figures use recorded token consumption and undiscounted provider list rates. They are comparative estimates, not cloud bills or prices for serving a production request. Likewise, observed 20/20 records what happened in these trials; it is not a promise of future success.
The released OpenEnv adapter returns a binary pass-or-fail reward for each finished episode and is designed for evaluation. The authors recommend checking final state, targeting recoverable errors with retries and requiring human approval for hard-to-reverse changes. They have not measured those interventions’ gains on this benchmark.
Loading discussion...
Join the conversation
Explain when broader ability would outweigh repeatable results.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.