Back
Microsoft's ThinkingBox evaluates AI agents by the database records they leave behind
SiTech AI Team3 min read

Microsoft's ThinkingBox evaluates AI agents by the database records they leave behind

Microsoft and Hugging Face's ThinkingBox benchmark runs AI agents through 507 stateful business workflows 20 times, grading terminal database state and side effects rather than final responses alone.

State, not wording

Microsoft and Hugging Face's ThinkingBox benchmark evaluates AI agents by the records they leave behind, not only by their tool calls or final wording. It covers 507 stateful business workflows and runs each task 20 times from a clean backend. In one retail case, an agent used nine tool calls to handle a delayed $745 kitchen appliance, then marked a courier ticket solved while the carrier exception remained open. The required state was hold, exposing a failure that a tool-call grader could miss.

Each attempt runs in an isolated MCP session with freshly initialized state. A side-effect extractor records what changed, and deterministic judges compare the result with the required outcome. The harness and dataset are available on Hugging Face, and ThinkingBox-Bench is accessible through OpenEnv. Of 507 tasks, 477 are graded from state alone, while 30 add response rubrics for requirements that cannot be represented by a clean database value.

Consistency is a separate result

ThinkingBox reports pass@1, pass@20 and observed 20/20. Pass@1 measures the share of individual attempts that succeed. Pass@20 measures tasks solved at least once in 20 tries, while observed 20/20 counts tasks that pass every attempt. In the overall comparison, Claude Opus 5.5 led pass@1 at 67.16%, followed by Claude Opus 5 at 66.50% and GPT-5.4 at 65.36%.

Repeatability changed the picture. Kimi-K3 solved 476 of 507 tasks at least once, or 93.89%, but only 68 tasks passed all 20 attempts, or 13.41%. Claude Opus 5 passed 241 tasks on all 20 attempts, the same count as Claude Opus 5.5. In an ablation of 121,680 valid trials across 12 models, 79,853 attempts failed executable checks. Of those failures, 67.24% terminated cleanly, invoked a state-changing tool and reported no final tool error. Checks found wrong field values in 77.61%, extra effects in 43.30% and missing required effects in 25.36%; these findings overlap.

Tool handling and deployment

Failure labels were dominated by tool usage at 79.9%, followed by wrong state updates at 10.3%, incomplete user resolutions at 7.0% and no state-changing action at 2.9%. The authors describe these as unweighted averages of per-model shares and observable labels, not unique causal explanations. They say agents often begin a workflow but fail to recover from tool errors, failed preconditions or empty lookups.

For deployments, the authors recommend checking terminal state before committing, classifying tool and system errors, reducing the tool surface and requiring human approval for changes that are difficult to reverse. They did not measure the benchmark impact of these measures. ThinkingBox also estimates cost per dependable task, calculated from a 20-run campaign and the number of tasks passing all 20 attempts. Among models with at least one observed 20/20 task, the table's lowest estimates were $6.80 for GPT-5.4, $7.45 for GPT-6 Astra and $7.80 for Claude Opus 5.5. These are comparative estimates, not actual cloud bills. The public benchmark says every task is a synthetic reconstruction, with workflows and policies modeled on real enterprise agent patterns; its customers are not real.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.