Skip to main content
AI Interview Question
INTERVIEW GUIDEAI News8 questions3 min readOct 4, 2026

ThinkingBox grades agents on the records they leave

Microsoft and Hugging Face posted ThinkingBox on 3 October 2026. 477 of 507 workflows are graded on state alone. 30 add response rubrics.

ThinkingBox grades agents on the records they leave

Microsoft and Hugging Face posted ThinkingBox on 3 October 2026. The post is by Tuhin Kundu at Microsoft. It grades agents on the database records they leave, not the sentences they generate.

The benchmark is 507 stateful business workflows, each run 20 times. 477 of 507 are graded on state alone. 30 add response rubrics. The harness and ThinkingBox-Bench are on Hugging Face behind OpenEnv. ThinkingBox code is MIT-licensed. The benchmark data is CDLA-Permissive-2.0. The OpenEnv environment ships under OpenEnv's BSD-3-Clause.

An ablation covers 121,680 valid trials across 12 models. 79,853 attempts failed executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. Of those failures, wrong field values are 77.61%, unintended extra effects are 43.30%, and missing required effects are 25.36%. Those state-check findings overlap.

One example is a $745 kitchen appliance case. The failing check is that the ticket status is solved where the required end state is hold. The task id is sandbox_external_retail_group1.py:test_case_ST003_006.

Claude Opus 5.5 leads overall at 67.16% pass@1. Claude Opus 5 is 66.50%. Both pass exactly 241 tasks on all 20 attempts. Opus 5 completes 47.53% of the benchmark on every attempt. That is not a claim that either model reliably finishes enterprise work.

Kimi-K3 solves 93.89% at least once, 476 of 507 tasks. Only 68 of 507, 13.41%, succeed in all 20 attempts. Its overall pass@1 is 57.37%. The 3 October post calls it the strongest open-weights model on that post. Solving 476 tasks at least once is not the same as succeeding on 68 tasks in all 20 attempts.

Claude Opus 4.6 scores 68.62% on retail and 8.30% on auto insurance. Across Table 2, retail averages 59.52% pass@1 and auto insurance averages 33.83%.

Failure labels, as an unweighted average of per-model shares, are tool usage 79.9%, wrong state updates 10.3%, incomplete user resolutions 7.0%, and no state-changing action 2.9%. These are unweighted averages of per-model shares and observable labels, not unique causal explanations. The 79.9% label is tool usage, not reasoning. The post has not measured the lift from suggested mitigations.

Cost on this post is a comparative efficiency index, not an invoice and not a price announcement. The OpenRouter cost snapshot was taken on 20 September 2026. Opus 5.5 pricing is as per the Anthropic site. GPT-5.6 Sol has the lowest cost per success, at $0.127. Cost per dependable task picks different models on that index: GPT-5.4 at $6.80 across 128 tasks (25.25%), GPT-6 Astra at $7.45 across 231 tasks (45.56%), and Claude Opus 5.5 at $7.80 across 241 tasks. GPT-6 Astra’s pass@1 is 58.31%. It does not lead.

A clean tool trace and a resolved reply can still leave the wrong record. Keep pass@1, solved at least once, and observed 20 of 20 as separate numbers. Opus 5.5 and Opus 5 both have 241 tasks at 20 of 20. Superpower Daily on 3 October 2026 and Pivot News on 4 October 2026 repeat these figures and do not re-run the benchmark. They are not an independent measurement. An earlier Microsoft page dated 19 August 2026 has different scores. Do not mix that table with the 3 October table. No outside lab is claimed here to have reproduced the leaderboard.

https://huggingface.co/blog/microsoft/thinkingbox https://commandline.microsoft.com/thinkingbox-bench-agent-benchmarking/

ThinkingBoxMicrosoftHugging FaceAgents

Questions in this guide

Deep explanations with architecture diagrams for every question below.