ThinkingBox: Microsoft + Hugging Face agent benchmark
507 stateful business workflows that grade AI agents on the database state they leave behind, not on what their final reply claims.
Why it matters
Exposes the exact failure mode of current agents: most wrong answers terminate cleanly with no tool error. Claude Opus 5.5 leads pass@1 at 67.16%, but a stronger open-weight showing from Kimi-K3 (57.37%) makes it the benchmark to watch for open agent stacks.
What you could build with it
A startup selling QA for enterprise AI agents could wrap ThinkingBox's scenario harness as a continuous regression service: every agent code change gets re-graded against the 507 workflows with executable state checks. The result: shipping agent updates with proof of no silent behavior drift.
Does it hold up?
Too early to judge: public release was Oct 3; the paper reports results across 18 models, but no third-party adoption or independent reproductions yet.
Built with ThinkingBox
- New Benchmark Catches AI Agents Lying About Finished WorkDEV Community · Deep dive into the headline finding: 67% of failed attempts terminated cleanly, and what terminal-state grading changes about agent evaluation.
- ThinkingBox: Opus 5.5 and Opus 5 Tie at 241 of 507 TasksFourWeekMBA · Analysis of the cost angle: Opus 5.5 passes the same 241/507 tasks as Opus 5 at $7.80 vs $13.30 per dependable task - accuracy gains buying no added reliability.
- microsoft/thinkingbox-datagithub · Official executable benchmark, tool servers, and MCP-server mocks; MIT code, CDLA-Permissive-2.0 data.
Learn more
First spotted on article: source.
More AI agents
Agent ReachOpen-source CLI and SKILL.md kit that gives AI agents read/search access to 15+ internet platforms (Twitter, Reddit…agents · JEV 0.8Google Agent SkillsGoogle's official catalog of 150+ agent skills for its products and technologies — Cloud, GKE, BigQuery, Gemini, Agent…agents · JEV 0.74hermes-agent (NousResearch)Self-improving personal AI agent across Telegram, Discord, Slack, and terminal with persistent memory.agents · JEV 0.73Paperclip: the app people use to manage agents at workOpen-source orchestration for teams of AI agents: org charts, budgets, governance, heartbeats and a ticket system in…agents · JEV 0.73AIHOT — open-source framework for building automated industry-news sitesFull-stack framework (Node 24 + PostgreSQL) that ingests RSS/webpages/X/WeChat, dedupes, double-scores, clusters with…agents · JEV 0.73marketingskillsA library of ~60 agent skills covering CRO, copywriting, SEO/AEO, ads, analytics, churn and growth engineering…agents · JEV 0.72
Get the week's best AI launches, plus 3 ideas worth building
One email every Saturday. Ranked by traction, not hype. Free.