Stop reading the agent's transcript, check what its tools did. microsoft's thinkingbox,…
stop reading the agent's transcript, check what its tools did. microsoft's thinkingbox, open-sourced september 11, verifies ai agents by inspecting tool side effects instead of the chat log, catching what the transcript hides.
the evidence moved from words to state.
Context
Microsoft's blog of 19 August 2026 describes ThinkingBox as running agents in isolated, stateful tool environments with a simulated user answering follow-ups, and inspecting the side effects left behind. Its example task passes only if the preference appears on the booking and the support ticket records the correct resolution, and the blog says the transcript explains the result while the assertions decide whether the task passes. ThinkingBox-Bench has 507 tasks. The microsoft/thinkingbox repository is MIT licensed and was created 22 April 2026. The blog names the tau-bench family as the closest comparison and adds a reusable lifecycle around MCP servers.
The method is first-party. The 11 September open-source date was not found: the first-party blog is dated 19 August, a benchmark release snippet 17 August, and the repo creation date is not a public release date. Verifies AI agents is a compression, since it evaluates agents against executable assertions, and the transcript is still used to explain results. Benchmark scores in the blog are vendor-reported and are not used. The comparison table is Microsoft's own. Stop reading the transcript is the author's framing.
Watch next
- A first-party tag or announcement dated 11 September and independent reruns of the 507 tasks.
Sources
- ThinkingBox-Bench: agent benchmarking (Microsoft, 19 Aug 2026)commandline.microsoft.com
- microsoft/thinkingbox (GitHub)github.com
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 21 September 2026 at 15:40 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →