AI News

ThinkingBox Checks Whether AI Agents Changed the Right Records

A customer-support agent can make valid tool calls, send a reassuring reply and still leave the wrong status in the database. Microsoft's ThinkingBox is built to catch that gap.

An October 3 joint Microsoft–Hugging Face post announced ThinkingBox's availability through Hugging Face and its OpenEnv evaluation interface. The underlying research is older: the paper was first submitted on August 20 and revised on October 1. The fresh development gives teams another route to run evaluations; it is not a newly discovered failure mode or a new model launch. Announcement and paper history.

The practical contribution is a way to examine an agent's final state changes, unwanted effects and consistency across repeated attempts. For anyone building an agent that updates tickets, bookings or account records, those are more useful acceptance criteria than whether its final message sounds finished.

What ThinkingBox actually checks

ThinkingBox is the sandbox and harness; ThinkingBox-Bench is the benchmark. The executable release contains 507 synthetic workflows across retail, travel, auto insurance, neobank support and consulting IT/HR. Each attempt begins in an isolated environment and is checked against the required backend outcome and side effects. Some tasks also assess requirements of the final response. Executable benchmark repository.

These are reconstructed scenarios, not transcripts proving that named businesses or real customers suffered the reported failures. That distinction matters when reading an example involving a support agent that closes a ticket even though a delivery issue remains unresolved.

The evaluator can reject an incorrect field value, a missing change or an extra change that the user never requested. It can also accept different sequences of tool calls that reach an allowed outcome. The test need not prescribe one exact conversation to know whether the resulting record is correct. Full paper, evaluation design.

There is a limit: 477 of the 507 tasks are graded on state and effects alone; 30 add response checks. A correct database transition with a misleading customer reply could therefore pass a state-only task. The paper acknowledges that tradeoff. Your own acceptance tests should include communication requirements when the reply itself is consequential.

The failure signal behind the announcement

The paper's retrospective analysis examined 79,853 failed attempts within a common set of 121,680 valid recorded trials across 12 models. Of those failed attempts, 67.24% terminated cleanly, performed a state-changing action and ended without an explicit error in the final tool response. They still failed the executable outcome checks. Paper, Table 16.

The denominator is failed attempts in that analysis. This does not mean 67.24% of all enterprise agents fail, or that a similar percentage will fail in your application. It shows why those three observable signs were weak substitutes for the required state in this experiment.

A useful response is to change what your test records. Keep the agent's narration, but retain the actual row identifiers, before-and-after values and effects that determine success. Our OpenClaw Enterprise pilot guide considers the control layer around agents; ThinkingBox addresses another part of deployment: proving that a particular attempted task achieved its allowed outcome.

Three repeat metrics that answer different questions

The announcement distinguishes average attempt success, solving a task at least once in twenty attempts, and the share of tasks that passed all twenty recorded attempts. Those measures should not be used interchangeably. Its “observed 20/20” measure is a literal count; the paper also discusses an estimated repeat-reliability metric. Keep the definitions attached to any number you cite. Metric definitions.

Here is an illustrative calculation, not a ThinkingBox result. If an attempt succeeds with probability 95%, independently each time under unchanged conditions, the chance of twenty consecutive successes is 0.95^20, about 35.8%. At 99% per attempt it is about 81.8%. Real failures can be correlated, and production conditions change, so neither number is a forecast.

Even observing twenty successes is not proof that attempt twenty-one will work. Repetition exposes some instability; it cannot certify every future input. For an evaluator, report the attempted task set, repetitions, successful outcomes and failure definitions so readers can understand what was actually measured.

Turn the idea into a concrete acceptance test

Consider a fictional support workflow: a shipment remains in a carrier exception. The permitted response is to open or update one ticket, place it on hold and tell the customer what remains unresolved. This is Kingy's proposed test design, not a task we ran or a reproduction of the benchmark's exact policy.

Check Evidence to retain Failure it catches
Correct target Order and ticket IDs bound to the request Updating another customer's record
Required outcome Ticket status and required notes after execution Marking an unresolved issue solved
Allowed effects Complete change list compared with a permitted set Extra tickets or unauthorized account edits
Retry behavior Transaction key and resulting record count after repetition Duplicating a ticket after a timeout
Reply Reply checked against the unresolved state Claiming the problem is fixed when it remains open

Start from a clean, known state and record it. Define the permitted end state before running the agent. Make ambiguous cases explicit: if policy requires more information, success may be asking for it without changing the record.

After an attempt, query the authoritative state rather than accepting the agent's account of what it did. Compare both required changes and forbidden effects. Repeat from an identical starting state to test consistency, then add separate failure-injection cases for timeouts and retries. A retry test should establish whether the earlier write happened before allowing another mutation.

This proposal has not been benchmarked by Kingy, and we are not claiming it improves ThinkingBox scores. Its value is making a deployment decision reviewable: you can point to the required result, the actual result and the changes between them.

The OpenEnv adapter still needs infrastructure

The public documentation says the adapter currently serves evaluation. Its container starts the OpenEnv API, while the Session Proxy, MCP servers, Typesense and agent, simulated-user and judge endpoints are managed separately. A running container does not establish that the whole evaluation stack is ready. OpenEnv deployment boundary.

The docs distinguish /health, which reports process liveness, from /ready, which checks observable dependencies. Scenario-specific Typesense readiness is outside what that process can observe. Verify it separately before trusting a result. The Hugging Face dataset is a viewer-friendly representation; the pinned GitHub release supplies executable benchmark assets. Environment and data documentation.

Keep infrastructure errors visible. The adapter separates them from binary episode outcomes, and the announcement says its reported analysis counted system errors as unsuccessful trials. Those are different reporting choices; a comparison needs an explicit denominator and treatment of errors. Silently dropping failed infrastructure attempts can make a system look more dependable than the full operation was.

A sensible first evaluation is one representative workflow with retained state, effects and error records. Expand only after you can explain why it passed. When comparing models later, keep the policy, tool surface, harness configuration and starting data fixed. ThinkingBox makes that evaluation approach accessible, but your own workflow, permissions and failure costs still determine what is safe to automate.