Outcome Monitors is a runtime layer for a specific agent failure: a tool can return a well-formed payload that describes the wrong outcome. The monitor checks the tool, arguments, and result against an outcome contract, preserves the original payload, and appends a nonbinding receipt with the violated property and visible recovery tools. 1
In frozen ToolMaze tests, the paper reports full-workflow completion rising from 10.9% to 28.1% across four models and 320 primary episodes. Retail τ-bench completion improved by 14 and 12 percentage points on two categorical-fault tiers. Removing the recovery-tool list returned performance to baseline in the tested controls. Detection outside the mined contract vocabulary fell to 46%, and 230 of 320 ToolMaze episodes still failed. 1
A first build can stay beside the current route: define contracts from task-disjoint traces or public schemas, run deterministic checks without another model call, preserve the result, and append a receipt that lists only workflow-visible recovery tools. Shadow-test completion, contract coverage, false positives, recovery choices, harmful recoveries, and p95 latency before routing users. The source says indexing code and the ToolMaze recovery mapping ship with its artifact, but it names no separate public repository or download URL. 1
Fuentes de referencia
- 1


Comentar