Evidence record EV-0004

Naming recovery tools is the active ingredient in failure receipts

Outcome Monitors (preprint) reports that receipts naming a violated outcome and the public recovery tools raised ToolMaze completion from 10.9% to 28.1% across four models, and that removing the list of recovery tools eliminated the gain.

Evidence class: Preprint. Unreviewed. Many are by one author or a small team, and some authors have a stake in the result. Most AX research is in this class today. Pattern tags: error-recovery, false-success.

Effect, as the source reports it

  • ToolMaze task completion with injected silent failures. Baseline: No outcome monitor: 10.9%. With the change: Outcome monitor receipts: 28.1%. Direction: increase. Size: +17.2 percentage points. Sample: Four models in two provider families, replicated in a third family.
  • Completion gain when the recovery-tool list is removed from the receipt. Baseline: Receipt with recovery-tool list. With the change: Receipt without it: measured gain eliminated; restoring the list recovered it. Direction: decrease. Size: Gain eliminated. Sample: Separate ToolMaze controls.
  • Completion when diagnostic detail or timing is varied. With the change: No detectable difference. Direction: no-change. Size: No detectable difference. Sample: Separate ToolMaze controls.
  • tau-bench retail completion. Baseline: Without monitors. With the change: With monitors: +14.0 and +12.0 points on two tiers. Direction: increase. Size: +14.0 and +12.0 percentage points. Sample: Two tiers.

Agent profile

  • Note on models: Four models in two provider families, replicated in a third. The abstract does not name them.

Conflicts of interest

None declared in the abstract. The full text was not checked for a competing-interest statement.

Source

Outcome Monitors: Recovery Affordances for Silent Tool Failures, arXiv, Sugam Panthi, Rabab Abdelfattah, 19 August 2026, arXiv:2608.19303. Retrieved ; verification: abstract-only.

Every number in this record was checked against the live arXiv abstract page on 2026-10-08. The full text was not re-checked.

Limitations

  • Preprint, not peer reviewed.
  • Failures were injected; detection outside the mined contract vocabulary fell to 46% on a suite from a published incident taxonomy.

For designers

When a tool result looks wrong, the most useful thing to send back is the name of the tool that can check or fix it. Extra diagnostic detail alone did not help in these tests.

Cite this record

Cite the original source for any number, and keep the evidence class and model set with the figure. To point at this record, use "AX evidence register, EV-0004" and this page's address, https://agentexperience.tech/evidence/ev-0004/. The record is also in /evidence.json. The register's licence will be confirmed before its source repository is published.