Evidence record EV-0003
Error text that names the next tool lifts recovery
A preprint testing five OpenAI models reports that an expired-credential error naming a terminal command left 45% of tasks recovered, naming the server's login tool instead raised recovery to 84%, and on rate limits naming the call to repeat raised recovery from 6% to 88%.
Preprint · Measurement · Retrieved
Evidence class: Preprint. Unreviewed. Many are by one author or a small team, and some authors have a stake in the result. Most AX research is in this class today. Pattern tags: error-recovery, description.
Effect, as the source reports it
- Tasks recovered after an expired-credential error. Baseline: Next step names a terminal command: 45% recovered. With the change: Next step names the server's login tool: 84% recovered. Direction: increase. Size: +39 percentage points. Sample: Five OpenAI models on Berkeley Function Calling Leaderboard tasks; sample counts are not in the abstract.
- Tasks recovered after a rate-limit error. Baseline: GitHub's "Wait before retrying.": 6% recovered. With the change: Step names the call to repeat: 88% recovered. Direction: increase. Size: +82 percentage points. Sample: As above.
- Recovery lost to a developer-addressed step, by model generation. Baseline: GPT-5.5: 18 points lost. With the change: GPT-6 Astra: 69 points lost. Direction: increase. Size: Loss grew from 18 to 69 points with the newer model. Sample: As above.
- Recovery on expired credentials with an agent-side fix. Baseline: Terminal-command step left in place: 45%. With the change: One-sentence prompt deletes the step before the model reads it: 82%. Direction: increase. Size: +37 percentage points. Sample: As above.
- Prevalence of next steps in MCP server error messages. Direction: not-applicable. Size: 949 of 3,001 messages give a next step; on credential errors 62 of 67 steps ask for a terminal command, configuration change or web page; on rate limits 20 of 30 say to wait without naming the call. Sample: 150 widely used MCP servers.
Agent profile
- Models: GPT-5.5, GPT-6 Astra.
- Note on models: Five OpenAI models; the abstract names these two.
- Harness: Berkeley Function Calling Leaderboard tasks; the agents act only through the tools, with no terminal.
Conflicts of interest
None declared in the abstract. The full text was not checked for a competing-interest statement.
Source
MCP Error Messages Written for Developers Hurt the Most Capable Agents Most, arXiv, Xiaonan Xu, Wenjing Wu, 28 September 2026 (v2, revised 2026-09-29), arXiv:2609.35381. Retrieved ; verification: abstract-only.
Every number in this record was checked against the live arXiv abstract page on 2026-10-08. The full text was not re-checked.
Limitations
- Preprint, not peer reviewed.
- OpenAI models only, so transfer to other model families is untested.
- The agent has no terminal by construction, which is the condition under which developer-addressed steps fail.
For designers
Write error messages for a caller that can only call your tools: name the tool or call that fixes the problem, not a command, a settings page or "wait". Newer models followed the step more literally, so a wrong step cost more.
Related checks
Cite this record
Cite the original source for any number, and keep the evidence class and model set with the figure. To point at this record, use "AX evidence register, EV-0003" and this page's address, https://agentexperience.tech/evidence/ev-0003/. The record is also in /evidence.json. The register's licence will be confirmed before its source repository is published.
