Evidence record EV-0024
Harnesses alter shell calls and the wrong action runs silently
A preprint reports that in 47,828 production shell calls, Claude Code's Bash tool changed 12.0% of calls carrying code, escape sequences or long text, and that for 80.7% of calls whose backslashes were changed the wrong action ran with no reported error.
Preprint · Measurement · Retrieved
Evidence class: Preprint. Unreviewed. Many are by one author or a small team, and some authors have a stake in the result. Most AX research is in this class today. Pattern tags: false-success, evaluation, measurement.
Effect, as the source reports it
- Calls changed in transit by the harness. With the change: Claude Code's Bash tool: 12.0% of calls carrying code, escape sequences or long text. Direction: not-applicable. Size: 12.0%; all 10 measured harnesses change some call. Sample: 47,828 shell calls in production sessions; 10 harnesses.
- Calls with changed backslashes where the wrong action ran with no error. Direction: not-applicable. Size: 80.7%. Sample: As above.
- Failures blamed on the model by trajectory-based judgement. Direction: not-applicable. Size: 95.1% blamed on the LLM, although the path caused more than half. Sample: Production failures.
- Token cost per passed task caused by the path. Baseline: Unaltered path. With the change: Altering path. Direction: increase. Size: 2.4 times (up to 12.3 times). Sample: IEC-Bench.
Agent profile
- Note on models: Not a model comparison; the abstract does not name the models in the production sessions.
- Harness: Claude Code (production sessions); 10 harnesses measured, four in IEC-Bench.
Conflicts of interest
The authors' repair, IntAct, is deployed in a commercial product.
Source
Do Tool Calls Execute as Intended? Measuring and Repairing Intent-Execution Correspondence in LLM Agents, arXiv, Boyang Yang, Zhenhao Li, Ziyao Yang, Kanghui Jia, Xin Yin, Mingmou Liu, Haoye Tian, 3 October 2026, arXiv:2610.04375. Retrieved ; verification: abstract-only.
Every number in this record was checked against the live arXiv abstract page on 2026-10-08. The full text was not re-checked.
Limitations
- Preprint, not peer reviewed.
- Harness versions are not stated in the abstract.
For designers
When a tool call fails, log what the tool actually received, not only what the model sent. A command that arrives altered can run the wrong action with no error.
Related checks
Cite this record
Cite the original source for any number, and keep the evidence class and model set with the figure. To point at this record, use "AX evidence register, EV-0024" and this page's address, https://agentexperience.tech/evidence/ev-0024/. The record is also in /evidence.json. The register's licence will be confirmed before its source repository is published.
