Evidence record EV-0026

Prompt injection split across tool channels evades defences

Across 12 frontier models and over 15,000 trials (preprint), models that resisted single-channel prompt injection exfiltrated data at up to 100% when the payload was split across two channels, such as a tool description and a tool result, and seven third-party MCP security tools failed to detect it.

Evidence class: Preprint. Unreviewed. Many are by one author or a small team, and some authors have a stake in the result. Most AX research is in this class today. Pattern tags: prompt-injection, description.

Effect, as the source reports it

  • Credential exfiltration compliance. Baseline: Single-channel injection: 0% for the named models. With the change: Two-channel fragmentation: up to 100%. Direction: increase. Size: 0% to up to 100%. Sample: 12 frontier models; three production clients; six payloads; over 15,000 trials.
  • Detection of fragmented payloads by third-party MCP security tools. With the change: All seven tools failed to detect them. Direction: not-applicable. Size: 0 of 7 tools detected them. Sample: Seven tools; three prompt-based defences (model-specific).

Agent profile

  • Models: GPT-4o, Llama 70B, Composer 2, Haiku 4.5.
  • Note on models: 12 frontier models; the abstract names these four as examples of models that resisted single-channel injection.
  • Harness: Three production clients (not named in the abstract); a VS Code MCP sampling override is described.

Conflicts of interest

None declared in the abstract. The full text was not checked for a competing-interest statement.

Source

Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines, arXiv, Murali Ediga, Sudipta Chattopadhyay, 16 September 2026, arXiv:2609.18217. Retrieved ; verification: abstract-only.

Every number in this record was checked against the live arXiv abstract page on 2026-10-08. The full text was not re-checked.

Limitations

  • Preprint, not peer reviewed.
  • Attack research with payloads designed by the authors.

For designers

Treat every tool description and tool result as untrusted input. Defences that limit what an agent can do, such as allowed destinations and narrow capabilities, hold better than defences that try to spot attacks.

Cite this record

Cite the original source for any number, and keep the evidence class and model set with the figure. To point at this record, use "AX evidence register, EV-0026" and this page's address, https://agentexperience.tech/evidence/ev-0026/. The record is also in /evidence.json. The register's licence will be confirmed before its source repository is published.