Evidence record EV-0028
A JSON input mode for a CLI raised agent cost 4x to 11x
Microsoft reports that on a synthetic CLI, run five times per model in GitHub Copilot Chat, every model was correct 5/5 with individual arguments, while a --json payload mode cost 4x to 11x more per task and dropped Claude Haiku 4.5 to 2/5 and MAI-Code-1-Flash to 3/5.
Vendor measurement · Negative result · Retrieved
Evidence class: Vendor measurement. The vendor chose the tasks, ran the evaluation and wrote it up. Nobody else has reproduced it unless the record says so. Pattern tags: cost, measurement.
Effect, as the source reports it
- Correct deployments out of five runs. Baseline: Individual arguments: 5/5 for every model. With the change: --json payload: 5/5 for Claude Sonnet 5, Claude Sonnet 4.6 and GPT-5.3-Codex; 3/5 for MAI-Code-1-Flash; 2/5 for Claude Haiku 4.5. Direction: decrease. Size: Down to 2/5 for the smallest model. Sample: Five runs per model per input mode; five models.
- Cost per task (GitHub Models pricing). Baseline: Arguments: GPT-5.3-Codex $0.05; MAI-Code-1-Flash $0.01; Claude Sonnet 4.6 $0.05; Claude Haiku 4.5 $0.03; Claude Sonnet 5 $0.08. With the change: JSON: $0.54 (11x); $0.08 (10x); $0.47 (9x); $0.23 (8x); $0.32 (4x). Direction: increase. Size: 4x to 11x. Sample: As above.
- Arguments-to-JSON cost gap by shell, Claude Sonnet 4.6. Baseline: Bash on macOS: 1.5x. With the change: PowerShell on Windows: 9x. Direction: increase. Size: 1.5x on Bash, 9x on PowerShell. Sample: One model, two shells.
Agent profile
- Models: Claude Haiku 4.5, Claude Sonnet 4.6, Claude Sonnet 5, GPT-5.3-Codex, MAI-Code-1-Flash.
- Harness: GitHub Copilot Chat.
- Operating system: Windows (PowerShell) for the main experiment; macOS (Bash) for the Sonnet 4.6 shell comparison.
Conflicts of interest
Microsoft authors evaluating with Microsoft tools (GitHub Copilot Chat, VS Code) and, where noted, Microsoft products. Microsoft reports the results itself. The CLI (podctl) is synthetic and was built for the test; MAI-Code-1-Flash is a Microsoft model.
Source
Don't rewrite your CLI for agents, Microsoft for Developers, Waldek Mastykarz, 7 July 2026. Retrieved ; verification: verified-live.
The full primary source was fetched on 2026-10-08 and every number in this record was found in it.
Limitations
- Vendor-run evaluation, not peer reviewed; no raw data or harness was found published with the post.
- Five runs per condition; Microsoft scopes each result to the agent profile it measured.
- One synthetic CLI and one deployment scenario with 30+ values.
For designers
Keep named arguments when you make a CLI agent-friendly; a JSON payload adds failure modes such as shell quoting. If you add --json, keep the arguments too and measure both on the shells your users run.
Related checks
Cite this record
Cite the original source for any number, and keep the evidence class and model set with the figure. To point at this record, use "AX evidence register, EV-0028" and this page's address, https://agentexperience.tech/evidence/ev-0028/. The record is also in /evidence.json. The register's licence will be confirmed before its source repository is published.
