Evidence record EV-0031
A model with cheaper tokens cost 3.7x more per run
Microsoft reports that on SharePoint Framework upgrade tasks Claude Sonnet 5 cost $2.01 per run against $0.55 for Claude Sonnet 4.6, 3.7x more despite 33% lower per-token prices, while on architecture tasks it was 12% cheaper.
Vendor measurement · Measurement · Retrieved
Evidence class: Vendor measurement. The vendor chose the tasks, ran the evaluation and wrote it up. Nobody else has reproduced it unless the record says so. Pattern tags: cost, measurement, evaluation.
Effect, as the source reports it
- Average cost per run, code upgrade tasks. Baseline: Claude Sonnet 4.6: $0.55. With the change: Claude Sonnet 5: $2.01. Direction: increase. Size: 3.7x. Sample: 3 SPFx scenarios; 5 runs per model per scenario.
- Average cost per run, architecture tasks. Baseline: Claude Sonnet 4.6: $0.54. With the change: Claude Sonnet 5: $0.47. Direction: decrease. Size: 12% cheaper. Sample: 12 architecture scenarios; 5 runs per model per scenario.
- Median token use, Sonnet 5 versus Sonnet 4.6. Baseline: Claude Sonnet 4.6. With the change: Claude Sonnet 5. Direction: increase. Size: 12x on architecture tasks; 10x on code upgrades. Sample: 150 runs in total.
- Quality and completion. Baseline: Claude Sonnet 4.6: idiomatic quality 90% (architecture); upgrade completion 60%. With the change: Claude Sonnet 5: idiomatic quality 78%; upgrade completion 100%. Direction: mixed. Size: Quality down 12 points on architecture; completion up 40 points on upgrades. Sample: 150 runs in total.
Agent profile
- Models: Claude Sonnet 4.6, Claude Sonnet 5.
- Harness: GitHub Copilot Chat in VS Code.
- Operating system: Windows.
Conflicts of interest
Microsoft authors evaluating with Microsoft tools (GitHub Copilot Chat, VS Code) and, where noted, Microsoft products. Microsoft reports the results itself. Costs are priced at GitHub Copilot's published rates; quality is scored by an LLM judge on Microsoft's own evaluation platform.
Source
Not all model upgrades are upgrades, Microsoft for Developers, Waldek Mastykarz, 6 July 2026. Retrieved ; verification: verified-live.
The full primary source was fetched on 2026-10-08 and every number in this record was found in it.
Limitations
- Vendor-run evaluation, not peer reviewed; no raw data or harness was found published with the post.
- Five runs per condition; Microsoft scopes each result to the agent profile it measured.
- Quality scored by an LLM judge on Microsoft's own, unpublished evaluation platform.
For designers
Compare models on cost per completed task, not on price per token. Run your own tasks, because the answer flipped between task types.
Cite this record
Cite the original source for any number, and keep the evidence class and model set with the figure. To point at this record, use "AX evidence register, EV-0031" and this page's address, https://agentexperience.tech/evidence/ev-0031/. The record is also in /evidence.json. The register's licence will be confirmed before its source repository is published.
