Evidence record EV-0009

Tools as code match or beat JSON tool calls for most models

The Bitter Lesson of Tool Calling (preprint) reports that exposing tools as typed Python stubs called through code matched or beat native JSON tool calling for 11 of 14 models on BFCL v4, and for 13 of 14 under parallel fan-out.

Evidence class: Preprint. Unreviewed. Many are by one author or a small team, and some authors have a stake in the result. Most AX research is in this class today. Pattern tags: code-mode.

Effect, as the source reports it

  • BFCL v4 accuracy, programmatic versus native JSON tool calling. Baseline: Native JSON tool calling. With the change: Programmatic tool calling matches or exceeds it in 11 of 14 models; GPT-5.6 family +10.6%. Direction: increase. Size: 11 of 14 models; best family +10.6%. Sample: 14 language models.
  • Accuracy under parallel fan-out. Baseline: Native JSON tool calling. With the change: Matches or outperforms in 13 of 14 models. Direction: increase. Size: 13 of 14 models. Sample: 14 language models.
  • Accuracy under context-rot conditions. Baseline: Native JSON: degrades 2.3% on average. With the change: Programmatic: holds stable. Direction: mixed. Size: JSON loses 2.3% on average; programmatic stable. Sample: 14 language models.

Agent profile

  • Models: GPT-5.6 family.
  • Note on models: 14 models across current and prior generations; the abstract names only the GPT-5.6 family.
  • Harness: BFCL v4 (Berkeley Function Calling Leaderboard).

Conflicts of interest

None declared in the abstract. The full text was not checked for a competing-interest statement.

Source

The Bitter Lesson of Tool Calling, arXiv, Ishan Patel, Sahil Sen, Elias Lumer, Vamse Kumar Subbiah, 6 August 2026, arXiv:2608.06370. Retrieved ; verification: abstract-only.

Every number in this record was checked against the live arXiv abstract page on 2026-10-08. The full text was not re-checked.

Limitations

  • Preprint, not peer reviewed.
  • One benchmark with simulated tools.
  • Measures call accuracy only; safety, approval and review cost of generated code were not measured.

For designers

Code-style access to your tools, such as typed stubs or an SDK in a sandbox, is a reasonable option next to one call per tool. Weigh it against the harder job of reviewing and approving generated code, which this study did not measure.

Cite this record

Cite the original source for any number, and keep the evidence class and model set with the figure. To point at this record, use "AX evidence register, EV-0009" and this page's address, https://agentexperience.tech/evidence/ev-0009/. The record is also in /evidence.json. The register's licence will be confirmed before its source repository is published.