ax-check rule
AXC-D017: Context footprint estimate
Reports the bytes and approximate tokens the whole surface costs an agent before it does any work.
ax-check is a checker being prepared for release. This page documents the rule ahead of that release; see all 50 rules.
| Severity | info |
| Kind | Heuristic. A pattern match: a prompt to look, not a verdict. |
| Mode | ax-check lint |
| Applies to | MCP tool lists, OpenAPI, SKILL.md, CLI help |
| Pattern tags | context-budget, cost, dynamic-tools |
| Fix in one line | If the footprint is large, load tools on demand, search first, or split rarely used items out. |
This rule is information, not a failure. It reports roughly how much text the whole surface puts in front of an agent before the agent does any work, in bytes and in approximate tokens. The token figure is a rough estimate (characters divided by four), not a count from any model’s tokenizer.
What it checks
ax-check measures the text an agent would load for the surface and reports one finding per surface with its size. There is no threshold and no pass or fail. The number is there so you can see the cost, compare versions and notice when it grows.
Why it matters
Every tool definition, operation description or skill summary an agent loads takes room in its context window, on every turn, whether or not it is used. Large catalogues leave less room for the task itself and cost more per request.
Two public reports show the scale. In the SCOUT paper (Saha, Wang and Manoharan, arXiv 2608.23992, 25 August 2026), the authors report that in their production MCP gateway at PayPal, tool definitions took 140.2k tokens (70.1% of context) before retrieval and 1.3k tokens (0.8%) after. That is the authors’ own production claim, from a vendor-affiliated team, and it measures tokens, not accuracy. Cloudflare’s post “Code Mode: give agents an entire API in 1,000 tokens” (20 February 2026) makes a vendor claim about covering a large API with a small footprint by changing how tools are exposed.
For skills, only the name and description are loaded up front; the body loads when the skill activates. Wang et al. (arXiv 2610.04832, 4 October 2026) report that loading a skill’s body raised episode tokens by 50% on average across four agents (the paper’s claim).
How to fix
You do not need to fix anything; this is a measurement. If the footprint is large for what the surface does:
- Load tools on demand: expose a search or discovery tool first and load matching definitions only when needed.
- Split rarely used items into a separate server, skill or command group.
- Shorten long descriptions (see AXC-D016) and move reference detail into schemas or documentation.
- For skills, keep the body short and move detail into
references/(see AXC-D026).
Example
Before
tools.json
info AXC-D017 This tool list costs about 41250 tokens (165000 bytes) for 120 tools before any work starts.
After
tools.json
info AXC-D017 This tool list costs about 2900 tokens (11600 bytes) for 8 tools before any work starts.
In this invented example, the 120 tools were split into a small core set and a search_tools tool that returns further definitions on demand.
How ax-check detects it
ax-check builds the text an agent would load and measures it:
- MCP: the whole
toolsarray, serialised as compact JSON. Names, descriptions, schemas and annotations all count. - OpenAPI: the whole document, serialised as compact JSON, including responses, components and examples. This over-counts compared with what a client turns into tools, which is usually the operations and their inputs only. Treat the OpenAPI figure as an upper bound.
- Agent Skills: one line per skill, its name and description, which is what is always loaded. The bodies are reported separately as the extra tokens added when skills activate.
- CLI help: the root
--helptext plus every subcommand help text ax-check captured.
Bytes are UTF-8 bytes. Tokens are the number of characters divided by four, rounded up. This is a coarse, model-independent estimate. Real token counts depend on the model’s tokenizer and on the text: JSON, code and languages other than English often come out quite differently. Use the figure to compare versions of the same surface, not as a bill.
The finding’s data also lists the three largest entries, measured by the size of their description, parameter names and descriptions, and schema.
If you do not want this information in your report, silence it with --disable AXC-D017.
Sources
- Paper: Saha, Wang and Manoharan, “Hybrid Semantic Tool Discovery for Enterprise MCP Gateway: Architecture and Implementation” (SCOUT), arXiv 2608.23992, 25 August 2026. https://arxiv.org/abs/2608.23992. The authors’ production claim (vendor-affiliated): tool-token use fell from 140.2k tokens (70.1% of context) to 1.3k tokens (0.8%).
- Vendor: Code Mode: give agents an entire API in 1,000 tokens, Cloudflare blog, 20 February 2026. https://blog.cloudflare.com/code-mode-mcp/. A vendor claim about the context footprint of exposing an API to agents.
- Paper: Wang et al., “Agent Skill Evolution: How Revisions Affect Coding Agents”, arXiv 2610.04832, 4 October 2026. https://arxiv.org/abs/2610.04832. The paper’s claim, across four agents: loading a skill’s body raised episode tokens by 50% on average.
- Guide: Agent tool catalogues: how to keep an MCP tool catalog small, agentexperience.tech, retrieved 2026-10-08. https://agentexperience.tech/insights/agent-tool-catalogs/. Why a small, focused catalogue is easier for agents to use.
Related evidence
Records in the AX evidence register that share a pattern tag with this rule. A shared tag means the record is about the same pattern, not that it tests this rule. Read the evidence class before the number.
- EV-0022: Search-and-execute meta-tools cut catalogue tokens by 99% in production (Preprint). PayPal authors report (preprint) that exposing two meta-tools, search and execute, over 2,000+ MCP tools cut tool-token consumption in production from 140.2k tokens (70.1% of context) to 1.3k tokens (0.8%).
- EV-0017: Injected skills lowered pass rates and raised token cost on average (Preprint). WebDev-Skills-Bench (preprint) found that injecting matched public skills reduced mean Pass@2 by 1.3% to 4.2% across four models and raised token cost by 72% to 394%, with gains in only 17% to 36% of skill-project pairs.
- EV-0010: Agent scaffolding, not MCP versus CLI, drove cost (Preprint). A controlled comparison (preprint) found that on one git task the agent scaffolding drove cost more than MCP versus CLI: CLI-only scaffoldings were 5.0x to 28x cheaper, and paired MCP-to-CLI cost ratios ranged from 0.43x to 29x.
- EV-0023: Hiding tools is not enforcing permissions (Preprint). Across 2,160 attempts with four frontier models (preprint), a server with only in-body permission checks exposed forbidden tools in 152 of 720 trials and permission-aware visibility cut that to 0 of 720, yet models named a hidden tool in up to 94% of settings when it was inferable from the prompt.
- EV-0027: A capable agent skipped the index and guessed the page (Preprint). A preregistered ablation on a 709-page Markdown wiki (preprint) found that a capable tool-using agent never loaded the compact catalogue index, inferring page paths from the question instead, while retrieval-based access kept answer quality non-inferior and cut cost by about a third to over half.
- EV-0028: A JSON input mode for a CLI raised agent cost 4x to 11x (Vendor measurement). Microsoft reports that on a synthetic CLI, run five times per model in GitHub Copilot Chat, every model was correct 5/5 with individual arguments, while a --json payload mode cost 4x to 11x more per task and dropped Claude Haiku 4.5 to 2/5 and MAI-Code-1-Flash to 3/5.
