ax-check rule
AXC-D016: Description length outlier
The description is far longer than the others in the same catalogue, which costs context on every turn.
ax-check is a checker being prepared for release. This page documents the rule ahead of that release; see all 50 rules.
| Severity | info |
| Kind | Heuristic. A pattern match: a prompt to look, not a verdict. |
| Mode | ax-check lint |
| Applies to | MCP tool lists, OpenAPI, SKILL.md, CLI help |
| Pattern tags | context-budget, cost |
| Fix in one line | Move reference detail out of the description into docs, a resource or a schema, and keep selection text short. |
One description is far longer than the others in the same catalogue. Descriptions are usually loaded into the agent’s context on every turn, so a very long one costs context all the time, even when the item is not used. This is information, not a failure: a long description can be justified, but it is worth a second look.
What it checks
ax-check compares the length of each description with the median length of the other descriptions on the same surface (one MCP tool list, one OpenAPI document, one set of skills or one CLI’s help). It reports descriptions that are far above the median, or very long in absolute terms when the catalogue is too small to have a useful median.
Why it matters
Tool and skill descriptions are part of the prompt. An MCP client typically gives the model every tool description before the first call, and a skill’s description is always loaded so the agent can decide whether to activate it. A 4,000 character description in a list of 300 character ones usually means reference material has crept into selection text: full error tables, long examples or a changelog.
Footprint matters at scale. In the SCOUT paper (Saha, Wang and Manoharan, arXiv 2608.23992, 25 August 2026), the authors report that in their production MCP gateway at PayPal, tool definitions took 140.2k tokens (70.1% of the context) before they added retrieval, and 1.3k tokens (0.8%) after. That is the authors’ own production claim, from a vendor-affiliated team, and it is about token accounting, not accuracy.
How to fix
- Keep the description to what the item does, when to use it, when not to and which neighbour to use instead.
- Move reference detail into the input schema (enums, formats, limits), into a resource or into documentation the agent can fetch on demand.
- For skills, move detail into the SKILL.md body or into files under
references/.
Example
Before
{
"name": "create_invoice",
"description": "Create a draft invoice for one customer. Use this when the user asks to bill a customer. Error codes: E100 means the customer was not found, E101 means the currency is not supported, E102 means ... (40 more error codes) ... Changelog: in version 2 the tax field became optional ... (2,800 more characters)"
}
After
{
"name": "create_invoice",
"description": "Create a draft invoice for one customer and return its invoice ID. Use this when the user asks to bill a customer for work already done. Do not use this to send an invoice; use send_invoice instead. Errors name the field to fix."
}
How ax-check detects it
ax-check measures each description in characters. Only descriptions that are present and are not placeholders take part. For OpenAPI, the operation’s summary and description are joined and measured together.
- With four or more descriptions on the surface, a description is reported when it is longer than three times the median length and also longer than 1,000 characters.
- With fewer than four descriptions, a description is reported when it is longer than 2,000 characters.
The finding gives the length and the median. The 1,000 character floor stops the rule firing on catalogues where every description is short and one is merely medium.
Known false positives: a tool whose job needs a long contract (for example a query language summary) may be reported although the length is deliberate. Known false negatives: a catalogue where every description is too long has a high median, so none stands out. AXC-D017 reports the total footprint for that case.
If the finding does not apply, silence it with --disable AXC-D016.
Sources
- Paper: Saha, Wang and Manoharan, “Hybrid Semantic Tool Discovery for Enterprise MCP Gateway: Architecture and Implementation” (SCOUT), arXiv 2608.23992, 25 August 2026. https://arxiv.org/abs/2608.23992. The authors’ production claim (vendor-affiliated): tool-token use fell from 140.2k tokens (70.1% of context) to 1.3k tokens (0.8%). A token-accounting result, not an accuracy result.
- Guide: Write tool descriptions an agent can act on, agentexperience.tech, retrieved 2026-10-08. https://agentexperience.tech/insights/tool-descriptions/. Keep the description to what, when, when not and which neighbour; put closed sets in the schema.
- Guide: Agent tool catalogues: how to keep an MCP tool catalog small, agentexperience.tech, retrieved 2026-10-08. https://agentexperience.tech/insights/agent-tool-catalogs/. Keep a catalogue focused.
Related evidence
Records in the AX evidence register that share a pattern tag with this rule. A shared tag means the record is about the same pattern, not that it tests this rule. Read the evidence class before the number.
- EV-0017: Injected skills lowered pass rates and raised token cost on average (Preprint). WebDev-Skills-Bench (preprint) found that injecting matched public skills reduced mean Pass@2 by 1.3% to 4.2% across four models and raised token cost by 72% to 394%, with gains in only 17% to 36% of skill-project pairs.
- EV-0022: Search-and-execute meta-tools cut catalogue tokens by 99% in production (Preprint). PayPal authors report (preprint) that exposing two meta-tools, search and execute, over 2,000+ MCP tools cut tool-token consumption in production from 140.2k tokens (70.1% of context) to 1.3k tokens (0.8%).
- EV-0010: Agent scaffolding, not MCP versus CLI, drove cost (Preprint). A controlled comparison (preprint) found that on one git task the agent scaffolding drove cost more than MCP versus CLI: CLI-only scaffoldings were 5.0x to 28x cheaper, and paired MCP-to-CLI cost ratios ranged from 0.43x to 29x.
- EV-0027: A capable agent skipped the index and guessed the page (Preprint). A preregistered ablation on a 709-page Markdown wiki (preprint) found that a capable tool-using agent never loaded the compact catalogue index, inferring page paths from the question instead, while retrieval-based access kept answer quality non-inferior and cut cost by about a third to over half.
- EV-0028: A JSON input mode for a CLI raised agent cost 4x to 11x (Vendor measurement). Microsoft reports that on a synthetic CLI, run five times per model in GitHub Copilot Chat, every model was correct 5/5 with individual arguments, while a --json payload mode cost 4x to 11x more per task and dropped Claude Haiku 4.5 to 2/5 and MAI-Code-1-Flash to 3/5.
- EV-0031: A model with cheaper tokens cost 3.7x more per run (Vendor measurement). Microsoft reports that on SharePoint Framework upgrade tasks Claude Sonnet 5 cost $2.01 per run against $0.55 for Claude Sonnet 4.6, 3.7x more despite 33% lower per-token prices, while on architecture tasks it was 12% cheaper.
