ax-check rule

AXC-D016: Description length outlier

The description is far longer than the others in the same catalogue, which costs context on every turn.

ax-check is a checker being prepared for release. This page documents the rule ahead of that release; see all 50 rules.

Severityinfo
KindHeuristic. A pattern match: a prompt to look, not a verdict.
Modeax-check lint
Applies toMCP tool lists, OpenAPI, SKILL.md, CLI help
Pattern tagscontext-budget, cost
Fix in one lineMove reference detail out of the description into docs, a resource or a schema, and keep selection text short.

One description is far longer than the others in the same catalogue. Descriptions are usually loaded into the agent’s context on every turn, so a very long one costs context all the time, even when the item is not used. This is information, not a failure: a long description can be justified, but it is worth a second look.

What it checks

ax-check compares the length of each description with the median length of the other descriptions on the same surface (one MCP tool list, one OpenAPI document, one set of skills or one CLI’s help). It reports descriptions that are far above the median, or very long in absolute terms when the catalogue is too small to have a useful median.

Why it matters

Tool and skill descriptions are part of the prompt. An MCP client typically gives the model every tool description before the first call, and a skill’s description is always loaded so the agent can decide whether to activate it. A 4,000 character description in a list of 300 character ones usually means reference material has crept into selection text: full error tables, long examples or a changelog.

Footprint matters at scale. In the SCOUT paper (Saha, Wang and Manoharan, arXiv 2608.23992, 25 August 2026), the authors report that in their production MCP gateway at PayPal, tool definitions took 140.2k tokens (70.1% of the context) before they added retrieval, and 1.3k tokens (0.8%) after. That is the authors’ own production claim, from a vendor-affiliated team, and it is about token accounting, not accuracy.

How to fix

  • Keep the description to what the item does, when to use it, when not to and which neighbour to use instead.
  • Move reference detail into the input schema (enums, formats, limits), into a resource or into documentation the agent can fetch on demand.
  • For skills, move detail into the SKILL.md body or into files under references/.

Example

Before

{
  "name": "create_invoice",
  "description": "Create a draft invoice for one customer. Use this when the user asks to bill a customer. Error codes: E100 means the customer was not found, E101 means the currency is not supported, E102 means ... (40 more error codes) ... Changelog: in version 2 the tax field became optional ... (2,800 more characters)"
}

After

{
  "name": "create_invoice",
  "description": "Create a draft invoice for one customer and return its invoice ID. Use this when the user asks to bill a customer for work already done. Do not use this to send an invoice; use send_invoice instead. Errors name the field to fix."
}

How ax-check detects it

ax-check measures each description in characters. Only descriptions that are present and are not placeholders take part. For OpenAPI, the operation’s summary and description are joined and measured together.

  • With four or more descriptions on the surface, a description is reported when it is longer than three times the median length and also longer than 1,000 characters.
  • With fewer than four descriptions, a description is reported when it is longer than 2,000 characters.

The finding gives the length and the median. The 1,000 character floor stops the rule firing on catalogues where every description is short and one is merely medium.

Known false positives: a tool whose job needs a long contract (for example a query language summary) may be reported although the length is deliberate. Known false negatives: a catalogue where every description is too long has a high median, so none stands out. AXC-D017 reports the total footprint for that case.

If the finding does not apply, silence it with --disable AXC-D016.

Sources

  • Paper: Saha, Wang and Manoharan, “Hybrid Semantic Tool Discovery for Enterprise MCP Gateway: Architecture and Implementation” (SCOUT), arXiv 2608.23992, 25 August 2026. https://arxiv.org/abs/2608.23992. The authors’ production claim (vendor-affiliated): tool-token use fell from 140.2k tokens (70.1% of context) to 1.3k tokens (0.8%). A token-accounting result, not an accuracy result.
  • Guide: Write tool descriptions an agent can act on, agentexperience.tech, retrieved 2026-10-08. https://agentexperience.tech/insights/tool-descriptions/. Keep the description to what, when, when not and which neighbour; put closed sets in the schema.
  • Guide: Agent tool catalogues: how to keep an MCP tool catalog small, agentexperience.tech, retrieved 2026-10-08. https://agentexperience.tech/insights/agent-tool-catalogs/. Keep a catalogue focused.

Records in the AX evidence register that share a pattern tag with this rule. A shared tag means the record is about the same pattern, not that it tests this rule. Read the evidence class before the number.