ax-check rule
AXC-D005: Does not say when not to use it
The description never says when this is the wrong choice, so near-miss tasks select it too.
ax-check is a checker being prepared for release. This page documents the rule ahead of that release; see all 50 rules.
| Severity | warn (info for CLI help, info for OpenAPI) |
| Kind | Heuristic. A pattern match: a prompt to look, not a verdict. |
| Mode | ax-check lint |
| Applies to | MCP tool lists, OpenAPI, SKILL.md, CLI help |
| Pattern tags | description, selection, skills |
| Fix in one line | Add a sentence such as "Do not use this for X; use <neighbour> instead." |
The description never says when this item is the wrong choice. Without a boundary, tasks that are close but not quite right also select it. A short “Do not use this for X” sentence draws the line an agent cannot infer on its own.
What it checks
ax-check looks for a negative boundary anywhere in the description of each tool, operation, skill and subcommand: a “do not use”, a “not for”, an “instead” or a similar phrase. If it finds none, it reports the item.
Why it matters
Most catalogues contain items that overlap. A search tool and a fetch-by-ID tool both return invoices. A refund tool and a void tool both stop a payment. A description that says only what an item does will match every near-miss task, and the agent picks whichever reads closest.
Anand and Chattaraj built “canary tools” to test this: decoys designed to look right for a task while being wrong, in six types including semantic decoys, capability mirages and granularity traps. The paper’s claim (8 models, 120 tasks, August 2026) is that models are misled by such decoys, that susceptibility varies about 36 times across the models tested, and that it drops sharply as models get more capable. The paper did not test descriptions that state a boundary. Treating a stated boundary as a defence is this rule’s design rationale: it is the cheapest change a tool author controls.
How to fix
Add a sentence such as “Do not use this for X; use <neighbour> instead.” Name the near-miss task, and name the item that handles it (see AXC-D006).
Example
Before
{
"name": "search_invoices",
"description": "Searches invoices by customer, status or date range and returns up to 50 matches. Use this when the user asks to find invoices."
}
After
{
"name": "search_invoices",
"description": "Searches invoices by customer, status or date range and returns up to 50 matches. Use this when the user asks to find invoices. Do not use this to fetch one invoice whose ID you already know; use get_invoice instead."
}
How ax-check detects it
The rule passes when the description contains any of these phrase families, matched case-insensitively and including their contracted forms:
- “do not”, “never” or “avoid”, followed by use, call, invoke, choose, pick, select, run, trigger, load, apply or rely;
- “not” followed by for, suitable, intended, meant, designed, appropriate, the right, a replacement, a substitute or needed;
- “instead”, “rather than”, “in place of”, “as opposed to”;
- “wrong choice”, “wrong tool”, “wrong skill”, “wrong command”;
- “only for”, “only when”, “only if”, “only to” (and “solely” in the same forms);
- “unless”;
- “does not”, “cannot”, “will not” or “is not”, followed by a verb such as support, handle, modify, send, return or delete;
- “out of scope”, “not supported”;
- “it is not” or “this tool is not”, followed by the, a, an, for, meant or intended, as in “It is not the prose guide; read that with get_guide”.
The rule does not run when AXC-D002 found a placeholder.
Known false negatives: the match is generous, so an incidental phrase can count as a boundary. “Returns only to the account owner” contains “only to” and passes, and “instead” counts wherever it appears. Read the description to confirm the boundary is about selection.
Known false positives: a boundary written in other words, such as “Read-only; for changes see update_invoice”, is not recognised. Rephrase it, or silence the rule with --disable AXC-D005.
Sources
- Paper: Anand and Chattaraj, “Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools”, arXiv 2608.04719, 5 August 2026 (preprint, not peer reviewed). https://arxiv.org/abs/2608.04719 . Defines six decoy types that mislead tool selection; the paper’s claim, across 8 models and 120 tasks, is that susceptibility varies about 36 times between models.
- Site guide: Write tool descriptions an agent can act on, agentexperience.tech. https://agentexperience.tech/insights/tool-descriptions/ . Asks every description to say when the tool is the wrong choice and to name the neighbouring tool.
Related evidence
Records in the AX evidence register that share a pattern tag with this rule. A shared tag means the record is about the same pattern, not that it tests this rule. Read the evidence class before the number.
- EV-0018: Skill rules that name a command or path change what agents do (Preprint). A study of 3,159 skills (preprint) found that adding a checkable rule raised the rate at which four coding agents took the required action by +0.23 on average, with the gain coming mainly from rules naming a command or path the old skill did not mention.
- EV-0019: Skill selection precision collapses as the skill pool grows (Preprint). A preprint reports that as the pool of available skills grew from 5 to 100, the precision with which agents actually used the right skill fell from 29.6% to 3.3%.
- EV-0020: Praise and list order move tool selection (Preprint). A preregistered preprint with two small OpenAI models found that stacked praise in a tool description raised its pick rate by about 43 percentage points, and that with identical listings the first-listed tool was picked about 72 points more often.
- EV-0001: An enum in the schema ends silent failures from example-only vocabularies (Preprint). SilentProbe (preprint) reports that a vocabulary a parameter description only exemplified ("e.g.") was missed on 88 of 88 attempts across twelve models, and that promoting it into the schema cut the failure to 0 of 89.
- EV-0003: Error text that names the next tool lifts recovery (Preprint). A preprint testing five OpenAI models reports that an expired-credential error naming a terminal command left 45% of tasks recovered, naming the server's login tool instead raised recovery to 84%, and on rate limits naming the call to repeat raised recovery from 6% to 88%.
- EV-0017: Injected skills lowered pass rates and raised token cost on average (Preprint). WebDev-Skills-Bench (preprint) found that injecting matched public skills reduced mean Pass@2 by 1.3% to 4.2% across four models and raised token cost by 72% to 394%, with gains in only 17% to 36% of skill-project pairs.
