ax-check rule
AXC-D026: SKILL.md over 500 lines
SKILL.md is longer than the 500 lines the specification recommends, so activation loads a lot of context.
ax-check is a checker being prepared for release. This page documents the rule ahead of that release; see all 50 rules.
| Severity | info |
| Kind | Conformance. A finding is a fact about the input. |
| Mode | ax-check lint |
| Applies to | SKILL.md |
| Pattern tags | skills, context-budget |
| Fix in one line | Move reference material into files under references/ and link to them. |
SKILL.md is longer than the 500 lines the Agent Skills specification recommends. When a skill activates, its whole body is loaded into the agent’s context, so a long body costs context on every use, including the parts the task does not need. This is information, not a failure: the fix is usually to move reference material into separate files.
What it checks
ax-check counts the lines of each SKILL.md file, frontmatter and body together, and reports a file longer than 500 lines.
Why it matters
The Agent Skills specification says: “Keep your main SKILL.md under 500 lines. Move detailed reference material to separate files.” Files in references/, scripts/ and assets/ are loaded only when required, so detail that lives there costs nothing until the agent needs it.
Skill bodies are not free. Wang et al. (arXiv 2610.04832, 4 October 2026) report that loading a skill’s body raised episode tokens by 50% on average across four agents (the paper’s claim). The same paper reports that the gain from adding rules to a skill came mainly from rules that name a command or path the old skill did not mention (also the paper’s claim).
How to fix
- Keep in the body only what the agent needs on every use: the steps, the commands and the decisions.
- Move tables, long examples, API references and background into files under
references/. - Link to each file from the step that needs it, with a relative path, so the agent loads it only then.
Example
Before
---
name: invoice-review
description: Checks a draft invoice before it is sent. Use this when asked to review an invoice.
---
# Invoice review
1. Open the draft invoice and read every line item.
2. Check the tax rate for the customer's country.
## Tax rates by country
(a 300-line table)
## Payment terms reference
(250 lines)
After
---
name: invoice-review
description: Checks a draft invoice before it is sent. Use this when asked to review an invoice.
---
# Invoice review
1. Open the draft invoice and read every line item.
2. Check the tax rate in [the tax table](references/tax-rates.md).
3. Check the due date against [the payment terms](references/payment-terms.md).
How ax-check detects it
ax-check counts every line of SKILL.md: the opening ---, the frontmatter, the closing --- and the body. It removes a leading byte order mark and converts Windows line endings first, and a final line break at the end of the file does not add an extra line. A file with more than 500 lines is reported. The finding gives the count and points at the first line of the body. If the frontmatter is missing or not closed, the whole file is counted in the same way.
What it ignores: files under references/ and other bundled files are not counted. AXC-D017 reports the approximate tokens that skill bodies add when they activate.
Known false positives: blank lines and lines inside code blocks count like any other line, so a file padded with long code samples may be reported although its instructions are short. Moving those samples into references/ is usually the right fix anyway.
If the finding does not apply, silence it with --disable AXC-D026.
Sources
- Specification: Agent Skills specification, agentskills.io, retrieved 2026-10-08. https://agentskills.io/specification. “Keep your main SKILL.md under 500 lines. Move detailed reference material to separate files.” Bundled files load only when required.
- Paper: Wang et al., “Agent Skill Evolution: How Revisions Affect Coding Agents”, arXiv 2610.04832, 4 October 2026. https://arxiv.org/abs/2610.04832. The paper’s claim, across four agents: loading a skill’s body raised episode tokens by 50% on average; added rules raised the rate of the required action by +0.23 on average, mainly from rules that name a command or path.
Related evidence
Records in the AX evidence register that share a pattern tag with this rule. A shared tag means the record is about the same pattern, not that it tests this rule. Read the evidence class before the number.
- EV-0017: Injected skills lowered pass rates and raised token cost on average (Preprint). WebDev-Skills-Bench (preprint) found that injecting matched public skills reduced mean Pass@2 by 1.3% to 4.2% across four models and raised token cost by 72% to 394%, with gains in only 17% to 36% of skill-project pairs.
- EV-0018: Skill rules that name a command or path change what agents do (Preprint). A study of 3,159 skills (preprint) found that adding a checkable rule raised the rate at which four coding agents took the required action by +0.23 on average, with the gain coming mainly from rules naming a command or path the old skill did not mention.
- EV-0019: Skill selection precision collapses as the skill pool grows (Preprint). A preprint reports that as the pool of available skills grew from 5 to 100, the precision with which agents actually used the right skill fell from 29.6% to 3.3%.
- EV-0022: Search-and-execute meta-tools cut catalogue tokens by 99% in production (Preprint). PayPal authors report (preprint) that exposing two meta-tools, search and execute, over 2,000+ MCP tools cut tool-token consumption in production from 140.2k tokens (70.1% of context) to 1.3k tokens (0.8%).
- EV-0027: A capable agent skipped the index and guessed the page (Preprint). A preregistered ablation on a 709-page Markdown wiki (preprint) found that a capable tool-using agent never loaded the compact catalogue index, inferring page paths from the question instead, while retrieval-based access kept answer quality non-inferior and cut cost by about a third to over half.
- EV-0030: Adding the context7 MCP server gave no lift; its tools went unused (Vendor measurement). Microsoft reports that adding the context7 MCP server to an anti-hallucination skill gave no meaningful lift on an SPFx upgrade: its tools did not load in 3 of 5 runs and were not called in the other 2, while telling the agent to use CLI for Microsoft 365 raised configuration correctness from 30/80 to 75/80.
