ax-check rule
AXC-F009: Claimed CLI flag does not exist
The file uses a flag that neither the CLI's --help nor the subcommand's --help documents.
ax-check is a checker being prepared for release. This page documents the rule ahead of that release; see all 50 rules.
| Severity | error |
| Kind | Observed. Depends on the program and environment at run time. |
| Mode | ax-check drift |
| Applies to | Instruction files and manifests |
| Pattern tags | drift, agents-md, skills |
| Fix in one line | Update the instructions to the current flag, or document the flag in --help. |
The file uses a flag with a CLI, and neither the CLI’s --help nor the help for the subcommand documents that flag. The flag was renamed or removed. An agent that follows the instruction gets an “unknown option” error or, worse, a command that does something slightly different.
What it checks
You name the program with --cli "<command>". ax-check reads the files for code spans and fenced code blocks that start with that program, and collects each flag (a token starting with -, with any =value part removed). It captures <command> --help for the root flags, and <command> <sub> --help for the subcommands listed there. A flag is accepted if it appears in the root help flags, or in the help flags of the subcommand it was used with. The rule fires when it appears in neither. --help and -h are always accepted.
Passing --cli is your explicit request to run that program, so it works without --online.
Why it matters
Flags change meaning and names between versions. Instruction files are written once and copied around. Coding agents work with these files a great deal (in one preprint of 557 agentic coding sessions, instruction files and working notes were 60.5% of documentation interactions against 1.3% for API references; paper’s claim, not peer reviewed). A second preprint reports that the benefit of adding a rule to a skill came mainly from rules that name a command or path (paper’s claim, four agents, not peer reviewed). Neither study tested stale flags. That an agent will use the flags written in these files is this rule’s design rationale. A dead flag turns a working workflow into a failing one, and the failure shows up far from the cause.
How to fix
- Update the instructions to the flag that exists now.
- Or document the flag in
--help, if it exists but is undocumented. An undocumented flag is a risk for the same reason. - If the CLI has changed a lot, re-read its full
--helpand update the whole section.
Example
Find it:
ax-check drift --cli invoicekit SKILL.md
Before
---
name: invoice-export
description: Exports invoice figures to CSV. Use when the user asks for monthly numbers.
---
# Invoice export
```bash
invoicekit export --month 2026-09 --out report.csv
```
invoicekit export --help documents --month and --output, but not --out.
After
---
name: invoice-export
description: Exports invoice figures to CSV. Use when the user asks for monthly numbers.
---
# Invoice export
```bash
invoicekit export --month 2026-09 --output report.csv
```
How ax-check detects it
The extraction and the rule logic are deterministic, but the result depends on what the CLI’s help prints on the machine and version you run it against, so the same file can give different findings on another day. JSON and SARIF reports record when and where the check ran. It uses the same command extraction as AXC-F008. For each token that starts with - and looks like a flag (-x or --long-name), it records the flag and, if there is a subcommand directly after the program, links the flag to that subcommand. It then compares the flag with the root help flags plus the flags in <command> <sub> --help. If the subcommand itself is not a real command, the flag is skipped, because AXC-F008 already reports that.
To stay safe, ax-check does not capture help for subcommands whose names suggest a change in the world, such as delete, deploy, publish, release, create, login or run. A flag used with one of those is compared with the root help flags only, so it may be reported when it is in fact documented under the subcommand. Silence the rule in that case.
Known limits:
- A flag that belongs to a different program in a pipeline can be wrongly assigned when the command line is complex.
- Flags that exist but are only documented on a web page, not in
--help, are reported. This is intended: an agent working in a terminal sees--help. - If
--helpcannot be captured, ax-check stops with an error rather than guess.
Without --cli, this rule does not run. Add --dry-run to list the claims without running the program. To silence the rule, use --disable AXC-F009.
Sources
- Paper: “From Agent Behaviour to Agent-Friendly Documentation”, Gao and Chen, arXiv 2608.20195, 20 Aug 2026. https://arxiv.org/abs/2608.20195 . Paper’s claim: in 557 agentic coding sessions, agent-facing artefacts were 60.5% of documentation interactions versus 1.3% for API references. Preprint, not peer reviewed.
- Paper: “Agent Skill Evolution: How Revisions Affect Coding Agents”, Wang et al., arXiv 2610.04832, 4 Oct 2026. https://arxiv.org/abs/2610.04832 . Paper’s claim: across four agents an added rule raised the rate of the required action by +0.23 on average, mainly from rules that name a command or path the old skill did not mention. Preprint, not peer reviewed.
- Guideline: Command Line Interface Guidelines. https://clig.dev/ . A discoverable CLI documents its options in
--help.
Related evidence
Records in the AX evidence register that share a pattern tag with this rule. A shared tag means the record is about the same pattern, not that it tests this rule. Read the evidence class before the number.
- EV-0032: Vendor claim: agents never invoked a docs skill in 56% of eval cases (Vendor claim). Vercel states that in its Next.js 16 evals the docs skill was never invoked in 56% of cases, so the skill matched the 53% pass rate of no docs, while an 8KB docs index in AGENTS.md reached 100%.
- EV-0040: Agent skills recommended a package that does not exist (Vendor measurement). Merged pull requests in Vercel's agent plugin repository corrected skill instructions that recommended an npm package that is not published, and plugin guidance that advertised deployment cards the production MCP server does not expose.
- EV-0011: Coding agents mostly read instruction files, not docs sites (Preprint). An observational study of 557 agentic coding sessions (preprint) found that instruction files and working notes made up 60.5% of agents' documentation interactions, against 10.6% for classical technical documentation and 1.3% for API references.
- EV-0017: Injected skills lowered pass rates and raised token cost on average (Preprint). WebDev-Skills-Bench (preprint) found that injecting matched public skills reduced mean Pass@2 by 1.3% to 4.2% across four models and raised token cost by 72% to 394%, with gains in only 17% to 36% of skill-project pairs.
- EV-0018: Skill rules that name a command or path change what agents do (Preprint). A study of 3,159 skills (preprint) found that adding a checkable rule raised the rate at which four coding agents took the required action by +0.23 on average, with the gain coming mainly from rules naming a command or path the old skill did not mention.
- EV-0019: Skill selection precision collapses as the skill pool grows (Preprint). A preprint reports that as the pool of available skills grew from 5 to 100, the precision with which agents actually used the right skill fell from 29.6% to 3.3%.
