ax-check rule
AXC-C004: Error does not name a next step
The error output names no next step: no valid values, no command to run, no pointer to help.
ax-check is a checker being prepared for release. This page documents the rule ahead of that release; see all 50 rules.
| Severity | warn |
| Kind | Heuristic. A pattern match: a prompt to look, not a verdict. |
| Mode | ax-check probe-cli |
| Applies to | Command-line programs, run as a subprocess |
| Pattern tags | error-recovery |
| Fix in one line | End every error with the exact command, flag or value to use next. |
The error output says that something went wrong but not what to do next. There is no valid value, no command to run and no pointer to help. An agent that reads it has to guess.
What it checks
The rule looks at the error probes, which run only with --probe errors: unknown subcommand, unknown flag and missing value. It reads the text the program wrote to stdout and stderr, and only when the probe exited non-zero (an exit of 0 is reported by AXC-C002 instead). It fires when the text contains none of these signals: --help or help, usage, did you mean, try, run, valid, available, expected, must be, one of, see, supported, allowed, possible, choose, instead, for example, e.g., pass, or a commands: heading. It also fires when the program printed nothing at all.
The unknown-subcommand probe only counts when the help output lists subcommands. The missing-value probe skips --help, --version, --json, --format, --output and -o, and prefers a flag that help shows with a <placeholder>.
Why it matters
A person can search the web or ask a colleague. An agent has only the text on screen. The Command Line Interface Guidelines say to “catch errors and rewrite them for humans” and describe discoverable CLIs as ones that “suggest what command to run next, suggest what to do when there is an error.”
Recent research points the same way, though it studies MCP servers rather than command-line tools. Xu and Wu (arXiv 2609.35381, a preprint that is not peer reviewed) looked at 150 widely used MCP servers. The paper’s claim is that 949 of 3,001 error messages tell the caller what to do next. With five OpenAI models acting only through tools, the paper reports that on expired credentials a terminal command in the message left 45% of tasks recovered, while naming the server’s login tool raised recovery to 84%. On a rate limit, “Wait before retrying.” left 6%, while naming the call to repeat gave 88%. These are the paper’s numbers for its own model set and tasks. They suggest that the exact next action matters, not that your tool will see the same figures.
How to fix
End every error with the exact command, flag or value to use next. Name real things: the flag spelled out, the list of valid values, or the command that shows them.
Example
Before
$ invoicer export --format </dev/null; echo $?
Invalid input.
1
After
$ invoicer export --format </dev/null; echo $?
error: --format needs a value. Valid values: csv, json, pdf.
Run: invoicer export --format json
1
A Node fix:
console.error(`error: --format needs a value. Valid values: ${formats.join(', ')}.`);
console.error('Run: invoicer export --format json');
process.exit(1);
How ax-check detects it
This is a deliberately broad keyword test, so it is cheap and predictable. It checks for the presence of a next-step signal, not whether the signal is correct. A message that says “see” or “run” in an unhelpful way will pass. A message with a good next step in different words, for example “pass a number between 1 and 5”, may be flagged.
To reproduce by hand: invoicer export --format </dev/null; echo $? and read the output. Empty output on a non-zero exit is also flagged, with its own message.
Silence the rule with --disable AXC-C004 if your messages use wording the heuristic cannot see.
Safety note: The error probes this rule needs are opt-in: run with --probe errors. Without it the rule is not checked, and the report lists it as not checked. Argument parsers normally reject an unknown subcommand, an unknown flag or a missing value before doing any work, but a program that does not validate its arguments will do its ordinary work instead. probe-cli executes your program under your own OS account. The temporary working directory and HOME it uses are not a sandbox: the program can still read and write anything your account can, and use the network. Probe a CLI you have not reviewed only inside a container or a throwaway virtual machine. Run ax-check probe-cli --dry-run -- invoicer first to see the probes that would run, and see the “Probe safety” section of the ax-check README for what is and is not isolated.
Sources
- Guideline: Command Line Interface Guidelines, clig.dev. https://clig.dev/ . Says to catch errors and rewrite them for humans, and that discoverable CLIs suggest what to do when there is an error.
- Paper: Xu and Wu, “MCP Error Messages Written for Developers Hurt the Most Capable Agents Most”, arXiv 2609.35381, submitted 28 Sep 2026. https://arxiv.org/abs/2609.35381 . The paper’s claim: 949 of 3,001 error messages in 150 MCP servers name a next step; naming the login tool raised recovery from 45% to 84%, and naming the call to repeat on a rate limit raised it from 6% to 88% (five OpenAI models).
- Site guide: Design for recovery, agentexperience.tech. https://agentexperience.tech/insights/design-for-recovery/ . Explains how to write errors an agent can act on.
Related evidence
Records in the AX evidence register that share a pattern tag with this rule. A shared tag means the record is about the same pattern, not that it tests this rule. Read the evidence class before the number.
- EV-0002: Constraints stated only in prose produce silent failures on live APIs (Preprint). SilentProbe (preprint) reports that only 7.5% of 2,501 public OpenAPI documents declare an enum, and that on live endpoints machine-checkable constraints returned an honest error in 111 of 111 cases while prose-only constraints failed silently in 44 of 61.
- EV-0003: Error text that names the next tool lifts recovery (Preprint). A preprint testing five OpenAI models reports that an expired-credential error naming a terminal command left 45% of tasks recovered, naming the server's login tool instead raised recovery to 84%, and on rate limits naming the call to repeat raised recovery from 6% to 88%.
- EV-0004: Naming recovery tools is the active ingredient in failure receipts (Preprint). Outcome Monitors (preprint) reports that receipts naming a violated outcome and the public recovery tools raised ToolMaze completion from 10.9% to 28.1% across four models, and that removing the list of recovery tools eliminated the gain.
- EV-0005: Idempotency keys cut duplicate writes from 28% to 4% (Preprint). LIMBO (preprint) reports that offering an idempotency key on every write cut duplicate side effects from 28% to 4% of episodes because agents use keys when they exist, and that agents reported success in 90% of the episodes in which they had duplicated an effect.
