ax-check rule
AXC-C002: Error exits with status 0
An unknown command, unknown flag or missing value exited with status 0, so an agent reads failure as success.
ax-check is a checker being prepared for release. This page documents the rule ahead of that release; see all 50 rules.
| Severity | error |
| Kind | Observed. Depends on the program and environment at run time. |
| Mode | ax-check probe-cli |
| Applies to | Command-line programs, run as a subprocess |
| Pattern tags | exit-codes, false-success |
| Fix in one line | Exit non-zero on every error, including usage errors and aborted operations. |
An unknown subcommand, an unknown flag or a flag with a missing value ended with exit status 0. The exit code is the first thing a script or an agent reads. Status 0 tells it that everything worked.
What it checks
The rule looks at three opt-in probes (--probe errors): <cmd> axcheck-unknown-subcommand, <cmd> --axcheck-unknown-flag and a documented value-taking flag passed with no value (for example <cmd> --since). It fires for each of these probes that exits 0.
The unknown-subcommand probe only counts when the help output lists subcommands. A CLI with no subcommands may reasonably take the word as an ordinary argument, so exit 0 would not be an error there. The missing-value probe skips --help, --version, --json, --format, --output and -o, and prefers a flag that help shows with a <placeholder>. If help documents no usable flag, that probe does not run.
Why it matters
An agent decides what to do next from the exit code, often before it reads any text. If a failed call returns 0, the agent records success and moves on. A CI job does the same. The Command Line Interface Guidelines put it plainly: “Return zero exit code on success, non-zero on failure. Exit codes are how scripts determine whether a program succeeded or failed, so you should report this correctly.”
This is a known problem in real tools. Cloudflare documents that, in a non-interactive session, a destructive cf command without --force prints “Aborted.” to standard error and exits with status 0, and that a successful exit does not mean the resource was deleted (vendor documentation). A public issue on the same project, cloudflare/cf issue #94, reports that a CI job checking the exit status treats the requested deletion as successful even though cf refused to run it.
A related paper, Zhu et al. (arXiv 2609.35732, a preprint that is not peer reviewed), measures how often tool-using models report success after a failure. The paper’s claim is that false-success reports fell from 22.8% in the baseline to 0.8% with a structured evidence contract, across six models and 100 tasks. It studies what models report, not exit codes, so it is background on why a false success costs something, not evidence about exit codes themselves.
How to fix
Exit non-zero on every error: usage errors, unknown commands, bad flag values and aborted operations. Use 2 for usage errors and 1 for general failures, or document your own table. Also exit non-zero when a destructive action is refused because confirmation was not possible.
Example
Before
$ invoicer axcheck-unknown-subcommand </dev/null; echo $?
Unknown command.
0
After
$ invoicer axcheck-unknown-subcommand </dev/null; echo $?
error: unknown command "axcheck-unknown-subcommand".
Run: invoicer --help
2
A Node fix:
program.on('command:*', (operands) => {
console.error(`error: unknown command "${operands[0]}".`);
console.error('Run: invoicer --help');
process.exit(2);
});
How ax-check detects it
The three error probes run, when enabled with --probe errors, with no TTY and stdin closed. The rule fires when the exit status is exactly 0. A probe that timed out is skipped here, because AXC-C001 reports it. It does not judge the message, and it does not run any subcommand your CLI really has, so it never triggers a real action.
To reproduce by hand: invoicer axcheck-unknown-subcommand </dev/null; echo $?, then repeat with --axcheck-unknown-flag.
Known false negatives: a CLI that exits non-zero for the wrong reason, such as a crash, passes this rule. Known false positives: a wrapper that deliberately forwards unknown arguments to another program and returns that program’s status. Silence the rule with --disable AXC-C002 if that is intended.
Safety note: The error probes this rule needs are opt-in: run with --probe errors. Without it the rule is not checked, and the report lists it as not checked. Argument parsers normally reject an unknown subcommand, an unknown flag or a missing value before doing any work, but a program that does not validate its arguments will do its ordinary work instead. probe-cli executes your program under your own OS account. The temporary working directory and HOME it uses are not a sandbox: the program can still read and write anything your account can, and use the network. Probe a CLI you have not reviewed only inside a container or a throwaway virtual machine. Run ax-check probe-cli --dry-run -- invoicer first to see the probes that would run, and see the “Probe safety” section of the ax-check README for what is and is not isolated.
Sources
- Guideline: Command Line Interface Guidelines, clig.dev. https://clig.dev/ . Says to return zero on success and non-zero on failure so that scripts can tell the difference.
- Vendor documentation: Use cf with coding agents, Cloudflare, updated 2026-09-29. https://developers.cloudflare.com/cf/agents/ . Documents that a destructive command without
--forcein a non-interactive session prints “Aborted.” and exits with status 0. - Vendor issue: cloudflare/cf issue #94, “workflows delete exits 0 when noninteractive confirmation aborts the operation”. https://github.com/cloudflare/cf/issues/94 . Reports that a CI job treats the aborted deletion as successful.
- Paper: Zhu et al., “Failure-Transparent Agents: Benchmarking Post-Failure Reporting in Tool-Using Language Models”, arXiv 2609.35732, 28 Sep 2026. https://arxiv.org/abs/2609.35732 . The paper’s claim: false-success reports fell from 22.8% to 9.3% to 0.8% across three conditions (six models, 100 tasks).
Related evidence
Records in the AX evidence register that share a pattern tag with this rule. A shared tag means the record is about the same pattern, not that it tests this rule. Read the evidence class before the number.
- EV-0036: A CLI exits 0 when a destructive command is refused (Independent measurement). Cloudflare's cf CLI documents that in a non-interactive session a destructive command without --force prints 'Aborted.' and exits with status 0, and a public issue reproduces this for workflows delete.
- EV-0037: Destructive and bulk commands that report success but do nothing (Independent measurement). Users of Cloudflare's cf beta report commands that exit 0 without doing the work: R2 deletes of keys containing '/' (a cleanup script reported 2,220 objects deleted that were all still present) and a secrets bulk update that deployed an empty change to 100%.
- EV-0001: An enum in the schema ends silent failures from example-only vocabularies (Preprint). SilentProbe (preprint) reports that a vocabulary a parameter description only exemplified ("e.g.") was missed on 88 of 88 attempts across twelve models, and that promoting it into the schema cut the failure to 0 of 89.
- EV-0002: Constraints stated only in prose produce silent failures on live APIs (Preprint). SilentProbe (preprint) reports that only 7.5% of 2,501 public OpenAPI documents declare an enum, and that on live endpoints machine-checkable constraints returned an honest error in 111 of 111 cases while prose-only constraints failed silently in 44 of 61.
- EV-0004: Naming recovery tools is the active ingredient in failure receipts (Preprint). Outcome Monitors (preprint) reports that receipts naming a violated outcome and the public recovery tools raised ToolMaze completion from 10.9% to 28.1% across four models, and that removing the list of recovery tools eliminated the gain.
- EV-0005: Idempotency keys cut duplicate writes from 28% to 4% (Preprint). LIMBO (preprint) reports that offering an idempotency key on every write cut duplicate side effects from 28% to 4% of episodes because agents use keys when they exist, and that agents reported success in 90% of the episodes in which they had duplicated an effect.
