ax-check rule
AXC-C007: Documented JSON output does not parse
The command succeeded with the documented JSON flag but stdout was not valid JSON or JSON lines.
ax-check is a checker being prepared for release. This page documents the rule ahead of that release; see all 50 rules.
| Severity | error |
| Kind | Observed. Depends on the program and environment at run time. |
| Mode | ax-check probe-cli |
| Applies to | Command-line programs, run as a subprocess |
| Pattern tags | non-interactive, false-success |
| Fix in one line | With the JSON flag, write only JSON to stdout and send everything else to stderr. |
The help text documents a JSON flag. The command succeeded with that flag, but what it printed to stdout was not valid JSON or JSON lines. This is worse than having no JSON flag, because the agent trusts the promise and then fails to parse.
What it checks
The json probe is opt-in: it runs only with --probe json, and only if the help text documents a JSON selector. It runs <cmd> --json (or --jsonl, --ndjson), or the documented selector with its value as separate arguments, such as --format json, --output json or -o json. The rule fires when the exit status is 0, stdout is not empty, and stdout is neither a single JSON document nor a series of JSON lines.
Why it matters
A JSON flag is a contract. A banner, a progress line or a “Done!” message on stdout breaks it, and the whole output stops parsing. The Command Line Interface Guidelines say to send messaging to stderr, which is how a program keeps stdout clean for the result. Failing to parse also tends to hide the real answer. If a program printed the data and then a status line, an agent may discard both.
How to fix
With the JSON flag, write only JSON to stdout. Send banners, progress, warnings and deprecation notices to stderr, or suppress them. Test the flag with a parser in your own tests: invoicer list --json | jq ..
Example
Before
$ invoicer list --json </dev/null
Fetching invoices...
[{"id":1042,"status":"paid"}]
Done.
$ invoicer list --json </dev/null | jq .
jq: error: Invalid numeric literal at line 1, column 9
After
$ invoicer list --json </dev/null
[{"id":1042,"status":"paid"}]
$ invoicer list --json </dev/null 2>/dev/null | jq length
1
A Node fix:
function log(message) {
// Progress goes to stderr so that stdout stays valid JSON.
console.error(message);
}
log('Fetching invoices...');
process.stdout.write(JSON.stringify(rows) + '\n');
How ax-check detects it
ax-check tries to parse stdout as one JSON value. If that fails, it accepts output of two or more non-empty lines that each parse as JSON. If neither works, the rule fires. It checks stdout only, so noise on stderr is fine.
It deliberately ignores a non-zero exit and an empty stdout. Those are different problems, and an empty stdout may be a valid result for a command with nothing to list. A command that needs extra arguments to produce data may exit non-zero on the json probe and never reach this check. If the help text documents no JSON flag, there is no json probe, and AXC-C006 applies instead. When the probe was not enabled, exited non-zero or printed nothing, AXC-C013 reports that the output was not verified, so a documented flag is never passed silently.
To reproduce by hand: invoicer list --json </dev/null 2>/dev/null | python3 -c 'import json,sys; json.load(sys.stdin)'.
Silence the rule with --disable AXC-C007 only if the flag is documented as something other than JSON.
Safety note: The json probe is opt-in (--probe json), because it runs your command with its JSON selector added, which on some CLIs does the command’s ordinary work. Point it at a read-only command such as invoicer list. Without it, AXC-C013 reports that the output was not verified. probe-cli executes your program under your own OS account. The temporary working directory and HOME it uses are not a sandbox: the program can still read and write anything your account can, and use the network. Probe a CLI you have not reviewed only inside a container or a throwaway virtual machine. Run ax-check probe-cli --dry-run -- invoicer first to see the probes that would run, and see the “Probe safety” section of the ax-check README for what is and is not isolated.
Sources
- Guideline: Command Line Interface Guidelines, clig.dev. https://clig.dev/ . Says to display output as formatted JSON if
--jsonis passed and to send messaging to stderr.
Related evidence
Records in the AX evidence register that share a pattern tag with this rule. A shared tag means the record is about the same pattern, not that it tests this rule. Read the evidence class before the number.
- EV-0036: A CLI exits 0 when a destructive command is refused (Independent measurement). Cloudflare's cf CLI documents that in a non-interactive session a destructive command without --force prints 'Aborted.' and exits with status 0, and a public issue reproduces this for workflows delete.
- EV-0001: An enum in the schema ends silent failures from example-only vocabularies (Preprint). SilentProbe (preprint) reports that a vocabulary a parameter description only exemplified ("e.g.") was missed on 88 of 88 attempts across twelve models, and that promoting it into the schema cut the failure to 0 of 89.
- EV-0002: Constraints stated only in prose produce silent failures on live APIs (Preprint). SilentProbe (preprint) reports that only 7.5% of 2,501 public OpenAPI documents declare an enum, and that on live endpoints machine-checkable constraints returned an honest error in 111 of 111 cases while prose-only constraints failed silently in 44 of 61.
- EV-0004: Naming recovery tools is the active ingredient in failure receipts (Preprint). Outcome Monitors (preprint) reports that receipts naming a violated outcome and the public recovery tools raised ToolMaze completion from 10.9% to 28.1% across four models, and that removing the list of recovery tools eliminated the gain.
- EV-0005: Idempotency keys cut duplicate writes from 28% to 4% (Preprint). LIMBO (preprint) reports that offering an idempotency key on every write cut duplicate side effects from 28% to 4% of episodes because agents use keys when they exist, and that agents reported success in 90% of the episodes in which they had duplicated an effect.
- EV-0006: An evidence contract cuts false-success reports after tool failures (Preprint). Failure-Transparent Agents (preprint) reports that after a required tool failed, six models falsely reported success in 22.8% of responses by default, 9.3% with a transparency instruction and 0.8% with a structured evidence contract.
