Lab
Checks for what agents read and run.
ax-check is a rule-based checker for agent-facing surfaces. It reads tool and skill descriptions the way an agent reads them before choosing, runs a command-line program the way an agent runs it, and checks that instruction files still match reality.
The rules are deterministic: they use no model, so the same text gives the same lint findings. CLI probes and online drift checks are different, because their results depend on the program, the machine and the network at the time they run. Each report records the ax-check version, the environment and the time. A rule that matches the same phrase every time is reproducible, not proof that a description is correct or that fixing it improves outcomes. See three kinds of rule.
Status: ax-check is a checker being prepared for release, so there is no install command here yet. These pages document its 50 rules (version 0.2.0) ahead of that release.
It gives no score. Each finding names a rule, the evidence for it and a fix, so you decide what matters for your product.
Jump to: Kinds of rule · D: descriptions and schemas (26) · C: CLI behaviour (13) · F: drift and facts (11)
How to read a rule
Each rule has an ID such as AXC-D007. The letter says what kind of check it is: D for descriptions and schemas, C for CLI behaviour, F for drift between instructions and facts.
Severity is the checker's default judgement of how serious a finding is. An error is a broken or missing contract, such as a failed command that exits with status 0. A warn is a gap that leaves the agent guessing, such as allowed values given only in prose. An info finding is worth knowing about but may be fine for your product. Some rules are milder on surfaces with less room, such as a one-line CLI summary.
Kind says what a finding can tell you: conformance, heuristic or observed. The next section explains each one.
Pattern tags link to the matching section of the AX evidence register, which holds the public findings behind the rules. Each rule page lists the register records that share its tags. For dated news on the same patterns, see What changed in AX.
Three kinds of rule
- Conformance (13 rules). Checks a stated contract or specification, such as a required field, a length limit in a published specification, or an example against its own schema. A finding is a fact about the file.
- Heuristic (19 rules). Matches phrases, patterns or thresholds, such as "Use this when" or a word count. It gives the same answer for the same text, but that answer can be wrong: each rule page lists its known false positives and negatives. Treat a finding as a prompt to look, not a verdict.
- Observed (18 rules). Depends on what a program, server or registry did when the check ran, such as an exit status, a hang or a missing page. The rule logic is fixed, but the program, the network and the registries are not, so a later run or another machine can give different findings.
None of the kinds measures whether agents succeed. A description can pass every rule and still mislead an agent. Where a rule cites a study, the numbers are that study's claim for its own models and tasks.
D: descriptions and schemas
Rules that read what an agent reads when it picks a tool: MCP tool lists, OpenAPI documents, SKILL.md files and CLI help text. Command: ax-check lint.
C: CLI behaviour
Rules that run a command-line program the way an agent runs it, with no terminal and stdin closed, and check what comes back. Command: ax-check probe-cli.
F: drift and facts
Rules that check whether instruction files such as AGENTS.md, llms.txt and SKILL.md still match reality: links, packages, tools, commands and flags. Command: ax-check drift.
| ID | Rule | Severity | Kind | Pattern tags |
|---|---|---|---|---|
| AXC-F001 | Claimed URL is broken | error | Observed | drift, docs-for-agents, llms-txt |
| AXC-F002 | Claimed URL could not be verified | info | Observed | drift, docs-for-agents |
| AXC-F003 | Claimed npm package does not exist | error | Observed | drift, prompt-injection |
| AXC-F004 | Claimed PyPI package does not exist | error | Observed | drift, prompt-injection |
| AXC-F005 | Claimed package version does not exist | error | Observed | drift, server-cards |
| AXC-F006 | Claimed MCP tool is not served | error | Observed | drift, discovery, server-cards |
| AXC-F007 | Served MCP tool is not mentioned | info | Observed | drift, discovery |
| AXC-F008 | Claimed CLI command does not exist | error | Observed | drift, agents-md, skills |
| AXC-F009 | Claimed CLI flag does not exist | error | Observed | drift, agents-md, skills |
| AXC-F010 | Linked local file is missing | error | Conformance | drift, agents-md, skills |
| AXC-F011 | Claims extracted but not verified | info | Heuristic | drift |
What the checks cannot tell you
The description rules match phrases and identifiers in English. They will miss a boundary written in unusual words and sometimes flag an incidental phrase; each rule page lists its known false positives and negatives. The CLI probes observe a program from outside. By default they run only --help and --version, so they cannot see what it does with real arguments or credentials. They run the program under your own account in a temporary folder, which is not a sandbox, so probe a CLI you have not reviewed only inside a container. None of the checks runs a model, so they cannot tell you whether an agent will pick your tool. Pair them with an evaluation of one real job, for example with the Open Agent-Readiness Rubric.
ax-check and its rule descriptions are by Himadri Mishra. The licence will be confirmed before the source repository is published.
