Lab

The AX evidence register.

A public record of measured and observed findings about agent experience: how AI agents discover, choose, call and recover from the tools, APIs, CLIs, docs and approval flows that products offer them. Each record states one finding, gives the numbers as the source reported them, and says what kind of evidence it is and what it does not show.

40 records, last updated . Every source was fetched and checked on the date shown in its record. The register records what others have published, with their caveats; it does not rank anyone, re-run any result or publish experiments of its own.

Machine-readable copy: /evidence.json, with its JSON Schema. Cite the original source for any number you use. The register's licence will be confirmed before its source repository is published.

Jump to: Read this first · Evidence classes · Browse by pattern · All records

Read this first

  • Most records are preprints. 27 of the 40 records come from papers that have not been peer reviewed. Many are by one author or a small team, and some authors have a stake in the result.
  • Vendor claims are claims. A figure from a launch post or a press release is labelled as the vendor's claim. It can show a direction, not a size.
  • Every number belongs to its agent profile. A result holds for the models, harness and dates the source tested. Newer models can behave differently, and one record shows the effect growing with a newer model.
  • Many records were checked against the abstract only. 27 records say "abstract only": the numbers are in the live abstract, but the full text was not re-checked.
  • Some patterns have no records. As of the last update there is no record that directly measures whether llms.txt, ARD, Server Cards or WebMCP change agent task success. An empty pattern below is information too.

Evidence classes

Read the class before you read the number. Each class lists its records, so this list also works as a filter.

  • Peer reviewed (0). Published after independent review at a journal or conference. Review checks the method, not the truth of the result. The result still belongs to the models and dates tested.
  • Preprint (27). A paper posted publicly, for example on arXiv, without peer review. Unreviewed. Many are by one author or a small team, and some authors have a stake in the result. Most AX research is in this class today. Records: EV-0001, EV-0002, EV-0003, EV-0004, EV-0005, EV-0006, EV-0007, EV-0008, EV-0009, EV-0010, EV-0011, EV-0012, EV-0013, EV-0014, EV-0015, EV-0016, EV-0017, EV-0018, EV-0019, EV-0020, EV-0021, EV-0022, EV-0023, EV-0024, EV-0025, EV-0026, EV-0027.
  • Independent measurement (4). Someone with no stake in the product measured or reproduced a behaviour and published how, for example a bug report with steps to reproduce. Often a single report on one version. Software changes, so check again before you rely on it. Records: EV-0036, EV-0037, EV-0038, EV-0039.
  • Vendor measurement (5). A vendor reports its own measurement with enough method to judge it, such as run counts, models and the harness. The vendor chose the tasks, ran the evaluation and wrote it up. Nobody else has reproduced it unless the record says so. Records: EV-0028, EV-0029, EV-0030, EV-0031, EV-0040.
  • Vendor claim (4). A vendor states a figure without a published method. A claim, not a measurement. The definitions behind it are usually unstated, and the figure often comes from a launch or funding announcement. It can show a direction, not a size. Records: EV-0032, EV-0033, EV-0034, EV-0035.
  • Anecdote (0). A single first-hand account without a method. Shows that something can happen, not how often it happens.

Each record also has a finding type: a measurement, a negative result (the change made things worse or did not work), a null result (no detectable effect), an observation (behaviour or prevalence, not an intervention) or a vendor claim. Negative and null results are recorded on purpose.

Browse by pattern

Records and ax-check rules grouped by the shared AX pattern vocabulary. A record can carry several tags.

PatternRecordsax-check rules
discoveryEV-0013, EV-0019, EV-0023, EV-0027AXC-D001, AXC-D014, AXC-D022, AXC-D023, AXC-C005, AXC-C009, AXC-F006, AXC-F007
selectionEV-0019, EV-0020, EV-0021AXC-D001, AXC-D002, AXC-D003, AXC-D004, AXC-D005, AXC-D006, AXC-D013, AXC-D014, AXC-D015, AXC-D020, AXC-D021
descriptionEV-0001, EV-0003, EV-0018, EV-0020, EV-0026AXC-D001, AXC-D002, AXC-D003, AXC-D004, AXC-D005, AXC-D006, AXC-D007, AXC-D008, AXC-D009, AXC-D013, AXC-D015, AXC-D020
schema-enumEV-0001, EV-0002AXC-D007, AXC-D008, AXC-D012, AXC-D018, AXC-D019
error-recoveryEV-0002, EV-0003, EV-0004, EV-0005AXC-D012, AXC-D018, AXC-C003, AXC-C004, AXC-C005
false-successEV-0001, EV-0002, EV-0004, EV-0005, EV-0006, EV-0007, EV-0008, EV-0024, EV-0036, EV-0037AXC-D007, AXC-D012, AXC-C002, AXC-C007, AXC-C013
exit-codesEV-0036, EV-0037AXC-C002, AXC-C003
non-interactiveEV-0036AXC-C001, AXC-C006, AXC-C007, AXC-C010, AXC-C011, AXC-C012, AXC-C013
dry-runEV-0038AXC-C008
idempotencyEV-0005AXC-D010
approvalEV-0014, EV-0015, EV-0016, EV-0034AXC-D009, AXC-D010, AXC-D011, AXC-C008
confirmationEV-0014, EV-0016, EV-0034, EV-0036AXC-D009, AXC-D010, AXC-C001
auth-scopesEV-0023none
secretsEV-0038none
claim-later-onboardingnone yetnone
context-budgetEV-0017, EV-0022, EV-0027, EV-0039AXC-D016, AXC-D017, AXC-D021, AXC-D024, AXC-D026, AXC-C010
dynamic-toolsEV-0022, EV-0023AXC-D017
code-modeEV-0009none
docs-for-agentsEV-0011, EV-0012, EV-0027, EV-0029, EV-0030, EV-0032AXC-D019, AXC-C006, AXC-C009, AXC-F001, AXC-F002
llms-txtnone yetAXC-F001
ardnone yetnone
server-cardsnone yetAXC-F005, AXC-F006
agents-mdEV-0011, EV-0032AXC-F008, AXC-F009, AXC-F010
skillsEV-0017, EV-0018, EV-0019, EV-0030, EV-0032, EV-0040AXC-D004, AXC-D005, AXC-D022, AXC-D023, AXC-D024, AXC-D025, AXC-D026, AXC-F008, AXC-F009, AXC-F010
driftEV-0039, EV-0040AXC-D025, AXC-C012, AXC-F001, AXC-F002, AXC-F003, AXC-F004, AXC-F005, AXC-F006, AXC-F007, AXC-F008, AXC-F009, AXC-F010, AXC-F011
handoffEV-0006none
evaluationEV-0006, EV-0007, EV-0008, EV-0010, EV-0024, EV-0025, EV-0031none
measurementEV-0008, EV-0010, EV-0013, EV-0024, EV-0025, EV-0028, EV-0031, EV-0033, EV-0035none
propensityEV-0021, EV-0029, EV-0030, EV-0032, EV-0033none
costEV-0010, EV-0017, EV-0022, EV-0028, EV-0031AXC-D016, AXC-D017
prompt-injectionEV-0015, EV-0026AXC-D011, AXC-D020, AXC-F003, AXC-F004
webmcpnone yetnone

All records

One line per record. Open a record for its effect sizes, agent profile, conflicts of interest, source and limitations.

IDFindingClassPatterns
EV-0001An enum in the schema ends silent failures from example-only vocabulariesPreprintschema-enum, description, false-success
EV-0002Constraints stated only in prose produce silent failures on live APIsPreprintschema-enum, error-recovery, false-success
EV-0003Error text that names the next tool lifts recoveryPreprinterror-recovery, description
EV-0004Naming recovery tools is the active ingredient in failure receiptsPreprinterror-recovery, false-success
EV-0005Idempotency keys cut duplicate writes from 28% to 4%Preprintidempotency, false-success, error-recovery
EV-0006An evidence contract cuts false-success reports after tool failuresPreprintfalse-success, handoff, evaluation
EV-0007Browser agents declare success on most of their failuresPreprintfalse-success, evaluation
EV-0008Declared completion and logged completion differ by up to 41 pointsPreprintfalse-success, evaluation, measurement
EV-0009Tools as code match or beat JSON tool calls for most modelsPreprintcode-mode
EV-0010Agent scaffolding, not MCP versus CLI, drove costPreprintcost, evaluation, measurement
EV-0011Coding agents mostly read instruction files, not docs sitesPreprintdocs-for-agents, agents-md
EV-0012Compact documentation did not help when the source was presentPreprintdocs-for-agents
EV-0013FAQ blocks and structured data showed no citation effect within a domainPreprintdiscovery, measurement
EV-0014Allow/ask/never policies blocked less overreach than per-action approvalPreprintapproval, confirmation
EV-0015Approvals that outlive their task raise attack successPreprintapproval, prompt-injection
EV-0016Approval records omit the effects a command goes on to triggerPreprintapproval, confirmation
EV-0017Injected skills lowered pass rates and raised token cost on averagePreprintskills, cost, context-budget
EV-0018Skill rules that name a command or path change what agents doPreprintskills, description
EV-0019Skill selection precision collapses as the skill pool growsPreprintskills, selection, discovery
EV-0020Praise and list order move tool selectionPreprintdescription, selection
EV-0021Agents leave available tools unused: the adoption gapPreprintpropensity, selection
EV-0022Search-and-execute meta-tools cut catalogue tokens by 99% in productionPreprintcontext-budget, dynamic-tools, cost
EV-0023Hiding tools is not enforcing permissionsPreprintauth-scopes, dynamic-tools, discovery
EV-0024Harnesses alter shell calls and the wrong action runs silentlyPreprintfalse-success, evaluation, measurement
EV-0025Run-to-run variance dominates, and a verification tool beats a verification promptPreprintevaluation, measurement
EV-0026Prompt injection split across tool channels evades defencesPreprintprompt-injection, description
EV-0027A capable agent skipped the index and guessed the pagePreprintdocs-for-agents, discovery, context-budget
EV-0028A JSON input mode for a CLI raised agent cost 4x to 11xVendor measurementcost, measurement
EV-0029A warning that names the failing plan redirected agents; a tip did notVendor measurementdocs-for-agents, propensity
EV-0030Adding the context7 MCP server gave no lift; its tools went unusedVendor measurementpropensity, docs-for-agents, skills
EV-0031A model with cheaper tokens cost 3.7x more per runVendor measurementcost, measurement, evaluation
EV-0032Vendor claim: agents never invoked a docs skill in 56% of eval casesVendor claimpropensity, skills, agents-md, docs-for-agents
EV-0033Vendor claim: agents were 48% of Wrangler CLI useVendor claimmeasurement, propensity
EV-0034Vendor claim: users approved about 93% of permission promptsVendor claimapproval, confirmation
EV-0035Vendor claim: 70% of new Supabase databases are created by agents or AI toolsVendor claimmeasurement
EV-0036A CLI exits 0 when a destructive command is refusedIndependent measurementexit-codes, non-interactive, false-success, confirmation
EV-0037Destructive and bulk commands that report success but do nothingIndependent measurementexit-codes, false-success
EV-0038Dry-run output printed secrets until a fix redacted themIndependent measurementdry-run, secrets
EV-0039Launch post and documentation disagree on agent output formatIndependent measurementdrift, context-budget
EV-0040Agent skills recommended a package that does not existVendor measurementdrift, skills

How to cite and propose a record

Cite the original source for any number; the register is a pointer, not the origin of a finding. To cite a record, keep its class and model set with the figure, and link its page, for example: "a preprint testing five OpenAI models reports recovery rising from 45% to 84% (AX evidence register, EV-0003)".

Know of a finding that belongs here, or a mistake in one? Send the source through the contact page. A record needs a public source, the numbers exactly as reported and an honest evidence class. Dated news on the same patterns is in What changed in AX.