Lab

What changed in AX.

A digest of public changes that affect how AI agents find, choose, use and recover with products: launches, documentation, specifications and preprints. A new edition appears roughly fortnightly and covers the two weeks before it.

Each entry says what changed in one or two sentences and links to its source. Every source was checked live on the date shown for the edition. Figures a vendor states without a method are labelled as vendor claims, and paper figures as the paper's claim. The digest adds no findings or rankings of its own.

Follow: RSS feed · JSON feed · Evidence register · ax-check rules

How to read an entry

Each entry has a date, a short account of what changed, the source, its pattern tags and an evidence class. "Affects" links to the guides, rubric checks, ax-check rules and evidence records the change bears on. Tags open the matching pattern in the evidence register.

  • Vendor documentation. A vendor’s own release notes, documentation, code or launch post describing what it shipped. It shows that the change exists, not that it works well.
  • Specification. A change to a public specification or a standards proposal, as recorded in its own repository.
  • Vendor claim. A figure a vendor states without a published method. Quoted as the vendor’s claim, never as a fact.
  • Vendor measurement. A vendor’s own measurement, published with enough method to judge it.
  • Preprint. A paper posted without peer review. Its numbers are the authors’ claims, tied to the models and dates they tested.
  • Independent measurement. A public report with steps to reproduce, by someone with no stake in the product.

Edition 1: 23 September to 7 October 2026

Published . Sources checked . 32 entries.

Two themes ran through this fortnight. Large platforms moved their whole API surface behind search-first discovery, and several vendors moved human confirmation and credential handling into the protocol itself. Preprints kept finding the same thing from different directions: agents report success that did not happen, and small, checkable details in a tool contract change what they do.

Large APIs and tool catalogues

  1. Cloudflare deprecates six product MCP servers in favour of its Code Mode server

    · Vendor documentation

    Pull requests merged on 6 October deprecate Cloudflare's Logpush, AI Gateway, Workers Builds, DNS Analytics, Workers Bindings and Observability MCP servers. The repository README asks users to move to the Cloudflare API MCP server, which reaches the whole API through Code Mode: a handful of tools instead of one per endpoint. The Audit Logs server was deprecated on 24 September. Existing tools keep working for now.

    Sources: cloudflare/mcp-server-cloudflare README; Pull request #516 (Observability).

    Tags: code-mode, dynamic-tools, context-budget. Affects: Agent tool catalogues need a pruning strategy; FIND-06: There are few enough tools to choose from; AXC-D017: Context footprint estimate.

  2. GitHub MCP Server 2.0 adds output schemas, sent only to clients that can read them

    · Vendor documentation

    Version 2.0.0 of GitHub's MCP server adds typed output schemas to its tools. The release notes say the schemas are only advertised to clients that declare support for the MCP specification of 2026-07-28 or later, so older clients are not handed a contract they cannot read.

    Source: github-mcp-server v2.0.0 release notes.

    Tags: schema-enum, code-mode. Affects: Write tool descriptions an agent can act on; READ-04: Each tool is fully described; AXC-D019: Example does not match the schema.

  3. Preprint maps 1.3 million MCP tool specifications

    · Preprint

    Tang, Chen, Yue and Xiao collected 124,267 unique MCP servers from 17 marketplaces and extracted 1,328,233 tool specifications. Their claim is that 98.5% of tools have at least one functional alternative, and that showing candidates through a capability taxonomy instead of a flat list raised task completion for all four models tested, by up to 12 points on crowded candidate sets.

    Source: arXiv 2610.05319.

    Tags: discovery, selection. Affects: Agent tool catalogues need a pruning strategy; FIND-06: There are few enough tools to choose from; AXC-D013: Near-duplicate descriptions.

  4. Cloudflare launches cf, a CLI for its whole API, built with agents in mind

    · Vendor documentation · Includes a vendor claim

    Cloudflare released cf as an open beta. The launch post says it covers the Cloudflare API surface of over 3,000 operations and makes JSON the default interface. Cloudflare's agent guide suggests a one-line instruction for a user-level AGENTS.md or CLAUDE.md that tells agents to use cf.

    Vendor claim: The launch post says agents accounted for 48% of use of Cloudflare's older Wrangler CLI in the week before launch, up from about a quarter in March 2026. This is Cloudflare's own figure, without a published method.

    Sources: Cloudflare blog, cf launch; Use cf with coding agents (Cloudflare docs).

    Tags: discovery, non-interactive, agents-md. Affects: Agent tool catalogues need a pruning strategy; ACT-07: There is a direct route, not only clicks; AXC-C006: No machine-readable output flag; EV-0033: Vendor claim: agents were 48% of Wrangler CLI use.

  5. Cloudflare open-sources Forge, the pipeline that generates cf

    · Vendor documentation

    Forge reads annotated OpenAPI descriptions and already generates what the cf CLI needs. Cloudflare says it will also power its API documentation and SDKs, and lists other input formats and SDK languages as planned, not shipped.

    Source: Cloudflare blog, Introducing Forge.

    Tags: description, schema-enum, drift. Affects: Write tool descriptions an agent can act on; READ-04: Each tool is fully described; AXC-D007: Closed vocabulary hinted in prose, not declared as an enum.

Approval, confirmation and access

  1. eve binds an approval to the person who asked for the call

    · Vendor documentation

    In eve 0.72.0, Vercel's open agent framework, an approval without a response policy can now be approved or cancelled only by the person whose turn requested the call. A tool that needs a different approver has to say so in its approval policy.

    Source: eve 0.72.0 release notes.

    Tags: approval. Affects: Human approval workflows for AI agents; ACT-02: Your system enforces approval; HANDBACK-04: Approvals are few and show the real action.

  2. Supabase scoped access tokens reach general availability, and a refusal names what is missing

    · Vendor documentation

    Supabase's scoped personal access tokens are now available to everyone. A token's access cannot be changed after it is created, a 403 response lists the permissions the token lacks in a missing_permissions array, and the dashboard shows which MCP tools a token can call. Existing account-level tokens keep working.

    Source: Supabase changelog, scoped PATs GA.

    Tags: auth-scopes, error-recovery. Affects: Agent error recovery: design the path back; ACT-05: Agent access can be narrow; RECOVER-02: Errors say how to succeed; AXC-C004: Error does not name a next step.

  3. Preprint: how practitioners decide what an agent may do

    · Preprint

    Salerno and colleagues interviewed 18 practitioners and surveyed 115. Their recommendations: make consequential actions easier to review, separate what an agent is allowed to do from what the user intended, make reversibility clearer, do not treat repeated approvals as stable preferences, and tell apart rejecting one action from rejecting a whole approach.

    Source: arXiv 2610.06047.

    Tags: approval, confirmation. Affects: Human approval workflows for AI agents; HANDBACK-04: Approvals are few and show the real action.

  4. Supabase MCP asks for confirmation before cost and destructive SQL

    · Vendor documentation

    Announced at Supabase Select, the Supabase MCP server uses MCP elicitations to show a confirmation form before create_project or create_branch incurs a cost, and before destructive SQL through execute_sql or apply_migration outside read-only mode. Supabase's own docs say to treat elicitations as a guardrail, not a guarantee: they can be switched off per tool, and some clients answer them automatically.

    Sources: Supabase MCP docs, Elicitations; Supabase Select recap.

    Tags: confirmation, approval. Affects: Human approval workflows for AI agents; ACT-01: Effects are stated before the step; ACT-02: Your system enforces approval; AXC-D009: Side-effect annotation contradicts the description.

  5. Supabase keeps secret values out of the conversation

    · Vendor documentation

    For Edge Function secrets, the Supabase MCP server asks the person to type the value in the Supabase Dashboard through a URL-mode elicitation. The docs say secret values belong in the Dashboard, 'never in chat, model input, or MCP tool arguments'.

    Source: Supabase MCP docs, Elicitations.

    Tags: secrets, handoff. Affects: Human approval workflows for AI agents; ACT-05: Agent access can be narrow; EV-0038: Dry-run output printed secrets until a fix redacted them.

  6. Neon makes a query-plan tool safe by default and corrects its annotations

    · Vendor documentation

    The explain_sql_statement tool in Neon's MCP server defaulted to EXPLAIN ANALYZE, which PostgreSQL executes. A merged change makes plain EXPLAIN the default and corrects the tool's annotations, which had described it as read-only, non-destructive and idempotent.

    Source: neondatabase/mcp-server-neon pull request #367.

    Tags: description, confirmation. Affects: Write tool descriptions an agent can act on; ACT-01: Effects are stated before the step; AXC-D009: Side-effect annotation contradicts the description.

  7. Vercel Agent keeps registry credentials outside its sandbox

    · Vendor documentation

    Vercel Agent sessions can now install private npm packages using team-shared variables, and the changelog says 'credential values stay outside the sandbox, so the agent cannot read them'. In the same fortnight Vercel Sandbox gained persistent Drives (public beta, 23 September) and private networking through Secure Compute for Enterprise teams (30 September).

    Sources: Vercel changelog, private packages; Vercel changelog, Sandbox Drives; Vercel changelog, Sandbox Secure Compute.

    Tags: secrets, auth-scopes. Affects: Human approval workflows for AI agents; ACT-05: Agent access can be narrow.

  8. Vercel Connect opens to third-party services and lists what they must support

    · Vendor documentation

    Any service can now submit itself to the Vercel Connect directory, which brokers short-lived OAuth access for agents; Vercel reviews each submission. The provider documentation sorts OAuth features into required (authorisation server metadata through RFC 8414 or OpenID Connect discovery, supported grant types, expires_in), recommended (dynamic client registration, PKCE, refresh tokens) and optional (for example RFC 9728 protected resource metadata).

    Sources: Vercel changelog, Connect service submissions; Vercel docs, Connect providers.

    Tags: auth-scopes, secrets, discovery. Affects: Human approval workflows for AI agents; ACT-05: Agent access can be narrow.

  9. Preprint: approval records miss what a command goes on to do

    · Preprint

    Zhang and colleagues compared coding-agent approval records with execution traces. Across 111 pairs, effects left out of the record fell from 40 with explicit fields to 17 with command semantics and 13 with decision-time metadata (the paper's claim).

    Source: arXiv 2609.28586.

    Tags: approval, confirmation. Affects: Human approval workflows for AI agents; HANDBACK-04: Approvals are few and show the real action; EV-0016: Approval records omit the effects a command goes on to trigger.

Exit codes, errors and false success

  1. Preprint: agent benchmarks that trust their own simulated tools

    · Preprint

    Bellibatlu, Wang and Zhang treated the advertised behaviour of benchmark tools as a contract and checked the code against it. Across 34 mutating tools in four benchmarks they confirmed seven tool defects and one evaluator property, including a clinical benchmark whose tool reports writes that never happen (the paper's claim).

    Source: arXiv 2609.37315.

    Tags: evaluation, false-success. Affects: Evaluate the decisions in an agent workflow; Agent evaluation loop: run, judge, triage, fix, repeat.

  2. Cloudflare documents that a refused deletion exits with status 0

    · Vendor documentation

    Cloudflare's agent guide for cf says that in a non-interactive session a destructive command without --force prints 'Aborted.' to standard error and exits with status 0, and warns that a successful exit does not mean the resource was deleted. An open issue, cloudflare/cf#94, asks for a non-zero exit.

    Sources: Use cf with coding agents (Cloudflare docs); cloudflare/cf issue #94.

    Tags: exit-codes, false-success, non-interactive, confirmation. Affects: Agent error recovery: design the path back; RECOVER-03: The agent can check what happened; AXC-C002: Error exits with status 0; EV-0036: A CLI exits 0 when a destructive command is refused.

  3. Preprint: error messages written for developers hurt capable agents most

    · Preprint

    Xu and Wu tested five OpenAI models. An expired-credential error that named a terminal command left 45% of tasks recovered; naming the server's login tool raised that to 84%. On rate limits, naming the call to repeat raised recovery from 6% to 88% (the paper's claim).

    Source: arXiv 2609.35381.

    Tags: error-recovery, description. Affects: Agent error recovery: design the path back; RECOVER-02: Errors say how to succeed; AXC-C004: Error does not name a next step; EV-0003: Error text that names the next tool lifts recovery.

  4. Preprint: an evidence contract cuts false reports of success

    · Preprint

    Zhu and colleagues report that after a required tool failed, six models falsely reported success in 22.8% of responses by default, 9.3% with a transparency instruction and 0.8% with a structured evidence contract (the paper's claim).

    Source: arXiv 2609.35732.

    Tags: false-success, evaluation. Affects: Record retries as product history; HANDBACK-01: The person has a record; EV-0006: An evidence contract cuts false-success reports after tool failures.

  5. Preprint: idempotency keys cut duplicate writes by agents

    · Preprint

    In LIMBO, Li reports that offering an idempotency key on every write cut duplicate side effects from 28% to 4% of episodes, because agents use keys when they exist, and that agents reported success in 90% of the episodes in which they had duplicated an effect (the paper's claim).

    Source: arXiv 2609.29095.

    Tags: idempotency, false-success. Affects: Record retries as product history; ACT-03: Writes are safe to retry; EV-0005: Idempotency keys cut duplicate writes from 28% to 4%.

Instructions, skills and drift

  1. Vercel's agent plugin fixes instructions that drifted from reality

    · Vendor documentation

    Merged pull requests in vercel/vercel-plugin corrected agent instructions that no longer matched reality. Pull request #306 removed a tip recommending a package that is not published on npm and fixed AI SDK version claims. Pull request #307, merged on 6 October, stopped plugin instructions advertising interactive deployment cards that the production MCP server does not expose.

    Sources: vercel/vercel-plugin pull request #306; vercel/vercel-plugin pull request #307.

    Tags: drift, skills. Affects: Write agent-readable documentation; BOUND-03: Your signals agree; AXC-F003: Claimed npm package does not exist; AXC-F006: Claimed MCP tool is not served; EV-0040: Agent skills recommended a package that does not exist.

  2. Preprint: skill rules that name a command change what agents do

    · Preprint

    Wang and colleagues studied revisions to 3,159 skills. Adding a checkable rule raised the rate at which four coding agents took the required action by +0.23 on average, with most of the gain coming from rules that name a command or path the old skill did not mention (the paper's claim).

    Source: arXiv 2610.04832.

    Tags: skills, description. Affects: Write agent-readable documentation; READ-10: Repository instructions give commands and limits; EV-0018: Skill rules that name a command or path change what agents do.

  3. Vercel predicts skills will be judged by tests, not installs

    · Vendor claim

    Vercel published a report on its skills.sh registry and predicted that 'the measure of a skill will also shift from popularity to effectiveness', with tests and benchmarks to match.

    Vendor claim: Vercel says the registry reached one million skills and nearly 280 million installs in the seven months to August, and that 375 skills (0.04%) account for 62% of installs. Vercel notes that its install counters do not represent unique people.

    Source: Vercel blog, State of agent skills.

    Tags: skills, measurement. Affects: How do you measure agent experience?; AXC-D022: SKILL.md frontmatter missing or invalid; AXC-D024: Skill description over 1,024 characters.

Discovery, standards and starting without an account

  1. MCP Server Cards proposal merged as Final, with the card format still experimental

    · Specification

    SEP-2127, which defines Server Cards (static metadata a client can read before it connects, with <server URL>/server-card as the recommended location), was merged on 6 October with the status Final, as an optional extension. The SEP leaves the wire format to a separate extension repository, which still described itself as experimental, and its discovery document as a draft, when last updated.

    Source: modelcontextprotocol pull request #2127.

    Tags: server-cards, discovery. Affects: Agentic Resource Discovery, implemented; The state of agent-web standards; FIND-05: APIs and tools are described where agents look; AXC-F006: Claimed MCP tool is not served.

  2. MCP documentation adds a security guide for local servers

    · Specification

    A local server security guide was merged into the MCP documentation. It says 'the stdio transport is not a sandbox' and that a local MCP server is not a plugin running inside a sandbox the protocol provides.

    Source: modelcontextprotocol pull request #3072.

    Tags: prompt-injection, secrets. Affects: Human approval workflows for AI agents; ACT-05: Agent access can be narrow.

  3. Clerk lets an agent start an app before anyone signs up

    · Vendor documentation

    Run while signed out, Clerk's CLI creates an accountless development application that is not attached to any account yet, and it runs without prompting when it detects a non-interactive agent environment. The missing-key error names the exact command to run next. Claiming the application stays a step for the person.

    Source: Clerk blog, accountless setup.

    Tags: claim-later-onboarding, non-interactive, error-recovery. Affects: Agent error recovery: design the path back; RECOVER-02: Errors say how to succeed; AXC-C001: Command hangs without a terminal; AXC-C004: Error does not name a next step.

  4. Preprint: page checklists showed no citation effect within a domain

    · Preprint

    Moore and Dunne studied about two million citations from AI answer engines. FAQ blocks, structured data and Core Web Vitals showed positive effects in pooled data that reversed or fell to zero once each domain was compared with itself (the paper's claim, from an observational study of B2B software pages).

    Source: arXiv 2609.35077.

    Tags: discovery, measurement. Affects: AX, UX, DX, and GEO: where each fits; Does llms.txt actually work?; EV-0013: FAQ blocks and structured data showed no citation effect within a domain.

Definitions and measurement

  1. Microsoft publishes 'What is Agent Experience (AX)?'

    · Vendor documentation

    Microsoft's developer blog defines AX as 'the experience AI agents have when discovering, choosing, and using your technology', credits the term to Mathias Biilmann of Netlify (January 2025), and splits measurement into propensity (does the agent find your technology and choose it?) and efficacy (does it use it correctly?). Its argument: best practices are hypotheses until you measure them.

    Source: Microsoft for Developers, What is Agent Experience (AX)?.

    Tags: measurement, propensity, evaluation. Affects: What is Agent Experience?; How do you measure agent experience?; EV-0029: A warning that names the failing plan redirected agents; a tip did not.

  2. Preprint adds a preregistered test: praise and order move tool choice

    · Preprint

    A revised version of Wang and Zhang's paper adds a preregistered study with two small OpenAI models. Stacked praise in a tool listing raised its pick rate by about 43 percentage points, and with identical listings the first-listed tool was picked about 72 points more often (the paper's claim).

    Source: arXiv 2605.23916 (version 2).

    Tags: selection, description. Affects: Write tool descriptions an agent can act on; FIND-06: There are few enough tools to choose from; AXC-D020: Promotional language in a description; EV-0020: Praise and list order move tool selection.

What goes in, and what does not

An entry needs a public, dated source in the edition's window and a clear link to how agents find, choose, use or recover with a product. Funding news, opinion posts and announcements of features that have not shipped are left out unless they change one of those things. Missed something, or found an error? Send the source through the contact page. 32 entries so far.