<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>What changed in AX | Agent Experience</title><description>A roughly fortnightly digest of public changes that affect how AI agents find, choose, use and recover with products: launches, documentation, specifications and preprints, each with its source and evidence class.</description><link>https://agentexperience.tech/</link><language>en-gb</language><item><title>eve binds an approval to the person who asked for the call</title><link>https://agentexperience.tech/changes/#change-eve-requester-approval</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-eve-requester-approval</guid><description>In eve 0.72.0, Vercel&apos;s open agent framework, an approval without a response policy can now be approved or cancelled only by the person whose turn requested the call. A tool that needs a different approver has to say so in its approval policy. (Vendor documentation. Source: https://github.com/vercel/eve/releases/tag/eve%400.72.0)</description><pubDate>Wed, 07 Oct 2026 00:00:00 GMT</pubDate><category>approval</category></item><item><title>Cloudflare deprecates six product MCP servers in favour of its Code Mode server</title><link>https://agentexperience.tech/changes/#change-cloudflare-mcp-code-mode</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-cloudflare-mcp-code-mode</guid><description>Pull requests merged on 6 October deprecate Cloudflare&apos;s Logpush, AI Gateway, Workers Builds, DNS Analytics, Workers Bindings and Observability MCP servers. The repository README asks users to move to the Cloudflare API MCP server, which reaches the whole API through Code Mode: a handful of tools instead of one per endpoint. The Audit Logs server was deprecated on 24 September. Existing tools keep working for now. (Vendor documentation. Source: https://github.com/cloudflare/mcp-server-cloudflare)</description><pubDate>Tue, 06 Oct 2026 00:00:00 GMT</pubDate><category>code-mode</category><category>dynamic-tools</category><category>context-budget</category></item><item><title>GitHub MCP Server 2.0 adds output schemas, sent only to clients that can read them</title><link>https://agentexperience.tech/changes/#change-github-mcp-output-schemas</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-github-mcp-output-schemas</guid><description>Version 2.0.0 of GitHub&apos;s MCP server adds typed output schemas to its tools. The release notes say the schemas are only advertised to clients that declare support for the MCP specification of 2026-07-28 or later, so older clients are not handed a contract they cannot read. (Vendor documentation. Source: https://github.com/github/github-mcp-server/releases/tag/v2.0.0)</description><pubDate>Tue, 06 Oct 2026 00:00:00 GMT</pubDate><category>schema-enum</category><category>code-mode</category></item><item><title>MCP Server Cards proposal merged as Final, with the card format still experimental</title><link>https://agentexperience.tech/changes/#change-mcp-server-cards</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-mcp-server-cards</guid><description>SEP-2127, which defines Server Cards (static metadata a client can read before it connects, with `&lt;server URL&gt;/server-card` as the recommended location), was merged on 6 October with the status Final, as an optional extension. The SEP leaves the wire format to a separate extension repository, which still described itself as experimental, and its discovery document as a draft, when last updated. (Specification. Source: https://github.com/modelcontextprotocol/modelcontextprotocol/pull/2127)</description><pubDate>Tue, 06 Oct 2026 00:00:00 GMT</pubDate><category>server-cards</category><category>discovery</category></item><item><title>Microsoft publishes &apos;What is Agent Experience (AX)?&apos;</title><link>https://agentexperience.tech/changes/#change-microsoft-what-is-ax</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-microsoft-what-is-ax</guid><description>Microsoft&apos;s developer blog defines AX as &apos;the experience AI agents have when discovering, choosing, and using your technology&apos;, credits the term to Mathias Biilmann of Netlify (January 2025), and splits measurement into propensity (does the agent find your technology and choose it?) and efficacy (does it use it correctly?). Its argument: best practices are hypotheses until you measure them. (Vendor documentation. Source: https://developer.microsoft.com/blog/what-is-agent-experience-ax/)</description><pubDate>Tue, 06 Oct 2026 00:00:00 GMT</pubDate><category>measurement</category><category>propensity</category><category>evaluation</category></item><item><title>Supabase scoped access tokens reach general availability, and a refusal names what is missing</title><link>https://agentexperience.tech/changes/#change-supabase-scoped-tokens</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-supabase-scoped-tokens</guid><description>Supabase&apos;s scoped personal access tokens are now available to everyone. A token&apos;s access cannot be changed after it is created, a 403 response lists the permissions the token lacks in a `missing_permissions` array, and the dashboard shows which MCP tools a token can call. Existing account-level tokens keep working. (Vendor documentation. Source: https://supabase.com/changelog/scoped-personal-access-tokens-ga)</description><pubDate>Tue, 06 Oct 2026 00:00:00 GMT</pubDate><category>auth-scopes</category><category>error-recovery</category></item><item><title>MCP documentation adds a security guide for local servers</title><link>https://agentexperience.tech/changes/#change-mcp-local-server-security</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-mcp-local-server-security</guid><description>A local server security guide was merged into the MCP documentation. It says &apos;the stdio transport is not a sandbox&apos; and that a local MCP server is not a plugin running inside a sandbox the protocol provides. (Specification. Source: https://github.com/modelcontextprotocol/modelcontextprotocol/pull/3072)</description><pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate><category>prompt-injection</category><category>secrets</category></item><item><title>Preprint: how practitioners decide what an agent may do</title><link>https://agentexperience.tech/changes/#change-permission-decisions</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-permission-decisions</guid><description>Salerno and colleagues interviewed 18 practitioners and surveyed 115. Their recommendations: make consequential actions easier to review, separate what an agent is allowed to do from what the user intended, make reversibility clearer, do not treat repeated approvals as stable preferences, and tell apart rejecting one action from rejecting a whole approach. (Preprint. Source: https://arxiv.org/abs/2610.06047)</description><pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate><category>approval</category><category>confirmation</category></item><item><title>Vercel&apos;s agent plugin fixes instructions that drifted from reality</title><link>https://agentexperience.tech/changes/#change-vercel-plugin-drift</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-vercel-plugin-drift</guid><description>Merged pull requests in vercel/vercel-plugin corrected agent instructions that no longer matched reality. Pull request #306 removed a tip recommending a package that is not published on npm and fixed AI SDK version claims. Pull request #307, merged on 6 October, stopped plugin instructions advertising interactive deployment cards that the production MCP server does not expose. (Vendor documentation. Source: https://github.com/vercel/vercel-plugin/pull/306)</description><pubDate>Mon, 05 Oct 2026 00:00:00 GMT</pubDate><category>drift</category><category>skills</category></item><item><title>Preprint maps 1.3 million MCP tool specifications</title><link>https://agentexperience.tech/changes/#change-mcpacific</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-mcpacific</guid><description>Tang, Chen, Yue and Xiao collected 124,267 unique MCP servers from 17 marketplaces and extracted 1,328,233 tool specifications. Their claim is that 98.5% of tools have at least one functional alternative, and that showing candidates through a capability taxonomy instead of a flat list raised task completion for all four models tested, by up to 12 points on crowded candidate sets. (Preprint. Source: https://arxiv.org/abs/2610.05319)</description><pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate><category>discovery</category><category>selection</category></item><item><title>Preprint: skill rules that name a command change what agents do</title><link>https://agentexperience.tech/changes/#change-skill-evolution</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-skill-evolution</guid><description>Wang and colleagues studied revisions to 3,159 skills. Adding a checkable rule raised the rate at which four coding agents took the required action by +0.23 on average, with most of the gain coming from rules that name a command or path the old skill did not mention (the paper&apos;s claim). (Preprint. Source: https://arxiv.org/abs/2610.04832)</description><pubDate>Sun, 04 Oct 2026 00:00:00 GMT</pubDate><category>skills</category><category>description</category></item><item><title>Supabase keeps secret values out of the conversation</title><link>https://agentexperience.tech/changes/#change-supabase-secrets</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-supabase-secrets</guid><description>For Edge Function secrets, the Supabase MCP server asks the person to type the value in the Supabase Dashboard through a URL-mode elicitation. The docs say secret values belong in the Dashboard, &apos;never in chat, model input, or MCP tool arguments&apos;. (Vendor documentation. Source: https://supabase.com/docs/guides/ai-tools/mcp#elicitations)</description><pubDate>Fri, 02 Oct 2026 00:00:00 GMT</pubDate><category>secrets</category><category>handoff</category></item><item><title>Supabase MCP asks for confirmation before cost and destructive SQL</title><link>https://agentexperience.tech/changes/#change-supabase-elicitations</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-supabase-elicitations</guid><description>Announced at Supabase Select, the Supabase MCP server uses MCP elicitations to show a confirmation form before `create_project` or `create_branch` incurs a cost, and before destructive SQL through `execute_sql` or `apply_migration` outside read-only mode. Supabase&apos;s own docs say to treat elicitations as a guardrail, not a guarantee: they can be switched off per tool, and some clients answer them automatically. (Vendor documentation. Source: https://supabase.com/docs/guides/ai-tools/mcp#elicitations)</description><pubDate>Fri, 02 Oct 2026 00:00:00 GMT</pubDate><category>confirmation</category><category>approval</category></item><item><title>Clerk lets an agent start an app before anyone signs up</title><link>https://agentexperience.tech/changes/#change-clerk-accountless</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-clerk-accountless</guid><description>Run while signed out, Clerk&apos;s CLI creates an accountless development application that is not attached to any account yet, and it runs without prompting when it detects a non-interactive agent environment. The missing-key error names the exact command to run next. Claiming the application stays a step for the person. (Vendor documentation. Source: https://clerk.com/blog/accountless-setup-clerk-init-nextjs)</description><pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate><category>claim-later-onboarding</category><category>non-interactive</category><category>error-recovery</category></item><item><title>Neon makes a query-plan tool safe by default and corrects its annotations</title><link>https://agentexperience.tech/changes/#change-neon-explain-default</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-neon-explain-default</guid><description>The `explain_sql_statement` tool in Neon&apos;s MCP server defaulted to `EXPLAIN ANALYZE`, which PostgreSQL executes. A merged change makes plain `EXPLAIN` the default and corrects the tool&apos;s annotations, which had described it as read-only, non-destructive and idempotent. (Vendor documentation. Source: https://github.com/neondatabase/mcp-server-neon/pull/367)</description><pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate><category>description</category><category>confirmation</category></item><item><title>Preprint adds a preregistered test: praise and order move tool choice</title><link>https://agentexperience.tech/changes/#change-tool-registry-rhetoric</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-tool-registry-rhetoric</guid><description>A revised version of Wang and Zhang&apos;s paper adds a preregistered study with two small OpenAI models. Stacked praise in a tool listing raised its pick rate by about 43 percentage points, and with identical listings the first-listed tool was picked about 72 points more often (the paper&apos;s claim). (Preprint. Source: https://arxiv.org/abs/2605.23916)</description><pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate><category>selection</category><category>description</category></item><item><title>Vercel Agent keeps registry credentials outside its sandbox</title><link>https://agentexperience.tech/changes/#change-vercel-agent-credentials</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-vercel-agent-credentials</guid><description>Vercel Agent sessions can now install private npm packages using team-shared variables, and the changelog says &apos;credential values stay outside the sandbox, so the agent cannot read them&apos;. In the same fortnight Vercel Sandbox gained persistent Drives (public beta, 23 September) and private networking through Secure Compute for Enterprise teams (30 September). (Vendor documentation. Source: https://vercel.com/changelog/vercel-agent-now-installs-private-packages-from-npm-and-custom-registries)</description><pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate><category>secrets</category><category>auth-scopes</category></item><item><title>Cloudflare documents that a refused deletion exits with status 0</title><link>https://agentexperience.tech/changes/#change-cf-aborted-exit-zero</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-cf-aborted-exit-zero</guid><description>Cloudflare&apos;s agent guide for `cf` says that in a non-interactive session a destructive command without `--force` prints &apos;Aborted.&apos; to standard error and exits with status 0, and warns that a successful exit does not mean the resource was deleted. An open issue, cloudflare/cf#94, asks for a non-zero exit. (Vendor documentation. Source: https://developers.cloudflare.com/cf/agents/)</description><pubDate>Tue, 29 Sep 2026 00:00:00 GMT</pubDate><category>exit-codes</category><category>false-success</category><category>non-interactive</category><category>confirmation</category></item><item><title>Preprint: agent benchmarks that trust their own simulated tools</title><link>https://agentexperience.tech/changes/#change-benchmark-contract-audit</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-benchmark-contract-audit</guid><description>Bellibatlu, Wang and Zhang treated the advertised behaviour of benchmark tools as a contract and checked the code against it. Across 34 mutating tools in four benchmarks they confirmed seven tool defects and one evaluator property, including a clinical benchmark whose tool reports writes that never happen (the paper&apos;s claim). (Preprint. Source: https://arxiv.org/abs/2609.37315)</description><pubDate>Tue, 29 Sep 2026 00:00:00 GMT</pubDate><category>evaluation</category><category>false-success</category></item><item><title>Stripe Link for personal agents: the person approves, then the agent pays</title><link>https://agentexperience.tech/changes/#change-stripe-link-agents</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-stripe-link-agents</guid><description>Stripe&apos;s Link wallet for personal agents, for consumers in the US and Canada, has the agent create a spend request that the customer approves on the Link website or mobile app. The launch post shows a failure returned with a structured next step the agent can act on. Vendor claim: Stripe says agentic purchases made with Link rose 38 times over the past month. Vendor claim, with no published method. (Vendor documentation. Source: https://stripe.com/blog/helping-personal-agents-shop-more-intelligently-and-reliably-with-link)</description><pubDate>Tue, 29 Sep 2026 00:00:00 GMT</pubDate><category>approval</category><category>error-recovery</category></item><item><title>Vercel Connect opens to third-party services and lists what they must support</title><link>https://agentexperience.tech/changes/#change-vercel-connect-providers</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-vercel-connect-providers</guid><description>Any service can now submit itself to the Vercel Connect directory, which brokers short-lived OAuth access for agents; Vercel reviews each submission. The provider documentation sorts OAuth features into required (authorisation server metadata through RFC 8414 or OpenID Connect discovery, supported grant types, `expires_in`), recommended (dynamic client registration, PKCE, refresh tokens) and optional (for example RFC 9728 protected resource metadata). (Vendor documentation. Source: https://vercel.com/changelog/vercel-connect-service-submissions)</description><pubDate>Tue, 29 Sep 2026 00:00:00 GMT</pubDate><category>auth-scopes</category><category>secrets</category><category>discovery</category></item><item><title>Cloudflare launches cf, a CLI for its whole API, built with agents in mind</title><link>https://agentexperience.tech/changes/#change-cloudflare-cf-cli</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-cloudflare-cf-cli</guid><description>Cloudflare released `cf` as an open beta. The launch post says it covers the Cloudflare API surface of over 3,000 operations and makes JSON the default interface. Cloudflare&apos;s agent guide suggests a one-line instruction for a user-level AGENTS.md or CLAUDE.md that tells agents to use `cf`. Vendor claim: The launch post says agents accounted for 48% of use of Cloudflare&apos;s older Wrangler CLI in the week before launch, up from about a quarter in March 2026. This is Cloudflare&apos;s own figure, without a published method. (Vendor documentation. Source: https://blog.cloudflare.com/cloudflare-cf-cli-launch/)</description><pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate><category>discovery</category><category>non-interactive</category><category>agents-md</category></item><item><title>Cloudflare open-sources Forge, the pipeline that generates cf</title><link>https://agentexperience.tech/changes/#change-cloudflare-forge</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-cloudflare-forge</guid><description>Forge reads annotated OpenAPI descriptions and already generates what the `cf` CLI needs. Cloudflare says it will also power its API documentation and SDKs, and lists other input formats and SDK languages as planned, not shipped. (Vendor documentation. Source: https://blog.cloudflare.com/forge-open-source-generation-pipeline/)</description><pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate><category>description</category><category>schema-enum</category><category>drift</category></item><item><title>Preprint: an evidence contract cuts false reports of success</title><link>https://agentexperience.tech/changes/#change-failure-transparent-agents</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-failure-transparent-agents</guid><description>Zhu and colleagues report that after a required tool failed, six models falsely reported success in 22.8% of responses by default, 9.3% with a transparency instruction and 0.8% with a structured evidence contract (the paper&apos;s claim). (Preprint. Source: https://arxiv.org/abs/2609.35732)</description><pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate><category>false-success</category><category>evaluation</category></item><item><title>Preprint: error messages written for developers hurt capable agents most</title><link>https://agentexperience.tech/changes/#change-developer-error-messages</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-developer-error-messages</guid><description>Xu and Wu tested five OpenAI models. An expired-credential error that named a terminal command left 45% of tasks recovered; naming the server&apos;s login tool raised that to 84%. On rate limits, naming the call to repeat raised recovery from 6% to 88% (the paper&apos;s claim). (Preprint. Source: https://arxiv.org/abs/2609.35381)</description><pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate><category>error-recovery</category><category>description</category></item><item><title>Preprint: page checklists showed no citation effect within a domain</title><link>https://agentexperience.tech/changes/#change-citation-study</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-citation-study</guid><description>Moore and Dunne studied about two million citations from AI answer engines. FAQ blocks, structured data and Core Web Vitals showed positive effects in pooled data that reversed or fell to zero once each domain was compared with itself (the paper&apos;s claim, from an observational study of B2B software pages). (Preprint. Source: https://arxiv.org/abs/2609.35077)</description><pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate><category>discovery</category><category>measurement</category></item><item><title>Vercel opens domain search without sign-in</title><link>https://agentexperience.tech/changes/#change-vercel-domain-search</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-vercel-domain-search</guid><description>Domain names, availability and pricing can now be checked through Vercel&apos;s CLI and API without signing in or an access token. Buying or managing a domain still needs authentication. (Vendor documentation. Source: https://vercel.com/changelog/search-domains-without-authentication)</description><pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate><category>claim-later-onboarding</category><category>auth-scopes</category></item><item><title>Preprint: approvals that outlive their task raise attack success</title><link>https://agentexperience.tech/changes/#change-consent-outlives-context</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-consent-outlives-context</guid><description>Zhang and colleagues report that approvals persisted beyond the context that justified them raised prompt-injection attack success by up to 35.1 percentage points on 508 AgentDojo cases, and by 24.9 points in live tests on three production coding agents (the paper&apos;s claim). (Preprint. Source: https://arxiv.org/abs/2609.33910)</description><pubDate>Sun, 27 Sep 2026 00:00:00 GMT</pubDate><category>approval</category><category>prompt-injection</category></item><item><title>Vercel predicts skills will be judged by tests, not installs</title><link>https://agentexperience.tech/changes/#change-state-of-agent-skills</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-state-of-agent-skills</guid><description>Vercel published a report on its skills.sh registry and predicted that &apos;the measure of a skill will also shift from popularity to effectiveness&apos;, with tests and benchmarks to match. Vendor claim: Vercel says the registry reached one million skills and nearly 280 million installs in the seven months to August, and that 375 skills (0.04%) account for 62% of installs. Vercel notes that its install counters do not represent unique people. (Vendor claim. Source: https://vercel.com/blog/state-of-agent-skills)</description><pubDate>Fri, 25 Sep 2026 00:00:00 GMT</pubDate><category>skills</category><category>measurement</category></item><item><title>Preprint: idempotency keys cut duplicate writes by agents</title><link>https://agentexperience.tech/changes/#change-limbo-idempotency</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-limbo-idempotency</guid><description>In LIMBO, Li reports that offering an idempotency key on every write cut duplicate side effects from 28% to 4% of episodes, because agents use keys when they exist, and that agents reported success in 90% of the episodes in which they had duplicated an effect (the paper&apos;s claim). (Preprint. Source: https://arxiv.org/abs/2609.29095)</description><pubDate>Thu, 24 Sep 2026 00:00:00 GMT</pubDate><category>idempotency</category><category>false-success</category></item><item><title>Vercel Connect sends a consent problem to the person, not the model</title><link>https://agentexperience.tech/changes/#change-vercel-connect-consent</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-vercel-connect-consent</guid><description>Vercel Connect&apos;s support for TanStack AI checks consent before the model runs. The changelog explains why: otherwise &apos;a consent error raised within a tool call would reach the model as an error string rather than the user as a redirect&apos;. (Vendor documentation. Source: https://vercel.com/changelog/vercel-connect-tanstack-ai)</description><pubDate>Thu, 24 Sep 2026 00:00:00 GMT</pubDate><category>auth-scopes</category><category>handoff</category></item><item><title>Preprint: approval records miss what a command goes on to do</title><link>https://agentexperience.tech/changes/#change-approval-laundering</link><guid isPermaLink="true">https://agentexperience.tech/changes/#change-approval-laundering</guid><description>Zhang and colleagues compared coding-agent approval records with execution traces. Across 111 pairs, effects left out of the record fell from 40 with explicit fields to 17 with command semantics and 13 with decision-time metadata (the paper&apos;s claim). (Preprint. Source: https://arxiv.org/abs/2609.28586)</description><pubDate>Wed, 23 Sep 2026 00:00:00 GMT</pubDate><category>approval</category><category>confirmation</category></item></channel></rss>