Evidence review

Does llms.txt actually work?

Two years after the proposal, the public evidence points one way for search visibility and a different way for coding agents. The distinction is the whole story.

Short answer, September 2026. There is no public evidence that llms.txt improves how often a site is retrieved or cited by AI search systems, and Google states plainly that Google Search ignores the file. There is modest evidence that coding agents fetch it. Publish one because it costs an hour and forces you to decide what your twenty most important pages are—not because you expect to be found more often. Anyone selling it as an AI-visibility tactic is ahead of the evidence.

What follows is a review of what the public studies actually measured, where they disagree, and which claims survive contact with their own methodology.

What llms.txt is, precisely

llms.txt is a Markdown file at a site’s root that gives an LLM a short description of the site and a curated set of links to its most useful documentation. Jeremy Howard proposed it on 3 September 2024; the current v2 page at llmstxt.org shows a last-modified date of 10 August 2026. The stated problem is real: web pages are built for people, wrapped in navigation and advertising, and expensive to extract from. The proposed fix is a hand-curated index in a format that costs few tokens to read.

Two things are worth fixing in place before looking at any data. First, it is a community proposal, not a standards-track document—there is no working group, no registry, and no conformance requirement. Second, it is a hint about where to look, not an instruction about what to do. It grants nothing, blocks nothing, and asserts nothing about permissions. That distinction gets lost regularly in marketing copy that positions llms.txt as the “robots.txt of AI”.

Adoption: the numbers disagree by a factor of five

Everyone cites an adoption figure. Almost nobody notes that the published figures are mutually incompatible.

Study Published Sample Adoption reported
Casey Burridge, HTTP Archive/CrUX analysis 2026-06-20 Millions of CrUX-origin sites, 12 months of crawls 6.28% of top 1,000; 5.61% of top 10,000; 5.07% of top 1M
SE Ranking reported 2025-11-20 ~300,000 domains 10.13%
Ahrefs 2026-06-15 137,210 domains with traffic in May 2026 28% (38,360 domains)
Originality.AI tracker 2026-07-03 3M+ site tracker 36,120 instances in May 2026, up from 4,088 in June 2025

A five-fold spread is not noise; it is a denominator problem. Ahrefs measured domains in its own Web Analytics product—a population of sites that installed an SEO analytics tool, and therefore a population unusually likely to have read an SEO blog post about llms.txt. Burridge measured the HTTP Archive’s CrUX-based crawl, which is closer to “sites real Chrome users visit”. Those are different universes, and the honest reading is that the low-single-digit-to-low-double-digit range is the general-web estimate and the 28% figure describes SEO-tool customers.

The growth trend is more consistent than the level. Burridge’s series moves from 1.04% of the top 10,000 in July 2025 to 5.61% in June 2026; Originality’s tracker shows roughly an 8.8x rise in absolute instances over a similar window. But a large share of that growth is not decisions—it is platform defaults. Burridge found 78.1% of Shopify sites in the top 10,000 carried the file following an automatic platform rollout, against 8.7% of WordPress sites. When one hosting platform can move the global adoption curve in a quarter, adoption stops being a proxy for belief.

Consumption: the file is published far more than it is read

Adoption is the easy number. The harder and more decisive question is whether anything fetches the file.

The best public dataset is Ahrefs’ server-log study of 137,210 domains, published 15 June 2026. Its headline: 97% of llms.txt files received zero requests in May 2026. Of the 3% that received any traffic, 96% of requests came from bots—and roughly 77% of those bots were not AI tools at all, with SEO audit crawlers among the largest single categories. Named AI user agents were a minority of a minority: GPTBot at 4.51% of AI requests, ClaudeBot at 0.80%, DeepSeek at 0.02%.

There is a widely circulated companion statistic: “over 500 million AI bot visits monitored across a 90-day window; only 408 targeted llms.txt directly.” It appears in Limy’s 2026 guide, attributed to the vendor’s own internal monitoring, with no published methodology, no sampling description, no date range, and no external link. It is directionally consistent with the Ahrefs log data, which is presumably why it spreads. But it is a vendor-internal figure that cannot be checked, and it should be cited as such or not cited at all. Treat it as an anecdote that agrees with better evidence, not as evidence.

One result in the Ahrefs data cuts against the general picture and deserves more attention than it got: among AI user agents, Claude Code—a coding agent—out-fetched the AI retrieval bots. That is a small but genuine signal, and it points at exactly the use case the file was designed for.

Effect: the one correlational study found nothing

Fetching is not the same as benefit. The most directly relevant attempt to measure effect is SE Ranking’s analysis of roughly 300,000 domains, reported by Search Engine Journal on 20 November 2025. It found no significant correlation between the presence of llms.txt and how often a domain was cited in responses from major LLMs. In their gradient-boosted model of citation behaviour, removing the llms.txt variable improved predictive accuracy—the feature was contributing noise rather than signal.

The methodological caveats are the obvious ones and they cut both ways: this is observational, domain-level, correlational, and cannot separate “the file does nothing” from “the file does something too small to detect against the enormous confound of site authority”. No published experiment—randomised, controlled, with a holdout—has been done at scale by anyone with the log access to do it properly. So the accurate statement is not “llms.txt has been proven useless for visibility”. It is: after two years, nobody has produced evidence that it helps, and the one large correlational attempt found nothing. In a field this eager for a tactic, absence of evidence after this much motivated searching is itself informative.

What the vendors actually say

This is where most write-ups get sloppy, because two arms of Google say things that sound contradictory and are not.

Google Search says it ignores the file. Google published its first official generative-AI optimisation guide in May 2026 and added a clarifying subsection on llms.txt the following month. The current text of Optimizing your website for generative AI features on Google Search (last updated 2026-07-10) states: “You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn’t use them.” And, more pointedly: “It’s completely fine if you decide to create and maintain LLMS.txt files (or other similar files) for other services or systems that use these files. Doing so will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them.” That is about as unambiguous as platform documentation gets.

Chrome ships a Lighthouse audit for it. Chrome’s agentic-browsing audit set includes an llms.txt check (docs last updated 2026-05-05). The audit flags server errors when retrieving the file and is marked Not Applicable on a 404, because publishing one is “currently optional”. Its stated rationale is efficiency, not ranking: “Without this file, agents may spend more time crawling the site to understand its high-level structure and primary content.”

These are not in conflict. Search is answering “does this change what we retrieve or rank?”—no. Chrome is answering “does this help an agent orient itself once it is on the site?”—plausibly yes. The conflation of those two questions is the single largest source of confusion in the llms.txt discourse, and vendors on both sides of the argument benefit from leaving it conflated.

Google spokespeople have been consistently dismissive on the search side. John Mueller has publicly compared it to the keywords meta tag—“AFAIK none of the AI services have said they’re using LLMs.TXT… it’s comparable to the keywords meta tag”—and separately described it as a “temporary crutch, perhaps to save some tokens” for AI coding tools reading developer documentation, explicitly “not done for search”. Gary Illyes is widely reported as saying at Search Central Live in July 2025 that Google does not support llms.txt and is not planning to. Both quotes reach us through secondary reporting (Ahrefs, 2026-06-15, and search-industry press) rather than from an official Google document, so weight them accordingly—but note that Mueller’s dismissal and Chrome’s audit describe the same narrow use case.

Other model vendors publish it without committing to reading it. Anthropic, OpenAI, and Perplexity all serve llms.txt files for their own developer documentation. None of them, as far as public documentation shows, has committed to fetching or acting on the file in their production crawling or retrieval systems; Anthropic’s own crawler-control guidance directs site owners to robots.txt. Publishing a file for the benefit of other people’s agents is not the same as consuming it, and the two are frequently reported as though they were.

What proponents claim, and what survives

“AI search engines use it to understand your site.” Not supported. No major AI search vendor documents consuming it; Google explicitly disclaims it; the log data shows retrieval bots barely fetch it.

“It improves your citation rate in AI answers.” Not supported. The one large correlational study found no relationship, and its model was more accurate without the variable.

“It saves tokens and reduces the work of extracting content from HTML.” Supported in principle, and this is the original argument. A curated Markdown index is genuinely cheaper to read than a rendered page tree. The catch is that the saving accrues only to a consumer that actually fetches it.

“Coding agents use it.” Modestly supported, and the most interesting live claim. Claude Code out-fetching the AI retrieval bots in Ahrefs’ data is a real observation from real logs, consistent with Mueller’s own characterisation of what the file is for. If you run a documentation site that developers point coding agents at, this is your use case.

“It’s the robots.txt of AI.” Wrong in a way worth correcting. robots.txt expresses access preferences and is near-universally supported—Cloudflare’s April 2026 scan of 200,000 major domains found 78% publish one. llms.txt expresses nothing about permissions and is supported by nobody in particular. If you want to state how your content may be used, that is robots.txt and the Content Signals extension, not this file.

So should you ship one?

Yes, with the right expectations and a small budget.

Ship it if you run documentation, a developer platform, a knowledge base, or any site whose content is genuinely worth reading in bulk. The file is a dozen lines. It is trivially reversible. It costs nothing to serve. And the exercise of writing it is the actual benefit: to produce a good llms.txt you must decide which twenty pages matter, describe each in one line, and confront the pages you cannot justify. Teams routinely discover their information architecture is worse than they thought at exactly this moment. That is worth an hour regardless of whether a single bot ever fetches the result.

Do not ship it if you are expecting it to move visibility, and do not let it be sold to you on that basis. Budget it as documentation hygiene, not as a growth tactic. If a vendor’s proposal line-items llms.txt under AI visibility with a projected outcome, that is a good moment to ask which study they are relying on.

Get the boring parts right first. Nothing in the evidence displaces the ordinary work: be crawlable and indexable, use real semantic structure, put decision-critical information in the page rather than behind interaction, publish dates, and keep robots.txt accurate. Those are prerequisites for every retrieval path that demonstrably exists. llms.txt is a rounding error on top of them.

And keep it honest if you keep it. A stale index is worse than none: it will confidently point an agent at pages you deleted six months ago. If you cannot commit to regenerating it when your docs change, do not publish it. That maintenance obligation, not the file itself, is the real cost.

What would change this assessment

Three things would move the analysis, and it is worth naming them in advance rather than reacting to the next blog post:

  1. A major retrieval vendor documenting consumption. Not publishing a file—documenting that their crawler fetches and uses one. That has not happened.
  2. A controlled experiment with holdouts. Matched sites, randomised assignment, measured citation and referral outcomes over a quarter. Everything published so far is observational or vendor-internal.
  3. Agent frameworks fetching it by default. The Claude Code signal in the Ahrefs logs is the seed of this. If mainstream agent harnesses started checking /llms.txt as a standard orientation step, the file’s value would come from the client side rather than from search, and Chrome’s rationale would be the one that matters.

Until one of those lands, the position that survives the evidence is narrow and slightly boring: llms.txt is a cheap, harmless, occasionally useful documentation index with one demonstrated audience and no demonstrated effect on visibility. Publish it as such.


Sources reviewed for this page were checked on 1 September 2026. Adoption and server-log figures come from third-party studies whose methodologies differ substantially; where studies disagree, both figures are shown rather than reconciled.

Ship one if your documentation deserves a good index. Do not ship one expecting to be found more often.

Read this guide as markdown