41% of B2B SaaS Companies Use llms.txt — but Almost Nobody Can Prove It Works

41% of B2B SaaS Companies Use llms.txt — but Almost Nobody Can Prove It Works
Sponsored

The adoption curve for llms.txt is moving faster than the evidence curve. In a scan of 100 B2B SaaS companies, AI visibility platform Treyci found that 41 had already published an llms.txt file — yet the company says few organizations can quantify whether their AI visibility work has actually changed how often answer engines recommend them.

That gap is more important than the adoption percentage itself. Publishing a small Markdown file at the root of a website is easy. Demonstrating that the file caused ChatGPT, Gemini, Perplexity or another system to cite or recommend the brand more often is a much harder measurement problem.

The 41% figure comes from a September 3 announcement from Treyci, which also describes its broader methodology for measuring brand mentions, recommendations and citations across multiple AI engines. The release does not present a controlled experiment comparing the 41 adopters with the 59 non-adopters, so it cannot establish whether llms.txt improves AI visibility.

What it does reveal is a familiar pattern in emerging SEO markets: implementation has become a KPI before impact has been established.

What 41% adoption actually tells us

Treyci's finding shows that llms.txt has achieved meaningful penetration within its 100-company B2B SaaS sample. For a convention proposed only in 2024, 41 installations out of 100 is substantial.

But adoption is not efficacy. Companies frequently implement low-cost technical changes because the downside appears limited and the potential upside is uncertain but attractive. That behavior can make a tactic spread rapidly even without evidence that it changes the outcome marketers care about.

The public Treyci release does not provide the complete company list or a matched comparison of citation performance between sites with and without the file. It therefore supports a statement about adoption, not a causal statement about effectiveness.

The useful question is not “How many companies have llms.txt?” It is “What happened after they added it?”

llms.txt was designed as a machine-readable orientation layer

The llms.txt proposal was introduced by Jeremy Howard and Answer.AI as a way for websites to provide LLM-friendly information in a predictable location. A file placed at /llms.txt can summarize the site and point machines toward important resources in Markdown.

The concept is understandable. Modern websites contain navigation, advertising, JavaScript, repeated interface elements and enormous numbers of URLs. A concise machine-oriented index could theoretically help an agent identify the most relevant documentation without processing an entire site.

That is a legitimate engineering use case. It is not the same claim as saying the file is a ranking signal, citation signal or recommendation signal inside consumer AI search products.

Much of the confusion around llms.txt comes from allowing those two propositions to collapse into one.

Google Search has explicitly separated itself from the file

For Google Search, the current answer is unusually clear. Google's official 2026 guidance for generative AI features says websites do not need special machine-readable files, AI text files, markup or Markdown to appear in Google Search, including its generative AI experiences, because Google Search itself does not use them.

The guidance specifically names llms.txt in its myth-busting section. That means publishing the file should not be presented as a way to improve visibility in Google AI Overviews or AI Mode.

Search Engine Journal also reported in June that Google's John Mueller described current llms.txt implementation as speculative from a Search perspective. The Search team has consistently emphasized conventional accessibility, indexing and content fundamentals instead.

This does not prove that no software anywhere can benefit from llms.txt. It establishes that one of the largest generative search surfaces says it does not use the file.

Google Search and agent tooling are different use cases

The distinction becomes important because other parts of the technology ecosystem may have different incentives. Developer tools and agents that need to navigate technical documentation can benefit from a curated, compact map of a site even if a search engine does not use that map for ranking or citation selection.

A file can therefore be useful as agent-readiness infrastructure without being useful as an AEO ranking tactic.

This is not merely semantic. The expected outcome changes depending on the use case. If the objective is to help an explicitly directed coding agent find API documentation efficiently, success can be measured through agent behavior and retrieval efficiency. If the objective is to increase unsolicited ChatGPT recommendations, the experiment needs to measure recommendations.

Calling both outcomes “AI visibility” makes it too easy to claim success without defining what success means.

Ahrefs found that most published files were never requested

A separate 2026 study from Ahrefs adds another layer of evidence. Ahrefs examined 137,210 domains using its Web Analytics product and found valid llms.txt files on roughly 28% of them.

More strikingly, 97% of the valid files received no requests during May 2026. Only about 3% were fetched at all in the observed logs.

The sample is not representative of the entire web — Ahrefs explicitly notes that its analytics users skew more technical and SEO-aware — and a one-month request log does not prove that a file can never influence any AI workflow.

Still, the finding introduces a basic measurement question that should precede citation analysis: is the system you hope to influence even requesting the file?

A file that is not fetched cannot explain a downstream recommendation

Causal reasoning matters here. If a company adds llms.txt and later receives more AI mentions, the timing alone does not show that the file caused the improvement.

The brand may also have published new product pages, earned press coverage, accumulated reviews, improved its documentation, launched a campaign or simply benefited from a model update.

Server logs can provide an initial filter. If the relevant crawler or agent never requests /llms.txt, it becomes difficult to argue that the contents of that file directly changed that system's answer through retrieval.

Even a fetch is only the beginning of the evidence chain. A request does not prove the file was used in retrieval, that retrieval influenced generation or that generation changed the brand's recommendation probability.

Treyci's own methodology points toward the missing experiment

The irony is that Treyci's broader AI visibility methodology provides a reasonable framework for testing the tactic more rigorously.

The platform says it runs approximately 100 buying-intent questions per category across ChatGPT, Gemini, Perplexity and Grok, repeating every prompt three times and scoring more than 1,200 resulting answers for mentions, recommendations and citations.

A company considering llms.txt could establish a baseline using a stable prompt set, publish the file, monitor whether relevant agents request it and continue the same repeated measurements afterward.

That still would not create a perfect controlled experiment because models and source indexes change over time. But it would be considerably stronger evidence than publishing the file and later pointing to a favorable screenshot.

Measure mentions, recommendations and citations separately

Another measurement problem is deciding what llms.txt is supposedly improving. “AI visibility” can refer to several different outcomes.

A mention means the brand appears somewhere in an answer. A recommendation means the system actively places it among the options a buyer should consider. A citation means a particular source is linked or attributed as evidence.

Those outcomes are related but not interchangeable. A technical file might theoretically make documentation easier for an agent to retrieve without increasing the frequency with which a brand is recommended in commercial prompts.

Any effectiveness claim should therefore name the outcome. “We implemented llms.txt” is an implementation metric. “Our citation rate increased from X to Y under a controlled prompt panel” is an outcome metric.

Recommendation lift is a harder standard than crawler access

The title question becomes especially demanding when the desired outcome is recommendation rather than retrieval.

An AI system can successfully access a company's site and still prefer a competitor. Recommendation can depend on category fit, product capabilities, pricing, reviews, third-party comparisons, reputation and the evidence retrieved from across the web.

Treyci's same announcement reports that cited sources in buying answers are dominated by review platforms, comparison articles and industry publications rather than vendor websites. If that pattern holds in a category, improving one machine-readable file on the vendor domain may address only a small part of the recommendation environment.

Technical accessibility can be necessary without being sufficient.

The external web may matter more than a self-authored summary

This is one reason AI visibility is difficult to engineer through a single on-site tactic. A company controls its product pages, documentation, structured data and llms.txt. It does not control how independent reviewers compare the product or which competitors appear in editorial roundups.

Answer engines can combine those sources into a market model. A self-authored file may tell the system what the company claims to offer, while third-party evidence tells the system whether other sources corroborate that positioning.

For commercial recommendations, corroboration can be more valuable than another version of the company's own description.

This suggests that AI visibility budgets should not automatically move toward new technical artifacts at the expense of product clarity, original research, digital PR, reviews and credible third-party coverage.

NetContentSEO data points toward the same measurement discipline

NetContentSEO's analysis of 366,000 generative-search impressions reached a related conclusion from a different dataset: AI visibility should be investigated from observed outcomes backward rather than from fashionable tactics forward.

That means identifying pages with unusually high or low generative visibility, comparing them with conventional search performance and then testing hypotheses about why they differ.

The opposite workflow starts with a tactic — add llms.txt, rewrite headings as questions, create AI summaries — and assumes the tactic matters before measuring whether the target systems behave differently.

The current evidence around llms.txt makes it a particularly useful case study in why that order matters.

Low implementation cost does not make the claim true

There is a reasonable argument for publishing llms.txt anyway. The file can be inexpensive to create, it may be useful to some agent or developer-tool workflows, and future systems could adopt the convention more widely.

That is an optionality argument, not an effectiveness argument.

A company can rationally spend a small amount of engineering time on a low-cost experiment whose upside is uncertain. What it should not do is report the deployment itself as evidence that AI visibility has improved.

The difference matters for budget allocation. Cheap experiments can coexist with skepticism, provided they remain experiments.

Keep the experiment clean enough to learn something

Companies that want to test llms.txt should resist changing ten other AI-search variables at the same time. If the site publishes the file while rewriting every product page, launching a PR campaign and adding thousands of review citations, any subsequent visibility movement becomes impossible to attribute.

A stronger test begins with a fixed panel of commercially relevant prompts and multiple runs per engine. Baseline mention, recommendation and citation rates are recorded before deployment. Server logs track whether relevant agents request the file. The same prompts are then repeated over a defined period.

Control groups can improve the design further. Similar sections, product lines or sites without the change can help distinguish a broad model update from a site-specific effect.

The objective is not laboratory perfection. It is to make the evidence better than “we added it and AI traffic went up later.”

Do not average away engine differences

Treyci's broader findings also show why an llms.txt experiment should report engines separately. In one B2B software category, the company observed brand mentions in 81% of answers on one AI engine and only 43% on another.

If the engines already disagree that sharply about which brands belong in an answer, there is little reason to assume they will respond identically to a new machine-readable file.

A combined AI visibility score could rise even if ChatGPT did not change, simply because Gemini or Perplexity moved. The reverse could also happen.

Any claim that llms.txt “works” should eventually specify where it works, for which outcome and under what conditions.

Adoption can create its own illusion of consensus

Forty-one installations in a 100-company SaaS sample can make a convention feel established. Once enough recognizable companies publish a file, others copy it because the adoption itself looks like evidence.

This is a common technology diffusion pattern. Teams assume peers know something they do not, vendors add the feature because customers request it, CMS platforms make implementation one click and the rising adoption rate becomes the justification for further adoption.

None of that requires the underlying mechanism to have been demonstrated.

The antidote is outcome measurement. If a tactic is widespread and effective, the industry should eventually be able to show repeatable differences in the metrics the tactic is supposed to improve.

The right status for llms.txt is “testable,” not “proven”

The evidence in 2026 does not support treating llms.txt as a universal AI ranking or citation lever. Google Search explicitly says it does not use the file. Ahrefs found that the overwhelming majority of files in its dataset were not requested during the month it studied. Treyci found substantial SaaS adoption but highlighted the industry's inability to quantify the effect of its AI visibility work.

None of those findings prove that llms.txt has zero value in every context. Agent tooling, developer documentation and future machine interfaces may provide legitimate use cases, and the convention can remain inexpensive to test.

But the burden of proof should now move from implementation to measurement.

Forty-one percent adoption is evidence that B2B SaaS companies are interested in llms.txt. It is not evidence that the file earns citations or recommendations. Until controlled measurements connect the two, the most accurate label is not “AI optimization best practice.” It is “an experiment that still needs a result.”

0%