The AI Crawl-to-Referral Metric Publishers Use to Block Bots May Be Missing Most Native-App Traffic

The AI Crawl-to-Referral Metric Publishers Use to Block Bots May Be Missing Most Native-App Traffic
Sponsored

Publishers deciding whether to allow or block AI crawlers increasingly rely on a deceptively simple metric: how many pages a bot fetches for every visitor its platform sends back. But one of the most widely quoted examples shows why that ratio can become misleading once it is separated from its methodology.

Anthropic’s crawl-to-referral ratio has appeared in analyses at values ranging from 70,900:1 to 2,237:1, a difference of more than 30 times. Both endpoints have been attributed to Cloudflare-derived measurements. As Duane Forrester explains in a September 10, 2026 Search Engine Journal analysis, the discrepancy is not evidence that one number must be fabricated. It is evidence that a ratio cannot be interpreted without knowing the time window, the sites measured, the crawlers grouped together and the traffic the denominator cannot see.

What the crawl-to-referral ratio actually measures

Cloudflare introduced its crawl-to-refer metric in 2025 as a way to compare AI crawler activity with the human traffic those platforms send to websites. The original Cloudflare methodology is straightforward: count HTML requests from user agents associated with a platform, then divide them by HTML requests carrying a Referer header associated with that platform. The result is normalized to one referral.

A ratio of 100:1 therefore means the measured crawler activity was 100 times the measured referral activity. The appeal for publishers is obvious. Traditional search engines historically crawl pages while also sending substantial traffic back to the sites they index. Generative AI systems can consume information and answer users directly, potentially creating a much more asymmetric exchange.

The problem is not the formula itself. The problem begins when a result produced under one set of conditions is repeated as though it were an enduring property of a company or bot.

Anthropic has been quoted anywhere from 70,900:1 to 2,237:1

Cloudflare’s launch analysis used the period from June 19 to June 26, 2025 and reported Anthropic at roughly 70,900 crawls for every measured referral. More recent Cloudflare-derived reporting has put the ratio near 2,237:1 for a rolling 28-day period ending July 21, 2026. Between those endpoints, other published figures have placed Anthropic at 38,000:1, 23,951:1, 11,122:1, 10,300:1 and 4,580:1.

Those values should not be averaged into a single “true” Anthropic ratio. They are measurements made under different conditions. A crawl-to-referral ratio can move because crawler volume changes, because referral volume changes or because both change at once. Even a different date range can materially alter the result.

Search Engine Journal points to Cloudflare data showing Google’s ratio moving 19.4% in a week after a change in GoogleBot crawling. That illustrates how sensitive the metric can be to short-term crawler behavior. A heavy crawling period can make a platform look dramatically more extractive even if the change is temporary.

The sites in the sample matter too

Cloudflare observes an enormous portion of web traffic, but its network is still a defined population rather than the entire web. The mix of publishers, ecommerce sites, forums, software companies and other properties behind Cloudflare affects the aggregate ratio.

A smaller commercial panel can produce a different answer during the same period because its sites attract different crawler behavior and different levels of AI referral traffic. Forrester cites an example in which changing the measured site population roughly doubled one platform’s ratio.

That makes operator-level figures useful as ecosystem indicators but dangerous as substitutes for a publisher’s own server data. A news organization, SaaS documentation site and ecommerce catalog may experience the same crawler very differently. Before making an access-control decision, publishers should compare public benchmarks with what is actually happening on their own infrastructure.

Training crawlers and user-triggered bots can be mixed together

Another complication is crawler grouping. Cloudflare’s platform-level analysis can aggregate multiple user agents associated with the same company. Those bots may serve fundamentally different purposes.

A training crawler can fetch content at scale without any expectation that an individual request will immediately create a referral. A user-triggered retrieval bot, by contrast, may access a page because someone asked an AI assistant a question and the system needs current information to answer it. Combining the two creates a convenient operator-level number, but it can obscure the behavior publishers actually care about.

This matters when a ratio is used to decide which bots to block. A publisher trying to stop large-scale training crawls may also restrict a retrieval crawler capable of surfacing the site as a source in a user-facing answer if crawler purposes are not distinguished correctly. The economic trade-off can therefore depend on user-agent-level policy rather than a single company-wide ratio.

Native apps create an even larger blind spot

The most consequential limitation sits in the denominator. Cloudflare counts a referral when an incoming request includes a Referer header associated with the AI platform. But not every AI-generated visit announces where it came from.

Cloudflare explicitly states that traffic from Claude’s native app does not include a Referer header and says it believes the same is true for native apps from other providers. Those visits therefore do not enter the referral side of the calculation even when a user actually arrived at a publisher after interacting with an AI assistant.

That means the metric is not strictly “all crawls divided by all referrals.” It is measured crawls divided by referrals that can be identified through the required web signal. If native-app traffic represents a meaningful share of user activity, the denominator is understated and the resulting crawl-to-referral ratio is overstated.

The size of that distortion is unknown. Cloudflare itself says it is unclear by how much the ratios may be overstated. That uncertainty is crucial: it does not prove that AI platforms secretly send enough native-app traffic to erase the crawl imbalance, but it does mean the published ratio cannot measure that missing traffic.

Publishers should not turn one public ratio into a permanent bot policy

Crawl-to-referral ratios have real operational value. A publisher paying for bandwidth and infrastructure has a legitimate reason to understand which automated systems consume resources and what observable value comes back. Cloudflare has also developed controls specifically to give site owners more choice over AI crawler access.

The danger comes from converting a time-bound aggregate benchmark into a permanent policy without examining what produced it. If a high ratio coincided with an unusually intensive training crawl, blocking based on that snapshot may solve a temporary problem. If the platform’s native app sends untracked traffic, the measured return may be lower than the real return. And if different crawler types are grouped together, an operator-wide block may remove useful retrieval alongside unwanted collection.

A better approach is to preserve the dimensions that disappear from viral statistics: date range, crawler identity, site population and referral-detection method. Publishers with sufficient infrastructure can then compare public ecosystem data with server logs, bot verification, analytics and conversion behavior on their own properties.

AI referral measurement has an attribution problem

The native-app issue also exposes a broader weakness in AI traffic analytics. Web analytics was built around mechanisms such as referrer headers, campaign parameters and browser navigation. AI discovery increasingly happens inside mobile and desktop applications that do not necessarily preserve those signals when users leave for the open web.

As a result, some traffic that originated with an AI assistant may appear as direct or unattributed traffic in analytics platforms. That makes it harder to calculate the economic value of AI visibility, compare assistants fairly or determine whether citation exposure eventually produces useful visits.

This is not unique to crawler ratios. Any report claiming to quantify AI referrals needs to explain how it identifies the source and what happens when that attribution signal is absent. Without that information, a precise-looking number may be measuring only the trackable portion of the channel.

The lesson is methodological, not a new universal benchmark

The 30-fold spread in reported Anthropic ratios should not be replaced with another supposedly definitive number. The lesson from the Search Engine Journal analysis is precisely the opposite: several figures can be accurate within their own collection boundaries while answering subtly different questions.

For any crawl-to-referral statistic, publishers should ask at least three questions before acting on it: what period was measured, which crawlers and sites were grouped together, and which referrals were technically detectable. If those details are missing, the ratio has lost the context required to interpret it.

AI crawlers may indeed consume vastly more pages than their platforms send back as measurable traffic. Cloudflare’s work has made that imbalance visible and given publishers a valuable framework for evaluating it. But the metric is not a universal exchange rate between crawling and value. In particular, native-app traffic can disappear from the denominator entirely, meaning the number used to justify blocking a bot may be missing an unknown share of the very referrals it is supposed to count.

0%