The Same Anthropic Crawl-to-Referral Metric Produced Seven “Correct” Answers

The Same Anthropic Crawl-to-Referral Metric Produced Seven “Correct” Answers
Sponsored

Ask for Anthropic's crawl-to-referral ratio and you can get seven dramatically different answers: 70,900:1, 38,000:1, 23,951:1, 11,122:1, 10,300:1, 4,580:1 or 2,237:1. All seven have been attributed to Cloudflare, and all circulated within roughly 13 months. The largest is more than 30 times the smallest.

The obvious conclusion would be that somebody measured the metric incorrectly. But Duane Forrester's September 11 analysis for Search Engine Journal makes a more important point: several apparently contradictory numbers can be valid because they are measuring different windows, different crawler groupings or different populations, while the referral side of the calculation itself is known to miss some traffic.

That turns a dispute over one AI crawler statistic into a broader lesson about AI visibility measurement. A ratio without its time period, bot scope, sample boundary and attribution limitations is not a stable property of an AI company. It is a snapshot produced by a particular methodology. Strip away that methodology and a precise-looking number can become misleading without ever becoming mathematically false.

What the crawl-to-referral ratio actually measures

Cloudflare introduced its crawl-to-refer metric in July 2025 as a way to compare how much content an AI or search platform retrieves with how much web traffic it sends back. The original Cloudflare methodology divides HTML requests made by user agents associated with a platform by HTML requests carrying a Referer header associated with that platform, then normalizes the result to one referral.

A ratio of 100:1 therefore means that the measured crawler activity produced roughly 100 HTML requests for every measured referred visit. A ratio of 1:1 would indicate one crawl request per referral. A ratio below one means the platform generated more measured referrals than crawler requests during that particular observation window.

In Cloudflare's launch analysis, covering June 19–26, 2025, Anthropic appeared at 70,900:1. At the opposite end, Mistral appeared at 0.1:1—roughly ten referrals for every crawl request. The enormous spread helped make crawl-to-refer one of the most memorable statistics in the debate over whether AI platforms consume more publisher content than they return in traffic.

The first source of disagreement is simply time

A ratio is inseparable from the dates used to calculate it. Cloudflare's initial 70,900:1 Anthropic figure covered one week in June 2025. Forrester points out that another Cloudflare publication from the same month reported Anthropic at approximately 73,000:1, while later analyses use monthly, quarterly and rolling 28-day periods.

Those are not interchangeable observations. Crawler behavior can change rapidly when a company modifies crawl schedules, launches products, expands an index or performs a large collection pass. Referral traffic can change at the same time as usage of the consumer product grows or falls. Moving either side of the fraction changes the ratio.

Forrester highlights a Cloudflare example in which Google's crawl-to-refer ratio moved 19.4% week over week after a decline in GoogleBot crawling that began on a particular day. Nothing about the formula had changed. The observation window simply captured a meaningful shift in crawler activity. Two analysts choosing adjacent but different periods could therefore publish different numbers while both calculations remained valid.

This is why a current Anthropic ratio should not be compared casually with the 70,900:1 launch figure as though they describe a fixed characteristic of Claude. The difference may reflect real changes in platform behavior, different date ranges or both.

“Anthropic” can represent several different crawlers

The second problem is aggregation. AI companies can operate different bots for different purposes: model training, search indexing and user-requested retrieval are not the same activity. Yet a network-level platform statistic may combine multiple user agents under one company name.

Forrester notes that Cloudflare's platform analysis aggregates a provider's training crawler and user-request crawler even though their economic relationship with a publisher can be very different. A training crawler may fetch large volumes of content without being designed to send a click. A user-request crawler may fetch a page because someone is actively asking a question and could generate a citation or visit.

Combine those behaviors and the resulting ratio describes the company-level traffic mix, but it does not necessarily describe any individual bot. That matters for robots.txt decisions. A publisher looking only at an aggregate ratio could conclude that an entire provider returns too little value and block crawlers indiscriminately, even though one crawler may be responsible for answer-time retrieval while another accounts for most of the high-volume collection.

The comparison with traditional search can also become structurally uneven. Providers do not all divide crawler functions in the same way. Comparing a multi-bot AI fleet with a more unified search crawler can therefore create a clean-looking table whose rows represent different operational objects.

The sample of websites can move the answer

The third variable is where the measurement was collected. Cloudflare's network is enormous, but it is still a particular population of websites. A separate commercial panel, an individual publisher's logs or another infrastructure provider will observe a different mix of industries, site sizes, geographies and crawler activity.

Forrester describes a comparison in which one platform's ratio roughly doubled when measured against a smaller commercial panel instead of Cloudflare's broader network during the same period. That does not imply either dataset was defective. It means sample composition was part of the metric.

This is especially important when publishers ask whether a global crawl-to-refer ratio describes their own site. It does not. A news publisher, ecommerce catalog, developer documentation site and niche B2B publication may receive very different crawler attention and very different AI referral behavior. Network-level statistics are useful for understanding the ecosystem, but server logs and first-party analytics remain necessary for domain-level policy.

The denominator misses referrals that do not identify themselves

The fourth limitation is arguably the most consequential: the referral count does not include every referral. It includes requests that arrive with a Referer header naming a hostname Cloudflare associates with the platform.

Cloudflare disclosed in the original methodology that traffic from Claude's native app does not carry a Referer header and said it believed the same issue affected other providers' native applications. As a result, the company warned that its calculations could overstate crawl-to-refer ratios because the referral denominator captures web-based traffic while omitting an unknown quantity of app-originated visits.

The size of that blind spot is not publicly quantified. If native-app usage becomes a larger share of AI activity, the missing-referrer problem can become more important even if the crawler side of the measurement remains perfectly accurate. A ratio may therefore say “70,900 crawls per measured referral,” but that is not necessarily equivalent to “70,900 crawls per human visit actually generated.”

This limitation is not unique to Cloudflare or evidence of bad methodology. It is a web attribution problem. When a client does not send identifying referral information, downstream measurement systems cannot recover it reliably from the request alone. The important analytical mistake happens later, when a carefully qualified measured-referral metric gets repeated as if it counted all referrals.

How seven “correct” Anthropic numbers emerged

Search Engine Journal lists Anthropic ratios of 70,900:1, 38,000:1, 23,951:1, 11,122:1, 10,300:1, 4,580:1 and 2,237:1, all presented downstream as Cloudflare-derived figures. The spread is not explained by seven competing formulas. It is the accumulated effect of measurements made at different times, with different aggregation choices and sometimes different panels, then repeated without all of their original context.

The 2,237:1 figure, for example, has circulated for a rolling 28-day period ending July 21, 2026, whereas the famous 70,900:1 figure refers to the one-week June 19–26, 2025 launch window. Putting those numbers into the same sentence can illustrate how much the reported ratio changed, but calling either one simply “Anthropic's ratio” erases the dates that make the comparison meaningful.

Forrester's larger concern is what happens as metrics travel through secondary reporting. The primary source may publish the window, crawler definitions and caveats. The next article shortens the explanation. Another source quotes that article and drops the date. Eventually a slide deck contains a company name and a ratio with none of the assumptions that produced it. The number still looks authoritative because the original attribution remains attached.

A precise number can be accurate and still be unusable

This distinction is increasingly important in AI search because companies are making policy decisions from immature measurement systems. Publishers use crawler-to-referral ratios to decide which bots to allow, block or potentially charge. Marketing teams use AI referral statistics to decide whether optimization for generative discovery deserves investment.

Those decisions can be reasonable, but a context-free ecosystem average is a weak foundation for them. A temporary training crawl can inflate the numerator. A native app can hide part of the denominator. Aggregating several crawler purposes can obscure which bot creates the imbalance. A network-wide sample can behave differently from the publisher making the decision.

The correct response is not to abandon the metric. Crawl-to-referral ratio expresses a genuinely important question about the exchange between publishers and AI systems: how much content does a platform retrieve relative to the measurable traffic it returns? The response should be to keep the dimensions attached to the answer.

Publishers should measure crawler value at the bot and site level

For operational decisions, first-party server logs provide the most relevant numerator. Publishers can identify crawler user agents, inspect request volume over time and separate bots by declared purpose where reliable documentation exists. That makes it possible to distinguish a training crawler from search or user-triggered retrieval rather than applying one company-level ratio to every bot.

The referral side should be treated with similar care. Analytics can identify visits carrying recognizable AI referrers, but direct and unattributed traffic may contain visits from apps that strip referral information. Tagged links can improve attribution where platforms use them, but publishers cannot force an external native app to transmit a Referer header.

A useful internal dashboard should therefore label the metric honestly—for example, crawls per measured identifiable AI referral—and attach the date range. Teams can then compare the trend on their own site rather than treating a global snapshot as a permanent benchmark. A ratio moving from 5,000:1 to 1,000:1 on the same property under the same measurement rules can be more actionable than discovering that another publisher reports 2,237:1.

The same problem is spreading across AI visibility analytics

Crawl-to-refer is only one example of a broader measurement challenge. AI citation share, answer visibility, sentiment, referral traffic and prompt coverage can all change depending on which prompts were tested, which models were included, when the test ran, whether answers were personalized and how a vendor defined a citation or mention.

That does not make AI visibility measurement meaningless. It means methodology is part of the result. Two tools can legitimately report different visibility scores for the same brand because they queried different prompt sets or sampled different engines. Just as with crawler ratios, the useful question is not only “what is the number?” but “what exactly entered the numerator and denominator?”

Forrester proposes a simple discipline: when someone presents a measurement, ask what period it covers, what was grouped together and where collection stopped. For crawl-to-referral ratios, a fourth question should be added immediately: what traffic could not be attributed at all?

Seven answers are a warning against false certainty

The Anthropic example is memorable because the range is so large. A ratio of 70,900:1 paints a very different intuitive picture from 2,237:1, even though both can trace back to Cloudflare-derived measurement. Without dates and definitions, readers naturally interpret the disagreement as error. With them, the disagreement becomes evidence that the metric is dynamic and conditional.

That is the real lesson for technical SEO teams and publishers. Do not ask for “the” crawl-to-referral ratio of an AI platform as though it were a permanent exchange rate. Ask for the ratio over a specific period, for specified crawler identities, across a defined site population, using a stated referral methodology.

Cloudflare's original work was valuable precisely because it documented those choices and acknowledged what the denominator could not see. The problem begins when the context disappears while the integer survives. In AI search measurement, seven different answers can all be numerically defensible. The dangerous one is the answer presented without enough methodology to know what it actually counted.

0%