Ranking on Google Doesn’t Mean ChatGPT Will Recommend Your Brand

Ranking on Google Doesn’t Mean ChatGPT Will Recommend Your Brand
Sponsored

A page can rank on the first page of Google and still disappear when a buyer asks ChatGPT which companies to consider. That gap is becoming one of the central measurement problems in answer engine optimization: traditional rankings tell marketers where a URL appears in search, but they do not tell them whether an AI system will mention, prioritize or recommend the brand.

A new sponsored Search Engine Journal guide from HubSpot proposes a practical way to measure that difference. Instead of collapsing AI visibility into one score, it separates three questions: how often the brand is mentioned, how often it appears among the leading recommendations and how much of the total brand conversation it captures relative to competitors.

The framework then uses Google Search Console as the starting point. Commercial queries for which the site already ranks become prompts that are repeatedly tested across ChatGPT, Gemini and Perplexity. The result is a bridge between an established SEO dataset and a newer, less stable AI discovery environment.

The methodology is useful precisely because it exposes something rankings cannot: a brand can be visible to Google without being a meaningful option in an AI-generated buying answer.

A mention and a recommendation are not the same outcome

The most important distinction in HubSpot's framework is between being named and being recommended.

A mention occurs whenever the brand appears somewhere in the generated answer. The model may list it in passing, include it as one of many examples or reference it while discussing the category.

A recommendation is stronger. HubSpot's proposed audit treats an appearance among the first three brands named as the recommendation signal, reflecting whether the brand occupies a prominent place in the buyer's consideration set.

This prevents an AI visibility dashboard from celebrating a large number of low-value mentions. A company can appear frequently while rarely being positioned as one of the options a buyer should seriously consider.

Mention rate answers the basic visibility question

The first metric is mention rate. HubSpot defines it as the number of prompts in which the brand appears divided by the number of prompts tested.

If a company tests 100 relevant buying prompts and appears in 35 answers, its mention rate is 35% for that sample. The calculation is straightforward, which makes it useful as a baseline.

But mention rate cannot describe the quality of the appearance. A brand mentioned tenth in a long answer and a brand introduced as the leading option both count as mentions.

That is why reporting the number alone can exaggerate commercial visibility. It answers “Are we present?” rather than “Are we being selected?”

Recommendation rate tries to measure prominence

HubSpot's second metric divides top-three appearances by the number of prompts tested. This recommendation rate is intended to capture whether the brand is being surfaced prominently enough to influence consideration.

The sponsored article argues that recommendation is the metric more closely connected to pipeline. That is intuitively plausible: being actively proposed as an option is more commercially meaningful than being named somewhere in the response.

However, the article does not present a controlled study establishing that a particular increase in recommendation rate causes a particular pipeline increase. Recommendation rate should therefore be treated as a stronger leading indicator of commercial prominence, not as a substitute for revenue attribution.

The distinction is important for AEO reporting. Visibility metrics describe what AI systems say. Business metrics still need to establish what customers do afterward.

Share of voice adds the competitive context

The third metric is share of voice. HubSpot calculates it by dividing the brand's mentions by all brand mentions returned across the tested answers.

This changes the question again. Mention rate asks how frequently a company appears. Share of voice asks how much of the category's AI attention the company captures relative to everyone else.

A brand could maintain the same mention rate while losing share of voice if competitors begin appearing much more frequently. Conversely, share of voice could rise even if absolute visibility remains modest, depending on how the rest of the competitive set changes.

The three measures therefore describe different dimensions of the same environment. Mention rate captures coverage, recommendation rate captures prominence and share of voice captures competitive presence.

Search Console becomes the prompt-research dataset

The most practical part of the HubSpot methodology is its starting point. Instead of inventing a list of prompts from scratch, the guide recommends exporting commercial-intent queries from Google Search Console.

The suggested process begins with the Performance report, a three-month date range and queries where the site's average position is 10 or better. The list is then narrowed toward commercial language such as “best,” “software,” “tool,” “platform,” “vs,” “alternative,” “pricing” and relevant category terms.

These are useful candidates because the site already has evidence of traditional search relevance. The AEO audit then asks whether that relevance carries into generative recommendations.

This creates a particularly actionable gap list: queries where Google already gives the brand first-page visibility but AI systems fail to mention it.

A keyword cannot simply be pasted into ChatGPT and called a prompt

The transition from Search Console query to AI prompt requires judgment. Search keywords are often compressed expressions of intent. Conversational prompts can contain constraints, context and decision criteria.

A query such as “project management software” might become a prompt asking for the best project management platforms for a small distributed team with a particular budget and integration requirement. A query such as “Shopify alternatives” can become a request for alternatives suited to a specific merchant type.

The purpose is not to make every prompt artificially long. It is to represent the kinds of questions a buyer would realistically ask an assistant while preserving the commercial intent visible in Search Console.

This matters because an AEO measurement program is only as useful as the prompt set it tracks. Measuring irrelevant prompts precisely still produces irrelevant intelligence.

One prompt run is not a measurement

HubSpot recommends running each prompt two or three times and recording the answer that appears most often, then repeating the exercise across each engine.

The repetition is essential because generative answers are stochastic. ChatGPT, Gemini and Perplexity can produce different brands, sources and ordering when the same question is asked again.

Independent research covered by Search Engine Journal has highlighted this instability. A July analysis on AI visibility measurement noise warned that rankings and citation shares can shift between runs, making individual observations poor evidence of a durable difference.

Two or three repetitions are better than one, but teams making high-stakes decisions should still treat small changes cautiously. Sample size, prompt selection and repeated measurement waves all influence the apparent trend.

Report the engines separately

The HubSpot framework calculates visibility for each engine separately and then allows an optional composite across ChatGPT, Gemini and Perplexity. Crucially, it also recommends reporting the per-engine figures alongside that aggregate.

That separation matters because there is no single AI search index. The platforms use different retrieval systems, source mixes, model behavior and product interfaces.

A brand can be a frequent recommendation in Perplexity while barely appearing in ChatGPT. Averaging those outcomes into one score makes the executive dashboard cleaner but can hide the diagnostic information needed to improve performance.

A composite is useful as a summary. It should not become a substitute for understanding where the visibility actually exists.

Log the answer, not just the score

HubSpot recommends recording six fields for every tested prompt: the engine, the brands mentioned, the order in which they appear, whether the company's brand appears, whether it is among the first three and which domains are cited.

This raw answer log is more valuable than a single visibility percentage because it preserves the evidence behind the metric.

If recommendation rate falls, the team can inspect which competitors replaced the brand. If mention rate stays stable while share of voice declines, the log can reveal which companies are gaining exposure. If a brand disappears from a particular prompt, the cited domains can show which sources are shaping the new answer.

The dashboard says that visibility changed. The transcript helps explain what changed with it.

The source mix is where measurement becomes an action plan

The sponsored guide groups common AI sources into review directories such as G2 and Capterra, online communities, third-party roundup articles, news coverage, reference pages and the company's own website.

That broader source mix helps explain why a first-page Google ranking does not guarantee a recommendation in an AI answer. A model assembling a buying response may draw evidence from several independent sites that discuss the category, not only from the company's highest-ranking page.

This does not mean traditional SEO becomes irrelevant. The HubSpot article explicitly says technical health, site structure, schema and content quality remain foundational.

The difference is that AI visibility can depend on what the wider web says about the brand as well as what the brand says about itself.

The most useful gap is “ranking but not recommended”

Once the Search Console and AI datasets are connected, one segment becomes especially interesting: commercial queries where the site ranks in Google's top 10 but the brand is absent from AI answers.

Traditional SEO has already cleared one relevance threshold. The brand has a page Google considers competitive for the query. Yet the generative layer is not carrying that visibility forward.

That narrows the diagnostic problem. The team can investigate whether competitors have stronger third-party validation, whether review sites describe them more clearly, whether the brand lacks comparison coverage or whether the AI engine is citing sources that barely mention the company.

The ranking becomes a control of sorts. It does not prove the site should be recommended by an LLM, but it shows that the visibility gap cannot be explained simply by saying the company has no search relevance at all.

AEO measurement should separate brand recognition from buyer consideration

This distinction is reinforced by other 2026 research. A separate sponsored study covered by Search Engine Journal found that AI systems could accurately describe most brands when asked directly, while many of those brands rarely appeared in category research prompts.

That is a useful warning against branded prompt testing. Asking “What is Company X?” mainly tests whether the model recognizes Company X. It does not show whether the company surfaces when a buyer has not already selected the brand.

Commercial category prompts are more revealing because the model has to choose among alternatives.

AEO visibility should therefore be measured on unbranded or category-level buying questions alongside branded tests, with the two groups kept analytically separate.

Zero-click visibility still needs a business layer

The HubSpot methodology is designed partly for interactions that produce no website visit. If an AI answer recommends a brand but the user does not click a citation, conventional analytics may record nothing.

Mention rate, recommendation rate and share of voice can help fill that observational gap. They show whether the company is participating in AI-mediated consideration even when no session appears in analytics.

But visibility is not the same as commercial impact. A company still needs downstream evidence: branded search changes, direct traffic, self-reported attribution, assisted conversions, CRM notes or other signals connecting AI exposure with buyer behavior.

Search Engine Journal's broader AI visibility measurement guidance similarly argues that prompt selection and cross-model visibility matter, while marketers need to connect those leading indicators to business outcomes rather than treating citations as revenue.

Sponsored methodology deserves the same scrutiny as any vendor metric

The Search Engine Journal guide is explicitly sponsored by HubSpot and written by a HubSpot senior marketer. The article also promotes HubSpot AEO, which automates prompt tracking, citation analysis, share of voice and recommendations across the three engines.

That commercial context does not make the methodology unusable. The distinction between mentions, recommendations and competitive share is genuinely helpful, and the Search Console-to-prompt workflow can be implemented manually without buying the product.

It does mean marketers should separate the measurement framework from claims about any particular platform's superiority or the precise business meaning of proprietary scores.

A good AEO measurement system should remain intelligible outside the vendor dashboard. Teams should be able to explain exactly what was tested, how often, on which engines and how every reported rate was calculated.

Freeze the prompt set before comparing periods

Repeated measurement only becomes meaningful when the underlying test remains sufficiently stable. If a team changes half its prompts between August and September, a change in visibility may reflect the new sample rather than a real change in how AI systems treat the brand.

Prompt suites should therefore be versioned. Core commercial prompts can remain fixed for trend measurement, while a separate exploratory set captures new products, emerging buyer language and changing market conditions.

The same principle applies to engines and model versions where they can be identified. AI products change rapidly, and a model rollout can alter answers even when the brand has changed nothing.

Without that experimental discipline, AEO dashboards can create an illusion of precision around a moving target.

Ranking and recommendation answer different questions

A Google ranking answers a relatively narrow question: where does this page appear for this query in a search environment? AI recommendation asks something broader: when a model synthesizes an answer from its available knowledge and retrieved sources, does it decide that this brand belongs among the options worth presenting?

The two systems overlap, but they are not equivalent.

That is why a company can rank well and still fail an AI visibility audit. It may have a strong page but weak third-party corroboration. It may be recognized but not considered. It may appear in one engine and disappear in another. It may be mentioned frequently without ever reaching the top of the recommendation set.

Traditional rankings remain valuable evidence. They are simply no longer a complete measure of search-mediated brand discovery.

AEO needs a measurement vocabulary before it needs another score

The most useful contribution of the HubSpot guide is not a proprietary “AI visibility score.” It is the insistence that different forms of visibility should be named separately.

Mention rate tells a team whether it appears. Recommendation rate estimates whether it is prominent enough to enter consideration. Share of voice shows how much of the category conversation it captures relative to competitors. Source analysis indicates where the evidence behind those answers comes from.

Search Console can then provide a grounded starting set of commercial intents, while repeated tests across ChatGPT, Gemini and Perplexity expose where traditional search visibility does and does not transfer.

The result is a more useful question than “Do we rank in AI?” There is no single AI ranking to report. The better questions are how often the brand appears, how often it is recommended, who appears instead, which sources support those choices and whether any of that visibility eventually contributes to business.

0%