There may be no meaningful single number called “AI visibility.” New measurement data from AI visibility platform Treyci shows why: in one B2B software category, the same tracked brands appeared in 81% of answers from one major AI engine but only 43% from another.
The nearly two-fold gap emerged from more than 1,200 scored answers generated across ChatGPT, Gemini, Perplexity and Grok. Treyci's methodology runs roughly 100 buying-intent prompts per category, asks each engine the same questions three separate times and records which brands are mentioned, recommended and cited.
The September 3 release does not identify which engine produced the 81% figure or which produced 43%, so those endpoints should not be attributed to ChatGPT, Gemini, Perplexity or Grok individually. What the published data does establish is the size of the cross-engine gap in that measured category.
For marketers, the result challenges two increasingly common habits: checking only ChatGPT and treating one generated answer as if it were a stable search ranking.
One brand can occupy four different AI markets
Traditional search trained marketers to think about visibility inside relatively coherent ecosystems. Rankings could vary by location, device and personalization, but a Google ranking was still fundamentally a measurement inside Google.
Generative discovery is more fragmented. ChatGPT, Gemini, Perplexity and Grok do not simply present four interfaces over one shared recommendation index. They use different models, retrieval systems, source graphs, ranking logic and answer-generation behavior.
Treyci's reported 81%-to-43% range makes that fragmentation measurable. A brand can appear highly visible when one engine is used as the benchmark and dramatically weaker when another engine receives the same buying questions.
The practical consequence is that an “AI visibility score” averaged across platforms can conceal the most important information: where the brand is actually winning and where it is effectively absent.
The study uses buying questions rather than branded prompts
Treyci says its measurement begins with approximately 100 buying-intent questions for a category. Examples include “best X for mid-size teams,” alternatives queries and pricing comparisons.
This is a stronger test of commercial visibility than asking an assistant directly about a known company. A branded prompt such as “What is Brand X?” tests whether the system can recognize and describe an entity that the user has already selected.
A category-level buying question forces the engine to construct the shortlist itself. The brand must compete for inclusion against alternatives without being named in the prompt.
That distinction is central to AEO measurement. Recognition answers “does the model know us?” Recommendation testing asks “does the model bring us into the conversation when the buyer has not already chosen us?”
Every prompt is run three times
The second important feature of the methodology is repetition. According to Treyci's published methodology, every question is run three separate times on each engine during the monthly measurement cycle.
With roughly 100 prompts across four engines, three runs produce more than 1,200 answers to analyze. Treyci then scores brand mentions, recommendations and citations and stores the underlying responses.
The three-run design addresses a problem that conventional ranking tools rarely face to the same degree: generative answers are probabilistic. Asking the same engine the same question again can produce a different vendor list.
A single screenshot is therefore an observation, not a reliable estimate of a brand's probability of appearing.
The engines disagree with themselves as well as each other
Treyci reports that repeated sessions routinely return different vendors for the same buying question. Its own product illustrates this as “run consistency” — for example, a brand might be recommended in two of three executions rather than all three.
This introduces two separate dimensions of instability. Cross-engine variance measures how differently ChatGPT, Gemini, Perplexity and Grok construct the market. Within-engine variance measures how consistently one of those systems reproduces its own answer.
A brand can therefore have high average visibility on an engine while still being unstable at the individual-prompt level. Another brand might appear less often overall but be extremely consistent whenever a particular use case is asked.
Those patterns require different interpretations. Aggregate visibility tells a marketer how broad the presence is. Run consistency tells the marketer how dependable that presence appears to be.
81% versus 43% is not a universal ranking of the four engines
The headline number needs an important limitation. Treyci says the result came from one measured B2B software category. The public announcement does not disclose enough underlying category-level data to treat the figures as universal properties of the engines.
It would therefore be incorrect to conclude that one unnamed platform always mentions brands 81% of the time while another always mentions them 43% of the time.
A different category could produce a different ordering. Consumer electronics, insurance, travel and local services have different source ecosystems, commercial language and information structures.
The useful conclusion is narrower: large platform-level differences can exist even when the prompt set, brands and measurement period are held constant.
The source is vendor research, not an independent benchmark
The findings also come from Treyci, a company selling AI visibility measurement software. The announcement was distributed by Helix Apps LLC, the company behind the platform.
That commercial context does not invalidate the methodology, but it should shape how the evidence is described. The public release provides the measurement design and headline findings but not the complete 1,200-plus-answer dataset needed for independent reproduction of the 81%-to-43% result.
The numbers should therefore be attributed to Treyci rather than presented as settled industry benchmarks.
The methodology itself is easier to evaluate independently: identical buying questions, multiple engines, repeated runs and answer-level scoring are all sensible responses to the known instability of generative systems.
NetContentSEO has observed the same cross-model problem at smaller scale
The pattern is consistent with experiments we have run on entity recognition. In our test asking six AI models what Net Content SEO is, ChatGPT, Gemini and Grok reconstructed the project reasonably accurately while Perplexity, Gemma and Llama either failed to identify it reliably or drifted toward generic interpretations.
The underlying entity did not change between prompts. The web did not suddenly contain a different NetContentSEO for each model. The difference came from how each system retrieved and reconstructed the available information.
Treyci's dataset moves the same issue into commercial recommendation at a much larger scale. The question is no longer merely whether systems understand the same brand differently, but whether they include the brand in the buyer's shortlist at all.
That makes cross-engine divergence a competitive measurement problem rather than only an interesting model-behavior problem.
Mentions, recommendations and citations should remain separate
Treyci scores three outcomes: whether a brand is mentioned, whether it is recommended and whether the answer cites a supporting source.
These should not be collapsed into one concept. A mention can place a brand among a long list of alternatives. A recommendation can frame it as particularly suitable for the user's requirements. A citation can show which external document contributed evidence to the answer.
NetContentSEO has seen a similar separation in insurance data. Our analysis of 20 Australian insurance brands found that raw visibility and AI preference did not produce the same leaderboard. A brand could appear frequently without being the strongest recommendation for a buyer segment.
AI measurement is developing the same funnel distinction that digital marketing already applies elsewhere: impressions are not clicks, and mentions are not recommendations.
Third-party sources may explain part of the divergence
Treyci also reports that review platforms, comparison articles and industry publications dominate citations around buying questions more than vendor websites do.
If engines retrieve different third-party sources, their vendor shortlists can diverge even when the brands' own websites are identical across platforms.
One system may rely heavily on a comparison page that includes five vendors. Another may retrieve a review directory containing a different competitive set. A third may generate more of the answer from model knowledge and use retrieval only selectively.
This means AI visibility cannot be optimized exclusively on the company domain. The external information environment — reviews, category pages, editorial comparisons, credible news coverage and other independent references — can influence which version of the market an engine reconstructs.
A brand-by-engine matrix is more useful than one AI score
For reporting, the simplest response to the 81%-to-43% gap is to stop hiding engine-level results.
A useful dashboard should show brands down one axis and engines across the other, with mention rate, recommendation rate and consistency visible separately. The combined score can still exist for executives who need a summary, but the matrix should remain available for diagnosis.
If a brand has 70% visibility on three engines and 20% on the fourth, a 57.5% average does not describe the strategic problem. The fourth engine is the problem.
Likewise, if a competitor dominates only Perplexity while another dominates ChatGPT, the company may need to investigate different source ecosystems rather than launch one generic “GEO” initiative.
Repeat runs should be treated as samples, not duplicates
Running the same prompt three times may look wasteful to teams accustomed to deterministic rank checking. In generative systems, the repetition is part of the measurement.
If a brand appears once in three runs, the useful result is not that the “correct” answer was the run where it appeared. The useful result is that the observed presence was unstable under identical prompting.
Three runs still form a small sample. For strategically important prompts, more repetitions or repeated measurement over time can produce a stronger estimate of the underlying distribution.
The general rule is that measurement confidence should increase with the importance of the decision. A casual monthly diagnostic may tolerate three runs; a claim that a multimillion-dollar brand campaign changed AI recommendations deserves stronger evidence.
Prompt sets need to remain stable over time
Cross-engine measurement becomes even more useful when the same prompt suite is repeated monthly. But longitudinal comparisons only work if teams control changes to the test itself.
If half the buying questions are replaced between August and September, an increase in visibility can be caused by easier prompts rather than a genuine improvement in how the engines represent the brand.
Core prompts should therefore be versioned and kept stable for trend measurement. New questions can be added to an exploratory set without rewriting the historical benchmark.
Model and product changes should also be documented when possible. An engine update can alter recommendation behavior even when the brand, its website and the surrounding source ecosystem remain unchanged.
The absence of an impression is a measurement blind spot
Traditional web analytics begin after a user reaches a property the company can observe. AI recommendations can influence the decision before that happens.
If an assistant recommends three competitors and never mentions the fourth brand, the excluded company receives no impression in Search Console, no website session and no lost-opportunity event in its analytics.
This is why synthetic prompt measurement has become attractive despite its limitations. It attempts to observe a part of the buyer journey that first-party analytics cannot see directly.
The tradeoff is that the resulting dataset is constructed rather than naturally occurring. The quality of the measurement depends heavily on whether the prompt set genuinely represents customer research behavior.
llms.txt adoption does not prove visibility improvement
Treyci adds another useful observation from a scan of 100 B2B SaaS companies: 41 had published an llms.txt file, while the company says relatively few could quantify whether their AI visibility efforts changed recommendation frequency.
The point is not that llms.txt is useless. It is that implementing an AI-oriented technical artifact and measuring a commercial outcome are separate tasks.
A company can publish a crawler-facing file, add structured content or rewrite product pages and still need a controlled before-and-after measurement to determine whether anything changed in actual answers.
AEO should be held to the same discipline expected from other marketing channels: implementation is not evidence of impact.
AI visibility is becoming a distribution, not a position
The deeper implication of Treyci's findings is conceptual. Traditional SEO reporting encourages a positional model: a page ranks third, fifth or tenth.
Generative recommendation behaves more like a probability distribution. A brand may appear in two of three runs, on three of four engines, strongly for comparison prompts but weakly for alternatives prompts, and frequently as a mention without often becoming the recommendation.
Trying to compress that behavior into “we rank number two in AI” removes most of the information that makes the measurement useful.
The better vocabulary is probabilistic: mention rate, recommendation rate, run consistency, engine coverage and competitive share across a controlled prompt set.
The same question does not produce one AI market
The 81%-versus-43% result is useful not because it identifies a winning answer engine — the public data does not even name the two endpoint platforms — but because it demonstrates how misleading a single-engine view can become.
ChatGPT, Gemini, Perplexity and Grok are separate discovery environments. They can retrieve different evidence, construct different shortlists and change their own answers between repeated sessions.
For brands, that means AI visibility cannot be established by asking ChatGPT a few questions and saving favorable screenshots. Measurement needs the same prompts across multiple engines, repeated enough to expose internal variability, with mentions, recommendations and citations tracked separately.
Search rankings taught marketers to monitor positions. Answer engines require them to monitor distributions. A brand that looks dominant in one system and nearly invisible in another does not have one AI visibility score. It has four different competitive realities.