A user asks ChatGPT one question. Behind that apparently simple interaction, the system may conduct a sequence of searches the user never sees.
A new dataset gives marketers an unusually detailed look at that hidden retrieval layer. SEO consultant MJ Cachón ran 189 branded prompts across three client projects and captured 1,797 sub-queries generated by ChatGPT. Greg Jarboe, writing in Search Engine Journal, argues that the pattern resembles an SEO technique he used more than two decades ago: building content around phrases nested inside longer phrases so it can remain relevant as search intent becomes more specific.
The modern data adds a twist Jarboe could not have observed in 2003. Cachón found that ChatGPT's search sequences often begin with ordinary language, then narrow toward specific domains with the site: operator and, in longer sequences, increasingly use quotation marks to verify literal wording. Comparing the first search in a run with the last, quote usage increased 25-fold.
That finding is strategically interesting, but it is not evidence that exact-match sentences are a ChatGPT ranking factor. The study covers three brands, three sectors, one model and one period in August 2026. It observes what ChatGPT searched for, not which sources ultimately earned citations or how a universal ranking system works.
189 prompts generated 1,797 searches
Cachón's original query fan-out study tested 189 branded prompts across healthcare, sports retail and hotel projects.
Because fan-out is probabilistic, each prompt was run multiple times. The experiment produced 723 total runs using GPT-5.5 in August 2026.
Of the 189 prompts, 174 produced fan-out at least once. Across all runs, 687 contained fan-out and 36 completed without a search.
The resulting dataset contained 1,797 machine-generated sub-queries that no human user had directly typed.
The invisible search layer can be much larger than the visible prompt
Across repeated runs, each branded prompt generated an average of 10.3 distinct sub-queries.
Within an individual run, the average was 2.6 searches, with a maximum of 13.
This distinction matters. A single ChatGPT session may expose only a small portion of the retrieval possibilities associated with a prompt.
Cachón found that repeating the same prompt revealed additional sub-query variations, which is why one-off AI visibility tests can provide an incomplete picture of how a model may research a brand.
Not every prompt triggers fan-out
The headline number can create the impression that every ChatGPT prompt automatically explodes into ten searches.
That is not what the dataset shows.
Fifteen of the 189 prompts produced no fan-out at all, and 36 of the 723 runs completed without any captured search. Among runs that did fan out, many used only one or two searches.
Fan-out is a capability the system uses when it needs retrieval, not a fixed number of searches attached to every user message.
Most search sequences are short
Cachón's data is particularly useful because it shows how often the dramatic multi-stage sequences actually occur.
Of the 687 runs with fan-out, 246 made one search and 193 made two.
That means 63.9% of fan-out runs were resolved in two searches or fewer.
The elaborate sequences that progress through domain restrictions and exact quotation checks are therefore informative, but they are not the default behavior in every run.
The average sub-query was seven words long
The generated searches were not dominated by traditional short head terms.
Cachón measured an average sub-query length of seven words, with project averages ranging from 6.7 to 8.0 words.
Only 7.8% of the sub-queries were three words or shorter, while 15.8% contained ten words or more.
AI retrieval can therefore create a substantial long-tail search layer even when marketers never see those exact phrases in Search Console or keyword-planning tools.
The queries do not simply get longer as the model searches
One nuance is important when describing the sequence.
The original study says the searches tend to become more specific as the run progresses, but they can actually become shorter.
The system begins by describing what it needs in relatively natural language. Later it can replace that descriptive phrasing with operators, exact strings, names, identifiers or concise verification searches.
The funnel therefore moves from broad exploration toward precision, not necessarily from short queries toward progressively longer ones.
The first search tends to sound natural
Cachón sorted the 1,797 sub-queries by their position within each run.
The first search typically resembled natural language without specialized operators or quotation marks.
At this stage, the model appears to be discovering where useful information may exist.
This looks relatively similar to the way a human researcher might begin with a broad question before deciding which sources deserve closer inspection.
Then the model starts narrowing with site:
From the second search onward, Cachón observed more frequent use of the site: operator.
Across the full dataset, 30.2% of sub-queries contained site:.
Most searches still went to an open results set: 69.8% contained no domain restriction. But when the model deliberately restricted a search, it often focused on digital properties connected with the brand.
The sequence suggests a move from discovering candidate information toward investigating a source more directly.
One in six sub-queries targeted the brand's own domain
Of all 1,797 sub-queries, 299 used site: against the brand's primary domain, representing 16.6% of the dataset.
Another 142 targeted domains belonging to the same corporate group, while 100 targeted third-party sites.
Among all site: searches, 55.1% pointed at the brand's main domain.
This gives owned web properties an important role in branded AI retrieval, although it does not guarantee that the owned page will become the source cited in the final answer.
Exact quotation marks appear during verification
The most distinctive pattern emerges later in the retrieval sequence.
Cachón found quotation marks in 23.8% of all sub-queries, representing 427 searches.
The model appeared to use literal strings after encountering information it wanted to verify, searching for names, job titles, codes, document references and exact phrases.
In the study's interpretation, ChatGPT was moving from finding information toward checking whether a source actually contained the wording or fact it expected.
Quote usage increased 25-fold from the first search to the last
When Cachón compared only the first and last searches of each sequence, the escalation became much clearer.
Use of site: multiplied by four.
Use of quotation marks multiplied by 25.
This does not mean every final query contains quotes. It means quotation-based verification was dramatically more common at the end of the observed sequences than at the beginning.
The pattern can be summarized as ask, narrow, verify, insist
Cachón describes four broad stages in the longer fan-out journeys.
The model begins with an open question in natural language. It then narrows toward a particular domain, moves into verification with quoted phrases and can finish with shorter, increasingly specific reformulations if earlier searches have not resolved the issue.
One anonymized example in the study moved from the brand's site to exact phrases, then to a third party, back toward the brand and eventually to a literal sentence from a press release.
That sequence is the part of the dataset that caught Jarboe's attention.
Jarboe saw an old SEO idea hiding inside a new AI behavior
Jarboe calls his historical approach “Russian nesting dolls.”
When optimizing press releases in the early 2000s, he would look for a useful longer phrase containing a shorter phrase inside it. A four-word phrase might contain a three-word search expression, allowing the same passage to remain relevant to multiple variations.
The strategy was designed partly around emerging searches that did not yet have established keyword volume.
In Cachón's ChatGPT data, Jarboe sees the process operating in the opposite direction: the model starts broadly and progressively narrows until it is sometimes searching for literal language.
Nested phrases were originally a long-tail strategy
Jarboe's historical logic was simple.
If a useful four- or five-word phrase naturally contains a shorter core phrase, writing the longer expression can cover both meanings without producing separate content for each variation.
That mattered for news and public relations because new events generate language that keyword tools cannot predict in advance.
A press release published when the event happens can contain terminology that later becomes the language users search for.
The old tactic should not be confused with exact-match keyword stuffing
There is an important difference between writing a complete, natural phrase and mechanically repeating exact strings.
Jarboe's strategic interpretation is that content should contain meaningful language at multiple levels of specificity.
It does not follow that publishers should insert every imaginable fan-out query into headings or body copy.
Doing so would turn an observation about retrieval behavior into a keyword-stuffing exercise unsupported by the study.
Google independently confirms that query fan-out is real
Cachón studied ChatGPT, but query fan-out is not unique to OpenAI's product.
Google's Search Central documentation says AI Overviews and AI Mode may use query fan-out by issuing multiple related searches across subtopics and data sources while constructing a response.
Google describes the technique as a way to identify a broader and more diverse set of supporting web pages than a traditional single search might retrieve.
This confirms the broader architectural shift: one visible user question can correspond to many machine-generated retrieval queries.
Google's version can search several subtopics concurrently
Google's description emphasizes parallelism.
The model identifies related information needs, generates several searches and retrieves relevant results from Google's index to support the response.
That differs from assuming the visible prompt itself is the only keyword publishers need to understand.
AI search introduces a second demand layer generated by the system between the user's question and the documents ultimately considered for the answer.
Deep Search can take fan-out much further
Google has said its deeper research experiences can expand the same technique substantially.
When introducing Deep Search for AI Mode, Google explained that the system can issue hundreds of searches and reason across disparate information before producing a cited report.
That makes the concept strategically important beyond the specific 1,797 searches in Cachón's experiment.
As AI systems perform more research on users' behalf, publishers increasingly compete for retrieval across machine-generated queries they will never see directly.
But Google's guidance explicitly rejects exact-match obsession
This is where the strategic interpretation needs an important counterweight.
Google's current guidance for generative AI Search says publishers do not need to rewrite content in a special way for AI features.
Google says its systems can understand synonyms and broader meaning even when the page does not contain the exact words used in a query.
It specifically tells site owners not to worry about acquiring every long-tail variation of how somebody might search.
Google also warns against building pages for every fan-out query
Once marketers learn that AI systems generate hidden searches, an obvious temptation is to manufacture a page for every discovered sub-query.
Google explicitly warns against that approach.
Its AI optimization guide says creating separate content for every possible search variation, including fan-out queries, primarily to manipulate rankings or generative AI responses can violate its scaled content abuse policy.
The useful lesson from fan-out is broader coverage of genuine user needs, not industrial-scale exact-query page generation.
Jarboe's “quotable sentence” is a hypothesis, not a ranking rule
Jarboe recommends writing important answers as clear standalone sentences that can be quoted without losing their meaning.
The idea follows logically from Cachón's observation that later-stage searches sometimes use exact phrases for verification.
Clear factual writing is also useful to human readers.
But the dataset does not demonstrate that adding a specially engineered quotable sentence causes ChatGPT to rank, cite or mention a page more often.
The study observed retrieval, not citations
This distinction is one of Cachón's own stated limitations.
The dataset captures what the model searched for.
It does not establish which retrieved pages were cited in the final answer, how much influence each search had on the response or whether a particular query formulation improved a site's probability of citation.
A page can be retrieved and never cited, just as a traditional search result can be viewed by a system without becoming the final source a user sees.
Retrieval visibility and citation visibility are different metrics
AI visibility platforms increasingly report citations, mentions and source appearances.
Cachón's work highlights another layer beneath those outputs: retrieval opportunity.
A brand might not appear in the final citation set because its pages were never retrieved, because they were retrieved but not considered useful enough, or because another source provided stronger verification.
Understanding those stages is essential before treating a citation count as a complete explanation of AI visibility.
The model frequently searched for the brand itself
Branded prompts remained strongly centered on the entity under investigation.
Cachón found that 93.4% of all sub-queries contained the brand name.
Only 0.8% named a competitor.
That makes this dataset very different from generic product-discovery research and limits how confidently its patterns can be generalized to prompts such as “best CRM for a startup” or “which running shoes should I buy?”
Reputation became part of the hidden search process
About 10.1% of the sub-queries were categorized around reputation and trust.
The system searched for reviews, complaints, ratings, incidents and other external validation signals.
This means a branded prompt can trigger something resembling a lightweight reputation audit even when the user did not explicitly ask for every underlying concern.
For companies, AI visibility therefore extends beyond the pages they publish themselves into the broader web evidence surrounding the brand.
Third-party sources enter when owned information is insufficient
Among domain-restricted searches, Cachón observed queries aimed at registries, review platforms, social networks, media organizations and industry databases.
Those sources can help the model verify claims the brand cannot credibly establish by assertion alone.
A company saying it has won an award is useful evidence; an independent organization documenting the award may be stronger verification.
This is one reason AI search optimization cannot be reduced to editing the corporate website.
Corporate pages can become retrieval assets
The study found searches for information that traditional SEO teams sometimes treat as secondary: legal details, company identity, leadership, values, sustainability, certifications and official social profiles.
For branded AI research, these facts can become retrieval targets.
A clearly structured About page or legal notice may not generate large volumes of traditional organic traffic, but it can help establish factual information a model is trying to verify.
That is a different definition of content value from ranking for a high-volume keyword.
Literal HTML can be easier to retrieve than information trapped in assets
Cachón recommends placing key facts in indexable HTML rather than leaving them only inside images or PDFs.
The reasoning is practical: if a model generates a search for a precise fact, the web needs to expose that information in a form search infrastructure can reliably discover and interpret.
This does not mean PDFs or visual assets are inherently invisible, but important corporate facts benefit from accessible, crawlable presentation.
The recommendation also aligns with conventional technical SEO and accessibility principles.
Consistency matters when the system searches literal claims
Exact-phrase verification raises an interesting brand-governance issue.
If a company's slogan, executive title, warranty language or sustainability claim differs across its website, press releases, social profiles and group domains, an AI system may encounter conflicting evidence.
Cachón argues that consistent wording can make those facts easier to resolve.
The goal should be factual coherence rather than artificially repeating a phrase everywhere for optimization.
Freshness is another recurring retrieval signal
Nearly one in five sub-queries in the study contained an explicit year.
That indicates the model often looked for time-specific evidence when researching a brand.
Pages with visible dates, current reports and clearly maintained factual information can therefore be easier to interpret when freshness matters.
Again, the dataset shows a search pattern rather than proving that adding the current year to a page produces higher AI rankings.
The study captured one model at one moment
Cachón is unusually explicit about the limits of the experiment.
The tests used GPT-5.5 in August 2026, immediately before the GPT-5.6 release.
AI retrieval systems change quickly, and different model versions can alter when searches happen, how queries are formulated and which sources are selected.
A tactic inferred from one snapshot should therefore be monitored rather than frozen into a permanent SEO rulebook.
The official API may not reproduce every consumer ChatGPT behavior
The data was extracted through OpenAI's official API.
Cachón notes that this capture method may differ from what users experience in the consumer ChatGPT product.
That does not invalidate the observed pattern within the experiment, but it limits claims about universal ChatGPT behavior.
Researchers should distinguish between the tested environment and every other OpenAI interface or product configuration.
Three brands cannot represent the entire market
The study covers three projects in three sectors.
The patterns repeated across those projects, but the sample is not representative of all industries, languages, markets or prompt types.
A multinational financial brand, a local restaurant and an open-source software project may trigger very different retrieval behavior.
Cachón explicitly says the study needs to be repeated and expanded.
Branded prompts behave differently from generic discovery prompts
Every prompt in the dataset was branded.
That naturally encourages the model to investigate the company's own site, corporate group and reputation footprint.
A generic prompt has no predetermined entity around which to organize those searches.
The 30.2% site: rate and 93.4% brand-name rate should therefore not be copied into forecasts for all AI queries.
The practical opportunity is semantic completeness
The strongest content lesson is not “repeat exact phrases.”
It is to make important topics, claims and entity facts clear enough that both humans and retrieval systems can understand them at several levels of specificity.
A strong page can explain the broad concept, define the important terms, state the critical facts plainly and provide evidence for deeper questions without spawning dozens of near-duplicate pages.
That kind of semantic completeness is useful whether the visitor arrives from classic Search, an AI citation or a direct link.
Think in question families rather than one target keyword
Traditional keyword planning often selects a primary query and a collection of secondary variants.
Fan-out suggests a more useful mental model: what family of questions might a researcher need to resolve before answering the original prompt confidently?
For a hotel brand, that family might include location, amenities, policies, reputation, accessibility, cancellation terms and current offers.
Content architecture can address those real information needs without trying to predict the exact machine-generated wording of every hidden search.
Use AI fan-out research diagnostically
When tools expose sub-queries, marketers can use them to discover missing information.
If a model repeatedly searches third-party sites for a fact that should be authoritative on the company's own website, that may reveal a genuine content gap.
If it repeatedly reformulates the same question, the available information may be ambiguous, inaccessible or insufficiently specific.
The useful action is to improve the underlying information environment, not simply paste the generated queries into the page.
Repeated searching may be a sign of uncertainty
Cachón found that 23.1% of runs with at least two searches contained sub-queries that were identical or only slightly different.
She interprets this as a sign that the model can keep trying alternative formulations when earlier retrieval has not resolved the question.
A long fan-out sequence may therefore indicate difficulty rather than greater importance.
That creates a useful monitoring hypothesis: repeated searches around the same brand fact may identify areas where the web evidence is weak or contradictory.
Query fan-out changes what “search demand” means
Keyword tools traditionally estimate demand from phrases humans type.
AI systems introduce another category: machine-generated retrieval demand.
These sub-queries may never appear in Search Console because no human entered them as a conventional Google search associated with the publisher's property.
Yet they can influence which documents an AI system considers while constructing an answer.
Invisible demand does not mean infinite content demand
The existence of thousands of hidden queries does not imply publishers need thousands of new pages.
One well-designed resource can answer many related questions when its structure, facts and context are clear.
Google's explicit warning against building content for every fan-out variation is useful here even though Cachón's experiment studied ChatGPT rather than Google.
AI retrieval increases the diversity of queries a page may need to satisfy, but modern semantic systems can also connect relevant content without exact lexical matches.
The old nesting-doll idea survives best as an editorial principle
Jarboe's 2003 framework remains interesting because it encourages writers to use natural, specific language that contains broader concepts inside it.
A precise sentence often does more work than a vague sentence because it can answer both the general and detailed question.
That is useful editorial practice independently of AI.
Where the idea becomes risky is when marketers turn it into a claim that exact wording mechanically controls retrieval or citation.
AI search is creating a funnel that marketers cannot see directly
Cachón's 1,797 sub-queries make one point difficult to ignore: the visible prompt is only the entrance to the retrieval process.
In the tested branded searches, ChatGPT frequently moved from broad discovery to domain-specific investigation and, when more verification was needed, toward literal phrases. Quote usage was 25 times higher in final searches than in first searches, while site: usage increased fourfold.
Jarboe sees that pattern as a modern version of his nested-phrase strategy. It is a useful way to think about writing clear, specific, quotable information, but it remains a strategic interpretation rather than a proven ranking formula.
The deeper lesson is more durable. AI systems do not merely answer questions; they can generate their own questions in order to answer ours. SEO strategy now has to account for that invisible research layer without chasing every machine-generated phrase. The opportunity is to make the underlying information so clear, complete and verifiable that whether the system searches broadly, narrows to a domain or checks a literal claim, the page still has something worth retrieving.