AI Search Removes Up to 60% of Uncertainty Language—While Keeping the Confident Tone

AI Search Removes Up to 60% of Uncertainty Language—While Keeping the Confident Tone
Sponsored

Adding web search to an AI answer can make the response sound less uncertain without making its language proportionally more confident in an absolute sense. That subtle shift matters because users may experience the resulting answer as more authoritative even when the system is still synthesizing incomplete, selective or contested evidence.

A new cross-system study calls this broader phenomenon an “answer bubble.” The paper, originally submitted to arXiv on March 17, 2026 and revised in August, examines 11,000 real Google search queries across traditional Google results, Google AI Overviews, GPT without search and GPT with web search. The updated version also includes a Perplexity Search robustness check using Grok.

The researchers find that generative search systems differ not only in which websites they cite but in how they transform those sources into a single authoritative-sounding answer. Search GPT's 100 most-cited domains overlapped only 25% with traditional Google Search and 24% with Google AI Overviews. Search grounding also reduced tentative language sharply: hedging fell 40% when GPT gained web search and 60% when comparing vanilla GPT with Google AI Overviews.

The crucial detail is that certainty markers declined more slowly than hedging. The ratio of certainty language to tentative language increased from 0.39 in vanilla GPT to 0.49 in both search-grounded systems. In the authors' interpretation, search does not simply compress all epistemic language. It selectively removes more of the linguistic signals that tell a reader the system is deliberating or uncertain.

The study starts with 11,000 real user search queries

The researchers—Michelle Huang, Agam Goyal, Koustuv Saha and Eshwar Chandrasekharan of the University of Illinois Urbana-Champaign—sampled queries from Google's Natural Questions corpus, which contains real information-seeking searches submitted by users.

They first classified 50,000 sampled queries into 11 topical categories and then selected 1,000 from each category, creating a balanced dataset of 11,000 searches spanning business and finance, education, entertainment, history, music, politics and news, religion, science, sports, technology, and travel or geography.

For each query, the main experiment collected a vanilla GPT-4o-mini response without web access, a GPT-4o-mini response using OpenAI's web_search tool and a Google AI Overview when Google produced one. Traditional Google organic results provided the link-based search baseline for source comparisons. The paper additionally tests Perplexity Search with Grok 4.1 Fast as a cross-system robustness check.

This design lets the researchers separate several questions that are often collapsed into one. Which sources does each system expose? How does search grounding change the language of the generated answer? And after sources are cited, how much of their content actually makes it into the final synthesis?

Search GPT and Google use surprisingly different source ecosystems

Traditional Google Search returns roughly nine to ten results per query in the dataset. Search GPT cites a median of three sources and provides at least one citation for 97% of queries. Google AI Overviews are more variable: they appeared for 57.8% of the study's queries and contained an average of seven references when present.

At the domain level, the divergence is substantial. Among each system's 100 most frequently cited or surfaced domains, Search GPT overlaps only 25% with traditional Google Organic and 24% with Google AI Overviews. By comparison, Google Organic and AI Overviews share 68% of their top 100 domains.

These are top-domain overlap statistics, not per-query citation overlap. They describe the systems' aggregate source preferences across the benchmark rather than saying that only one quarter of the URLs for any individual search are shared.

The pattern still has an important GEO implication. Search-grounded GPT does not appear to be simply reproducing Google's dominant source mix. A site that is highly visible in conventional Google results may occupy a very different position in a web-search-enabled assistant, while sources favored by the assistant can gain exposure without matching Google's aggregate ranking profile.

Wikipedia dominates all three source environments

Wikipedia is the most visible domain across the systems the researchers compare. It appears in 81% of traditional Google queries, 49% of Search GPT responses and 28% of Google AI Overviews.

Search GPT also shows a strong preference for encyclopedic and editorially curated sources. Encyclopedic and reference domains account for 27.3% of its citations, compared with roughly 10% in the two Google systems. News and media sources account for 7.2% of Search GPT citations, versus 2.4% in AI Overviews and 1.7% in organic Google.

Social sources move in the opposite direction. Search GPT draws only 0.1% of citations from social platforms, while AI Overviews draw 8.5% and Google Organic results 13.4%. Facebook, Quora and Reddit appear prominently in Google's systems but much less in Search GPT's source mix.

The researchers describe these differences as distinct source-selection philosophies. Search GPT appears more concentrated around encyclopedic and editorial sources, while Google's systems expose a broader range of community and user-generated material.

Source selection is only the first filter

Being cited does not mean a source contributes equally to the generated answer. That distinction is central to the paper's “answer bubble” concept.

A traditional search engine gives users a ranked list and leaves them to decide which pages to open. A generative system makes two editorial decisions on the user's behalf. First, it chooses the sources. Then it chooses which pieces of those sources deserve to survive into the synthesis.

The second filter can amplify the first. A system may cite a diverse group of pages while building most of the actual answer from only a subset. Citation diversity can therefore overstate information diversity if the final text disproportionately reflects a few sources.

For publishers, this creates an important measurement distinction between citation presence and substantive influence. A URL can appear in the references while contributing little identifiable information to the answer itself.

Web search reduced hedging by 40% in GPT

The researchers use Linguistic Inquiry and Word Count categories and other linguistic measures to compare the generated responses. One of the clearest differences involves tentative language: words and constructions such as “maybe” and “perhaps” that signal epistemic caution.

Vanilla GPT produced a tentative-language score of 0.025. With web search, Search GPT fell to 0.015, a 40% reduction. Google AI Overviews scored 0.010, 60% below vanilla GPT.

Other cognitive language also declined, but less rapidly. Overall cognitive-process language fell 23% in Search GPT and 49% in Google AI Overviews compared with vanilla GPT. The disproportionate decline in hedging is what makes the finding notable.

Search grounding did not merely make answers shorter or stylistically cleaner. It changed the relative presence of linguistic cues that communicate uncertainty.

The headline needs one important qualification: certainty language also fell

It would be inaccurate to summarize the study as “web search removes doubt but keeps every confidence marker unchanged.” Absolute certainty language also declined. The paper reports certainty markers falling from 0.010 in vanilla GPT to 0.007 in Search GPT and 0.005 in Google AI Overviews.

But hedging fell faster. As a result, the certainty-to-tentative ratio increased from 0.39 in vanilla GPT to 0.49 in both search-grounded systems. The relative balance shifted toward assertiveness.

This distinction makes the finding more interesting, not less. The systems are not simply injecting words such as “definitely” or “always” into web-grounded answers. Instead, they are removing caution faster than they remove certainty.

For a reader, that can create a more definitive surface even when the underlying evidence still contains ambiguity. The study characterizes the language of that transformation; it does not establish that every more assertive answer is factually wrong or unjustifiably confident.

Less hedging is not the same thing as more hallucination

This is where the research should not be overextended. The paper measures epistemic language and source-summary coverage. It does not run a universal factuality benchmark proving that search-grounded answers with less hedging are less accurate.

Web retrieval can improve factual accuracy, particularly for current information that a model's parametric knowledge does not contain. A system may sometimes be justified in becoming less tentative after finding strong evidence.

The concern arises because users cannot easily observe whether reduced uncertainty reflects stronger evidence or simply a stylistic consequence of the synthesis process. A confident tone is not itself a reliability metric.

For high-stakes information, the ideal behavior would be calibrated confidence: strong claims when evidence is strong, explicit uncertainty when sources conflict or remain incomplete. The study shows that search grounding systematically changes those signals, making calibration worth measuring directly.

Google AI Overviews and Search GPT have different writing styles

The linguistic analysis also shows that “AI search” is not one homogeneous writing style. Search GPT responses averaged 140 words, compared with 83 for Google AI Overviews and 87 for vanilla GPT.

Search GPT used longer sentences, longer words and more formal language than Google AIO. AI Overviews favored shorter sentences, simpler vocabulary and a compressed style suited to an information panel embedded in a search page.

These product choices affect how sources are experienced. A longer assistant response has more space to contextualize evidence, explain alternatives or incorporate multiple source perspectives. A compressed overview has stronger pressure to decide what can be omitted.

The paper does not argue that longer is inherently more accurate. It shows that interface and synthesis style are part of information exposure, not merely cosmetic differences.

The uncertainty effect changes by topic

The reduction in epistemic markers is not uniform across the benchmark. Technology and travel or geography show some of the largest differences between search-grounded systems. In technology queries, Google AI Overviews used 64% less hedging than Search GPT in the study's topic-level comparison.

Education, history, and business or finance showed much smaller differences in hedging and broader cognitive language between the two search-grounded systems. These were also categories in which AI Overviews appeared relatively frequently.

This topical variation matters for audits. A global average can hide a system that behaves cautiously in one domain and tersely in another. Organizations monitoring AI representations should evaluate the query categories that matter to them rather than relying only on platform-wide language statistics.

Long sources received disproportionately more influence

For a subset of 1,100 queries per generative system, the researchers decomposed answers into atomic content units and compared those units with the full text of cited sources. This allowed them to estimate how strongly each source was represented in the generated summary.

Document length produced the largest coverage disparity. Sources shorter than 200 words were underrepresented by 13.4 percentage points in Search GPT and 17.7 points in Google AI Overviews. Sources longer than 800 words were overrepresented by 3.9 and 4.6 points respectively.

The authors caution that their max-over-chunks entailment method itself can be sensitive to document length, because a longer document provides more opportunities for a chunk to match an answer claim. The finding therefore should not be converted into a simplistic GEO recommendation to make every page longer.

It does, however, show that citation presence and synthesis weight interact with document structure. A source containing more extractable relevant material may have more opportunities to influence the generated answer than a short page cited for one narrow fact.

Wikipedia is not merely cited frequently—it is used disproportionately

Wikipedia exhibits what the researchers call a compounding effect. It is already the most frequently surfaced domain, and once cited, its content receives more representation in the final synthesis than an equal-coverage baseline would predict.

Wikipedia is overrepresented by 2.6 percentage points in Search GPT and 5.4 points in Google AI Overviews. Encyclopedic sources as a broader category are also overrepresented in both systems.

That matters because source concentration can happen twice. A domain can win at retrieval and then win again during generation. Its influence over the final answer becomes larger than its citation count alone suggests.

For GEO measurement, this creates a reason to distinguish “cited” from “used.” A domain appearing in a source panel is not necessarily shaping the answer as much as another domain cited beside it.

Social and forum sources can be cited without materially shaping the answer

The opposite pattern appears for social and discussion content. Search GPT cites almost no social sources to begin with, so the researchers could not make the same coverage comparison there. Google AI Overviews cite social and forum sources more frequently, but those sources are underrepresented in the generated summary by 22.1 percentage points.

This means an AIO can display community sources among its references while relying much less on their content when constructing the visible answer.

The distinction is particularly relevant for marketers monitoring Reddit, forums and social platforms as AI-search sources. Citation tracking may show that these communities matter to retrieval, but their actual contribution to the generated wording can be considerably smaller.

Social visibility and social influence are therefore separate metrics. A source can help the system discover a perspective without that perspective surviving the compression step.

Negative sources were also underrepresented

The synthesis analysis finds another asymmetry around source sentiment. Negatively framed sources were underrepresented by 3.7 percentage points in Search GPT and 13.8 points in Google AI Overviews, while positive sources showed small positive or neutral effects.

The authors present this as a descriptive coverage difference rather than proof of intentional positivity bias. A negative source could be redundant, less relevant to a particular answer or structured differently from other evidence.

Still, the aggregate pattern raises a practical reputation question. If critical sources are retrieved but their content is less likely to survive into the synthesis, an AI answer can present a different balance of information from the source set visible in its citations.

Brand-monitoring systems should therefore inspect the generated narrative, not assume that the presence of a critical source means its criticism was incorporated.

Google AIO cited more sources but covered each less deeply

In the source-summary analysis, Google AI Overviews cited an average of five sources compared with 2.7 for Search GPT. Yet their mean source-content coverage was lower: 0.447 versus 0.500.

On the subset of queries where both systems produced scoreable results, the coverage gap remained significant. The researchers interpret this as a tradeoff in which AI Overviews cite a broader set of sources but synthesize each one less thoroughly.

Again, broader citation diversity is not automatically broader informational representation. A compact answer can reference many sources while extracting only a small amount from each and weighting some far more than others.

That distinction is central to understanding answer bubbles. The visible source list tells users what was consulted or attributed; the generated prose determines what they actually learn.

An answer bubble is more than a filter bubble

The paper borrows conceptually from the idea of filter bubbles but describes a different mechanism. A filter bubble is typically associated with personalization that changes which information a person encounters. An answer bubble can exist even when two users submit the identical query without personalization.

The system itself selects a source pool, chooses which parts of those sources matter and renders the result in its own linguistic voice. Another system can make different decisions at every one of those stages.

Two people asking the same question to Google AIO and Search GPT can therefore receive structurally different information realities even when neither answer contains an obvious factual error. One may rely on community sources; another on encyclopedias and wire services. One may compress uncertainty more aggressively. One may cite more pages but use each less deeply.

The user sees a finished answer, not the sequence of editorial choices that produced it.

This has consequences for GEO beyond citation share

Most GEO measurement currently emphasizes whether a domain or brand is cited. The answer-bubble framework suggests at least three separate layers need to be measured.

The first is source selection: does the system retrieve or cite the domain at all? The second is source utilization: how much information from that source reaches the final response? The third is transformation: how does the system change the framing, uncertainty, sentiment or emphasis of the source material?

A publisher can perform well on the first layer and poorly on the second. A brand can be accurately sourced but represented with a tone that removes important qualifications. A social discussion can be cited without materially influencing the answer.

Visibility alone is therefore an incomplete GEO outcome. Representation fidelity matters too.

Publishers should write uncertainty explicitly when it matters

The study does not test content-optimization interventions, so it cannot prove that a particular writing technique will survive AI synthesis. But its findings suggest a sensible editorial principle: important qualifications should not be buried in vague prose.

If evidence is preliminary, say so explicitly. If estimates have a range, provide it. If experts disagree, identify the disagreement and its basis. If a result applies only to one population or time period, make the limitation concrete.

Generative systems inevitably compress source material. Clear factual qualifiers have a better chance of remaining machine-detectable than subtle rhetorical caution spread across several paragraphs.

This is good editorial practice even outside GEO. The goal is not to write for a model's extraction algorithm but to make the epistemic status of information legible to both humans and machines.

The 60% number describes language, not truth

The headline finding is powerful enough that it deserves careful wording. Google AI Overviews used 60% less tentative language than vanilla GPT under the study's LIWC measure. Search GPT used 40% less. The certainty-to-tentative ratio rose in both search-grounded systems because certainty markers declined more slowly.

Those results do not mean AI search is 60% more overconfident, 60% less accurate or 60% more likely to hallucinate. None of those quantities was measured by that statistic.

What the study establishes is a change in how uncertainty is communicated. Search-grounded answers can present fewer markers of deliberation and caution relative to markers of certainty. Whether that rhetorical shift is justified depends on the evidence behind each individual answer.

That is precisely why the finding matters. Users often infer reliability from presentation. If the interface collapses many sources into one fluent response, the tone becomes part of the information system.

AI search is becoming an editorial layer, not merely a retrieval layer

Traditional search already exercises enormous influence by deciding which links appear and in what order. Generative search adds a second form of power: it decides what the selected sources are allowed to say in the final answer.

The Answer Bubbles study shows these decisions are system-specific. Search GPT's dominant source domains differ sharply from Google's. Wikipedia and long documents can receive disproportionate influence during synthesis. Social and forum content can be visible in citations but faint in the answer. Epistemic language can become more assertive as hedging is compressed faster than certainty.

None of this proves that one system is universally better. It demonstrates that an AI answer is not a neutral window onto the retrieved web. It is a constructed information product shaped by source selection, synthesis rules, model behavior and interface constraints.

For SEO, GEO and publishers, the practical consequence is that ranking and citation tracking are only the beginning. The next measurement problem is understanding what happens after a source is selected: whether its information survives, how heavily it is weighted and whether the AI preserves the uncertainty that made the original source intellectually honest.

0%