ChatGPT Retrieved 63% of Expert-Selected Clinical Studies—Gemini Found Just 17%

ChatGPT Retrieved 63% of Expert-Selected Clinical Studies—Gemini Found Just 17%
Sponsored

Ask three leading AI assistants the same evidence-based medical question and the studies they retrieve can differ dramatically. A new preprint comparing ChatGPT GPT-5.5, Claude Sonnet 5 and Gemini 3.1 Pro against studies selected by expert Cochrane reviewers found that ChatGPT retrieved 63.1% of the included studies on average, compared with 37.0% for Claude and just 17.3% for Gemini.

The study, submitted to arXiv on August 13, 2026, was conducted by Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt and Zhiyong Lu. Rather than asking whether AI assistants fabricate citations, the researchers examined a different question: when a chatbot supplies real clinical studies, does it retrieve the same evidence that expert systematic reviewers considered relevant?

720 AI answers were benchmarked against Cochrane reviews

The researchers adapted clinical questions from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews. They then prompted each of the three chatbots while simulating three different user roles—a patient, a clinician and an evidence-synthesis researcher—and repeated each condition four times. The design produced 720 chatbot responses in total.

Each assistant was asked to support its answer with primary clinical citations. Those citations were then compared with both the included and excluded study sets assembled by the corresponding Cochrane reviews, giving the researchers an expert-curated benchmark against which to measure retrieval.

Across all models and roles, an individual chatbot response retrieved an average of 39.2% of the studies that Cochrane reviewers had included. The assistants also cited an average of 5.0% of studies that the reviews had excluded. The overall number therefore conceals substantial variation in which expert-selected evidence actually reached the generated answer.

The model mattered enormously

The largest difference was between the AI systems themselves. ChatGPT GPT-5.5 achieved mean recall of 63.1%, Claude Sonnet 5 reached 37.0% and Gemini 3.1 Pro reached 17.3%. The authors report that the model effect was statistically significant.

Recall in this context should not be confused with clinical accuracy. A higher score means that the assistant retrieved a larger proportion of the studies included by the relevant Cochrane review; it does not establish that every generated medical statement was correct, that the assistant interpreted every paper appropriately or that the answer would lead to a better clinical decision.

The result nevertheless matters because citations can create an impression of comprehensive evidence retrieval. Two assistants may both provide polished answers with legitimate-looking references while relying on markedly different slices of the available expert-selected literature. A user who sees several citations therefore cannot assume that the system has surfaced most of the important evidence.

How the user identified themselves also changed retrieval

The study found another significant effect: the role embedded in the prompt influenced which evidence the chatbots returned. When the user presented themselves as an evidence-synthesis researcher, average recall reached 42.8%. The clinician condition produced 38.6%, while patient-style prompts averaged 36.1%.

The gap is smaller than the difference between models, but it demonstrates that source retrieval is not determined solely by the medical question. How a person frames their identity and information need can affect the evidence an assistant chooses to surface, even when the underlying clinical topic remains comparable.

That finding has practical implications for anyone evaluating AI search systems. Benchmarking a medical assistant with a single prompt formulation may provide an incomplete picture because a change in user framing can alter retrieval. It also raises a user-experience question: patients and researchers asking about the same treatment may not be exposed to the same evidence base.

Larger clinical studies were more likely to be retrieved

The researchers also examined characteristics that might predict whether an eligible clinical study appeared in an AI response. After controlling for publication year, citations per year and open-access status, sample size was the only independently significant predictor they identified.

For every one-unit increase in log sample size, the odds of retrieval increased by a factor of 1.80, according to the paper. In simpler terms, the chatbots showed a measurable tendency to retrieve larger clinical trials rather than selecting studies independently of their size.

That bias is not automatically undesirable. Larger studies can provide greater statistical power and may often be influential. But systematic evidence synthesis does not determine relevance by sample size alone. Smaller studies can still be important for understanding specific populations, interventions, outcomes or gaps in the evidence. A retrieval system that systematically favors scale could therefore produce a different picture from an expert review process.

Citations are not the same as systematic evidence retrieval

The study highlights a broader problem as general-purpose AI assistants become research interfaces. Citation features can improve transparency by giving users somewhere to check a claim, but the presence of references answers only one question: what did the model cite? It does not reveal which relevant evidence was never retrieved.

Systematic reviews are deliberately designed to reduce that blind spot. Reviewers define eligibility criteria, search multiple databases, screen candidate studies and document inclusion and exclusion decisions. A conversational assistant optimizes for a different experience: producing a useful response quickly. The comparison in this paper shows that those processes can yield very different evidence sets.

This distinction is especially important in medicine, where omitted evidence can matter as much as an incorrect citation. If an assistant consistently retrieves only a subset of the studies experts consider relevant, users may receive an answer that looks well-supported while remaining incomplete.

The numbers apply to specific models and a specific test

The headline differences should not be treated as permanent rankings of ChatGPT, Claude and Gemini. The researchers evaluated particular model versions—GPT-5.5, Claude Sonnet 5 and Gemini 3.1 Pro—and the experiment used a defined set of clinical questions adapted from 20 Cochrane reviews. The testing context dates to July 2026, while AI products and their retrieval systems can change rapidly.

The benchmark also measures overlap with studies selected by Cochrane reviewers, not every dimension of medical answer quality. A model could retrieve a relevant study and summarize it badly, or fail to retrieve one included paper while still provide a clinically useful answer supported by other strong evidence. The paper is best read as an audit of evidence retrieval behavior rather than a comprehensive clinical safety ranking.

Even with those limitations, the study offers a useful warning for researchers, clinicians and patients increasingly tempted to treat an AI assistant as a literature-search shortcut. The average chatbot response retrieved only 39.2% of the expert-included studies, and performance changed sharply depending on both the model and the way the user presented the question.

The takeaway is not that medical AI citations are useless. It is that a citation-rich answer should not be mistaken for a systematic search. In this experiment, ChatGPT came substantially closer than the other tested assistants to the evidence set assembled by expert reviewers, but even its 63.1% recall left a meaningful share of Cochrane-included studies unretrieved. When clinical decisions depend on comprehensive evidence, the difference between finding some good studies and finding the relevant body of research remains critical.

0%