Academic papers whose abstracts contain writing patterns associated with large language models receive about 7% more citations than otherwise comparable papers, according to a new large-scale study published in Scientometrics. The result raises a provocative possibility: generative AI may be changing not only how research is written, but also how scientific impact is measured.
The open-access study, by Sofia Paklina, Petr Parshakov and Elena Rapoport, analyzes 234,073 research papers published between 2019 and 2025 by linking arXiv records with bibliometric information from Crossref. Papers classified as potentially written with LLM assistance received approximately 7.04% more age-adjusted academic citations on average than papers without those detected signals, and the association remained statistically significant after the researchers controlled for several established predictors of citation performance.
But the headline requires two major qualifications. The study measures citations from one scholarly work to another, not citations inside ChatGPT, Google AI Overviews or other generative answers. More importantly, the researchers do not know which authors actually used an LLM: AI involvement is inferred through a text classifier with 84.8% validation accuracy. The paper therefore documents an association between LLM-like writing patterns and citation outcomes, not proof that using ChatGPT or another model causes a paper to receive 7% more citations.
The study connects arXiv text with Crossref citation data
The researchers constructed a multidisciplinary corpus by matching arXiv research outputs with external DOIs and Crossref bibliometric records. After normalization, deduplication and filtering, they identified 256,032 eligible papers published between 2019 and 2025.
The final regression sample is smaller. Because the primary dependent variable is based on the natural logarithm of citations per year, the authors excluded 21,959 papers with zero citations. That produced the 234,073-paper dataset used in the main regression analysis.
This filtering decision is important for interpretation. The final sample is not a random representation of every arXiv submission, and it excludes precisely the papers that had accumulated no citations by the August 2026 citation refresh. The authors explicitly acknowledge that the dataset represents arXiv-linked research outputs for which they could observe both abstract text and verified publication-level bibliometric information.
The study adjusts citation counts for paper age because an article published in 2019 has had substantially more time to accumulate references than one published in 2025. The average paper in the regression corpus had 25.19 raw citations, while the mean age-adjusted citation count was 5.99.
AI use was inferred from writing style, not observed directly
The most important methodological issue is how the study decides whether a paper used an LLM. Researchers generally do not have reliable disclosure data showing which authors used ChatGPT, Claude, Gemini or another system while drafting or editing an abstract, so the authors created an indirect measurement strategy.
They trained a BERT-based binary text classifier using human-written abstracts from the pre-LLM era and synthetic rewrites generated by nine different language models. The models included GPT-4, Claude 3.5 Haiku, Gemini 2.0 Flash, GPT-3.5 Turbo, Qwen, Phi, Mistral and Cohere systems, among others.
The models were instructed to rewrite existing abstracts while preserving technical accuracy and scientific rigor. After cleaning failed or unusable cases, the labeled training corpus contained 9,848 abstracts. The researchers then fine-tuned the classifier to distinguish the original human texts from the AI rewrites.
On its validation set, the classifier achieved 84.8% accuracy, with a true-positive rate of 0.92 and a true-negative rate of 0.78. Those numbers are useful but far from perfect. A 15.2% validation error remains, and the authors explicitly describe the resulting classification as a probabilistic proxy rather than definitive evidence that a particular researcher used an LLM.
About 4.3% of papers were classified as likely LLM-assisted
When the trained classifier was applied to the 234,073 abstracts, approximately 4.3% were categorized as likely containing LLM-generated or LLM-rewritten text. The share increased sharply in the years following the public release of ChatGPT, reaching 23.5% among papers in the dataset published in 2025.
That trajectory is consistent with other research showing growing linguistic signatures associated with generative AI in academic publishing. It still should not be read as a direct measurement of actual adoption. Human authors can naturally use phrases or stylistic structures that resemble model output, while researchers who use AI lightly or substantially rewrite generated material may leave fewer detectable signals.
Detection also becomes a moving target. Language models change, researchers adapt their prompting and editing practices, and public awareness of stereotypical AI vocabulary can alter human writing. A classifier trained on one generation of model output may therefore become less reliable as writing behavior evolves.
The main estimate is a 7.04% citation advantage
The study’s primary regression estimates a coefficient of 0.068 for the LLM classification variable. Because the dependent variable is logarithmic, the authors convert that coefficient into an approximate 7.04% increase in citation counts using the standard exponential transformation. The result is statistically significant at p<0.01.
The association survives a substantial set of controls. The model accounts for the number of authors, reference count, abstract length, authors’ cumulative prior citation performance, publication and license type, and journal fixed effects. Those fixed effects are intended to reduce the possibility that the result merely reflects LLM-associated papers appearing in more prestigious or citation-heavy publication venues.
The researchers initially observed thousands of journal and proceedings labels and grouped smaller venues into an “Other” category for computational purposes, leaving 304 journal categories in the regression. The resulting estimate therefore captures variation within publication environments more conservatively than a simple comparison of average citation counts.
Even after those adjustments, abstracts carrying the classifier’s LLM-associated signals were linked to modestly higher citation performance.
A second model also found citations above historical expectations
The authors did not rely solely on the OLS regression. They also built a CatBoost machine-learning model trained on papers from 2021 and 2022, before widespread LLM-assisted academic writing became common. That model learned historical relationships between characteristics such as journal, references, authorship and citation outcomes.
It was then used to predict expected citations for papers published in 2023 and 2024. The researchers excluded 2025 from this particular prediction exercise because those papers had substantially less time to accumulate citations.
Papers classified as likely LLM-assisted subsequently received citation levels above the model’s predictions. Their mean deviation between actual and predicted log age-adjusted citations was positive and statistically significant, providing a second form of evidence consistent with the main regression result.
This does not turn the analysis into a randomized experiment. Historical prediction models can control for observed patterns, but they cannot eliminate every unobserved difference between researchers who produce LLM-like abstracts and those who do not. The counterfactual exercise strengthens the association without establishing causality.
The effect varies substantially by academic field
The pooled 7.04% estimate also hides considerable disciplinary variation. When the researchers interacted LLM classification with broad scientific fields, the strongest statistically significant associations appeared in Economics and Finance, Mathematics and Computer Science.
The estimated citation association reached 31.9% in Economics and Finance, 28.1% in Mathematics and 9.1% in Computer Science. Estimates in other fields were smaller and less precise, and the study reports no broad field with a statistically significant negative association.
Those field-level numbers deserve more caution than the headline average because smaller subgroups produce wider uncertainty and disciplinary publishing practices differ considerably. They do, however, reinforce an important point: there is unlikely to be one universal “AI writing effect” across science.
Fields vary in abstract conventions, publication speed, conference importance, language norms, citation density and adoption of generative tools. The same stylistic changes could therefore have different visibility effects depending on the research community.
Why might LLM-associated writing correlate with more citations?
The study proposes a visibility and presentation mechanism. Abstracts are critical interfaces between a research paper and the systems that help people discover it. They are indexed by scholarly databases, displayed in search results, processed by recommendation engines and often read before a researcher decides whether to open the full paper.
Language models are particularly good at producing fluent, standardized academic prose. They can restructure arguments, improve readability, surface familiar terminology and make an abstract conform more closely to disciplinary writing conventions. If those changes make a paper easier to discover or evaluate quickly, they could plausibly increase readership and eventually citations even when the underlying scientific contribution is unchanged.
But this remains an interpretation rather than a demonstrated causal pathway. The study does not randomly assign some authors to use LLMs and others to write unaided, then track subsequent citation behavior. Researchers who use AI may differ in other ways that are difficult to observe: they may be more technologically sophisticated, more internationally connected, more active in fast-moving topics or more deliberate about promoting their work.
Statistical controls reduce some alternative explanations, but observational data cannot remove all of them.
The detector cannot tell us who actually used ChatGPT
This is the most important limitation for interpreting the headline. The study’s independent variable is not documented LLM use. It is a classification based on textual resemblance to AI-generated rewrites.
A paper labeled “LLM-assisted” may have been edited by an AI system, but the classifier cannot prove that. A highly polished human abstract could trigger similar signals. Conversely, an author may have used an LLM extensively and then edited the result enough that the final abstract no longer resembles the detector’s learned patterns.
The 84.8% validation accuracy makes the method useful for studying population-level patterns, but it is not sufficient for accusing individual authors of undisclosed AI use. The paper itself is careful about this distinction and frames the measure as predicted or potential LLM assistance.
That caution is especially important because automated AI detection has often been misused as a binary authorship test. A statistical classifier can support aggregate research while still being inappropriate as evidence about a particular person’s behavior.
This has nothing to do with citations inside AI answers
The word “citation” can also create confusion in the current search environment. GEO research frequently measures whether ChatGPT, Perplexity, Gemini or Google AI Overviews cite a website as a source. This Scientometrics study examines something completely different.
Its outcome is scholarly citation impact: whether later academic publications reference the research paper. The 7.04% figure therefore cannot be used to claim that LLM-written abstracts are 7% more likely to appear as sources in AI-generated answers.
The distinction matters because the mechanisms are different. Academic citations emerge through researchers reading and referencing prior work. AI citations emerge through retrieval, ranking, grounding and source-selection systems operated by AI platforms. A writing style that improves one type of visibility does not automatically improve the other.
Excluding zero-citation papers could matter
The removal of 21,959 uncited papers is another meaningful limitation. The authors explain that zero-citation observations cannot be included directly in their primary log-citation specification. That is a legitimate modeling constraint, but it changes the population being analyzed.
The final 234,073-paper regression therefore estimates the relationship among papers that had received at least one citation. It does not directly tell us whether LLM-associated writing increases the probability that an otherwise uncited paper receives its first citation.
If LLM-style writing is associated differently with the zero-citation group, excluding those papers could influence the estimated magnitude. Readers should therefore resist translating the 7.04% figure into a universal premium across every research paper published between 2019 and 2025.
The larger issue is what citation metrics actually reward
The study’s most interesting implication may not be whether researchers should use AI to polish abstracts. It is what happens to bibliometric evaluation when writing technology changes.
Citation counts are widely used as proxies for scientific impact in hiring, promotion, funding, journal evaluation and institutional rankings. Yet citations have always reflected more than underlying research quality. Journal prestige, collaboration networks, field size, language, discoverability and presentation can all affect how often a paper is referenced.
If generative AI makes academic writing more readable, standardized or discoverable, then access to and proficiency with these tools could become another factor embedded inside citation metrics. A researcher might receive a visibility advantage from presentation technology even if the scientific contribution itself is unchanged.
The authors argue that this possibility should make institutions more cautious when interpreting bibliometric measures. A 7% difference may appear modest, but cumulative citation advantages can influence thresholds, rankings and career-level indicators over time.
A useful association, not a recipe for citation growth
The new Scientometrics paper provides unusually large-scale evidence that LLM-associated writing patterns correlate with higher academic citation counts. Its multidisciplinary arXiv–Crossref corpus, rich regression controls and historical counterfactual analysis make the result more informative than a simple comparison between obviously AI-written and human-written abstracts.
But the study does not establish that asking an LLM to rewrite an abstract will increase citations by 7.04%. It cannot verify which authors used AI, its classifier makes errors, the main regression excludes uncited papers, and unobserved differences between groups may remain despite extensive controls.
The most defensible conclusion is narrower and more consequential: papers whose abstracts look like the kinds of academic text produced by contemporary LLMs are associated with a modest citation advantage after accounting for several conventional predictors of scholarly impact.
That is enough to raise a serious question for research evaluation. If presentation can increasingly be optimized by generative systems, citation counts may capture not only the influence of scientific ideas, but also the technologies used to package those ideas for discovery.