NTH

Most biomedical publications show signs of LLM-assisted writing

AuthorsLena Holzwarth, Rita González-Márquez, Dmitry Kobak

August 17, 2026 2 min read
Watch on YouTube
The one-line take

A large-scale study estimates that most biomedical papers now contain signs of LLM-assisted writing, especially in their Discussion sections.

Key results

1.194287M
PMC paper corpus

English biomedical papers analyzed from PubMed Central

379
LLM marker vocabulary

Non-content marker words used to detect excess vocabulary

89%
Full-paper prevalence

Estimated papers showing LLM-assisted writing signs in December 2025

68%
Discussion crop prevalence

Estimated usage in comparable 255-word Discussion crops

32%
Methods crop prevalence

Estimated usage in comparable 255-word Methods crops

0.02
Simulation absolute error

Upper bound on absolute estimation error across simulated usage rates

What the paper found

This study introduces an assumption-light estimator for LLM-assisted writing that tracks excess frequencies of vocabulary associated with language-model editing. It fits linear regressions to word frequencies during 2018–2022, before OpenAI released ChatGPT, extrapolates a human-writing counterfactual through 2025, and optimizes across sets of 379 marker words rather than relying on prompts, synthetic reference text, or unreliable AI detectors. Applied to 1.194287M English biomedical papers from PubMed Central, the method estimates that 89% of full papers showed signs of LLM assistance by December 2025; this indicates detectable editing or drafting, not necessarily wholesale machine authorship. Section-level analysis using comparable 255-word crops found the highest prevalence in Discussion paragraphs, at 68%, versus 32% in Methods paragraphs, although Methods still exceeded 50% when entire sections were analyzed. A simulation with known ground truth recovered the true usage rate with absolute error below 0.02 across the tested range. The authors argue that earlier frequency-gap methods systematically underestimated prevalence, while the observed convergence toward LLM-associated vocabulary—especially among non-native English-speaking authors—raises policy issues involving disclosure, hallucinated citations, scientific homogenization, and research integrity.

Original abstract

Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decisions, it is necessary to monitor the prevalence of LLM-altered texts in scholarly publications. Despite some recent progress in this direction, no existing method can produce reliable estimates. Here we suggest and validate a new unbiased approach to estimate LLM usage in a corpus of texts based on changing word frequencies. We apply our method to the full texts of open-access biomedical papers from Pubmed Central, and show that by the end of 2025, 89% of papers show excess of LLM-associated vocabulary. We also find that LLMs are twice as likely to be used when writing a paragraph in the Discussion section (68%) compared to a paragraph in the Methods section (32%), but even inside the Methods section, the overall prevalence of LLM usage is over 50%. We believe that our estimates are crucial to shape future guidelines and policies.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →