Most biomedical publications show signs of LLM-assisted writing
AuthorsLena Holzwarth, Rita González-Márquez, Dmitry Kobak
Resources
A large-scale study estimates that most biomedical papers now contain signs of LLM-assisted writing, especially in their Discussion sections.
Key results
English biomedical papers analyzed from PubMed Central
Non-content marker words used to detect excess vocabulary
Estimated papers showing LLM-assisted writing signs in December 2025
Estimated usage in comparable 255-word Discussion crops
Estimated usage in comparable 255-word Methods crops
Upper bound on absolute estimation error across simulated usage rates
What the paper found
This study introduces an assumption-light estimator for LLM-assisted writing that tracks excess frequencies of vocabulary associated with language-model editing. It fits linear regressions to word frequencies during 2018–2022, before OpenAI released ChatGPT, extrapolates a human-writing counterfactual through 2025, and optimizes across sets of 379 marker words rather than relying on prompts, synthetic reference text, or unreliable AI detectors. Applied to 1.194287M English biomedical papers from PubMed Central, the method estimates that 89% of full papers showed signs of LLM assistance by December 2025; this indicates detectable editing or drafting, not necessarily wholesale machine authorship. Section-level analysis using comparable 255-word crops found the highest prevalence in Discussion paragraphs, at 68%, versus 32% in Methods paragraphs, although Methods still exceeded 50% when entire sections were analyzed. A simulation with known ground truth recovered the true usage rate with absolute error below 0.02 across the tested range. The authors argue that earlier frequency-gap methods systematically underestimated prevalence, while the observed convergence toward LLM-associated vocabulary—especially among non-native English-speaking authors—raises policy issues involving disclosure, hallucinated citations, scientific homogenization, and research integrity.
Original abstract
Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decisions, it is necessary to monitor the prevalence of LLM-altered texts in scholarly publications. Despite some recent progress in this direction, no existing method can produce reliable estimates. Here we suggest and validate a new unbiased approach to estimate LLM usage in a corpus of texts based on changing word frequencies. We apply our method to the full texts of open-access biomedical papers from Pubmed Central, and show that by the end of 2025, 89% of papers show excess of LLM-associated vocabulary. We also find that LLMs are twice as likely to be used when writing a paragraph in the Discussion section (68%) compared to a paragraph in the Methods section (32%), but even inside the Methods section, the overall prevalence of LLM usage is over 50%. We believe that our estimates are crucial to shape future guidelines and policies.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.