Most biomedical publications show signs of LLM-assisted writing

arXiv:2608.10715 · cs.CL, cs.AI, cs.CY, cs.DL, cs.SI · Submitted 2026-08-11 · Read on arXiv

Lena Holzwarth, Rita González-Márquez, Dmitry Kobak

Hertie Institute for AI in Brain Health, University of Tübingen · Department of Mathematics, Computer Science, and Statistics, Ghent University · VIB Center for AI and Computational Biology, VIB

cs.CL, cs.AI, cs.CY, cs.DL, cs.SI

Submitted: 2026-08-11

Updated: 2026-08-12

Code: https://github.com/kobaklab/llm-usage-in-pmc

Project page: https://pmc.ncbi.nlm.nih.gov/tools/openftlist

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 100/100

The gist: Based on the paper "Most Biomedical Publications Show Signs of LLM-Assisted Writing" by Holzwarth, González-Márquez, and Kobak: Summary This paper presents a new method to estimate the prevalence

Terminology

Summary

Based on the paper Most Biomedical Publications Show Signs of LLM-Assisted Writing by Holzwarth, González-Márquez, and Kobak:

Summary

This paper presents a new method to estimate the prevalence of LLM-assisted writing in biomedical publications and applies it to a large corpus of open-access papers from PubMed Central (PMC). The authors analyzed 1,194,287 English-language papers published between 2017 and 2025 that contained sections titled Introduction, Methods, Results, and Discussion.

The method builds on prior work identifying 379 non-content marker words (e.g., these, potential, delves) that showed increased usage in PubMed abstracts after the release of ChatGPT in November 2022. For each marker word, the authors computed its usage frequency q(t) (fraction of documents containing the word) over time. They fit a linear regression to the pre-LLM period (2018–2022) and extrapolated to 2023–2025 to obtain p̂human, the counterfactual frequency expected in human-written texts without LLMs. The frequency gap Δ = q − p̂human provides a lower bound on LLM usage. The authors then derived a tighter lower bound: β̂ LB = (q − p̂human)/(1 − p̂human), which assumes pLLM ≤ 1. To optimize this, they selected sets of marker words G(T) containing all words with frequency below a threshold T, and chose the threshold yielding the maximum β̂ LB after discarding estimates with standard error above 0.025. They validated this procedure with simulations using 100,000 documents per year and realistic parameter ranges, achieving β̂ − β < 0.02 for all simulated β ∈ [0,1].

Applying this method, the authors found that by the end of December 2025, 89% of PMC papers showed signs of LLM-assisted writing or editing (β̂ = 0.89 for the full paper). For the entire year 2025, the estimate was 77% for full papers and 53% for abstracts. Monthly resolution showed continually increasing usage starting in early 2023.

When comparing sections using random 255-word crops (to control for section length), the Discussion had the highest LLM usage, reaching 0.68 in December 2025, closely followed by abstracts at 0.67. The Methods section had the lowest usage at 0.32 in December 2025. The authors note that "LLMs are twice as likely to be used when writing a paragraph in the Discussion section (68%) compared to a paragraph in the Methods section (32%), but even inside the Methods section, the overall prevalence of LLM usage is over 50%."

The authors also analyzed LLM usage by country of first author's affiliation. In 2025, the highest estimates were for South Korea (0.85), China (0.82), and Taiwan (0.80), while the lowest was for the UK (0.28). When grouping countries by whether they have a majority of native English speakers, the authors obtained β̂ values of 0.37 for native English-speaking countries and 0.72 for other countries. They observed that "The pre-2023 frequency of LLM marker words was much lower in non-natively English-speaking countries, but it became very close to the frequency in natively English-speaking countries by 2025, indicating strong linguistic convergence driven by the LLM usage."

The authors compare their estimates to existing literature. They note that previous word-frequency-gap studies reported 2024 estimates of 12–16%, much lower than their 31% (abstracts) and 52% (full papers) for 2024, because frequency gaps systematically underestimate true LLM usage. A distribution-overlap study found 57% in 2025, also an underestimation. Mixture-model studies produced highly variable results (18–35%) depending on prompts and models. The authors note their estimates are consistent with surveys, e.g., Wiley (2026) found over 70% of respondents using LLMs for writing assistance in mid-2025.

The paper discusses limitations, primarily the assumption that linear extrapolation of p̂human remains valid for 2025. The authors acknowledge that human vocabulary may shift due to exposure to LLM writing styles, or conversely, people may consciously avoid marker words. They argue that any effects arising from changes to human vocabulary are smaller and slower than the effects arising from direct LLM-assisted editing.

The authors discuss policy implications, noting both benefits (e.g., overcoming language barriers, increasing efficiency) and dangers of LLM-assisted writing, including hallucinations, fraud, reduced review quality, homogenization of language and reasoning, and potential loss of individual creativity. They conclude that widespread reliance on LLMs can pose serious challenges to academia and that these need to be addressed, both on an individual level by critical reflection, and on an institutional level by developing guidelines and increased safeguards against fraud and misconduct.

Improvements for AI systems

Improvements to AI Systems Based on This Paper:

  1. Adaptive Writing-Style Calibration for Scientific Text Generation
  • Improvement: Train LLMs to detect and reduce over-reliance on the 379 identified marker words (e.g., delves, potential, these) in biomedical writing.

  • Capability: The AI can generate or edit manuscripts with a style that statistically matches pre-LLM human writing distributions, reducing detectability and preserving linguistic diversity across sections (e.g., Methods vs. Discussion). This helps researchers who need to avoid homogenization or who face journal policies against undisclosed LLM use.

  1. Section-Aware Writing Assistance with Usage-Awareness
  • Improvement: Implement a module that tracks the estimated LLM-assistance probability (β̂) per section in real-time during drafting.

  • Capability: The AI can alert users when a section (e.g., Discussion) is likely to exceed a threshold (e.g., 50% LLM-style), prompting them to add more original phrasing, technical detail, or personal insight—especially for Methods sections, which currently show lower LLM usage and may benefit from more human-specific procedural nuance.

  1. Cross-Lingual Linguistic Convergence Mitigation
  • Improvement: Use the paper’s finding that non-native English speakers converge to LLM marker-word frequencies faster than native speakers to build a “linguistic fingerprint” for non-native users.

  • Capability: The AI can offer targeted suggestions to preserve the user’s original voice and idiomatic diversity, reducing the risk of homogenized academic English and helping non-native speakers maintain distinct, authentic phrasing while still improving clarity.

  1. Counterfactual Frequency Estimation for Plagiarism and Authenticity Checks
  • Improvement: Integrate the β̂ LB estimator into AI-based plagiarism or authorship-verification tools.

  • Capability: The AI can provide a lower-bound probability that a given manuscript (or section) was LLM-assisted, using time-series extrapolation from pre-2023 baselines. This enables journals to flag suspicious submissions with a quantified confidence interval, rather than relying on binary classifiers.

  1. Dynamic Marker-Word Set Selection for Style Control
  • Improvement: Use the paper’s threshold-based optimization (G(T) selection) to dynamically adjust which stylistic markers the AI avoids or emphasizes, based on target audience or venue.

  • Capability: The AI can tailor its output to match the expected human style of a specific journal or field (e.g., lower marker frequency for UK-based journals, higher for South Korean or Chinese submissions), improving acceptance rates and reducing unintended LLM-signature.

  1. Longitudinal Bias Correction in Language Models
  • Improvement: Incorporate the paper’s linear extrapolation method to correct for temporal drift in human vocabulary (e.g., post-2022 shifts) when fine-tuning on recent biomedical corpora.

  • Capability: The AI can distinguish between genuine human vocabulary evolution and LLM-induced changes, preventing the model from overfitting to LLM-contaminated text and maintaining accuracy in tasks like literature summarization or evidence extraction.

  1. Policy-Compliance Assistant for Academic Writing
  • Improvement: Build a tool that uses the paper’s prevalence estimates (e.g., 89% by 2025) to inform users about institutional guidelines and ethical disclosure requirements.

  • Capability: The AI can proactively remind authors to declare LLM assistance, based on the detected likelihood of LLM usage in their draft, and suggest rewrites that lower the estimated β̂ if they wish to avoid disclosure (while encouraging transparency).

Abstract

Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decisions, it is necessary to monitor the prevalence of LLM-altered texts in scholarly publications. Despite some recent progress in this direction, no existing method can produce reliable estimates. Here we suggest and validate a new unbiased approach to estimate LLM usage in a corpus of texts based on changing word frequencies. We apply our method to the full texts of open-access biomedical papers from Pubmed Central, and show that by the end of 2025, 89% of papers show excess of LLM-associated vocabulary. We also find that LLMs are twice as likely to be used when writing a paragraph in the Discussion section (68%) compared to a paragraph in the Methods section (32%), but even inside the Methods section, the overall prevalence of LLM usage is over 50%. We believe that our estimates are crucial to shape future guidelines and policies.

Sources

Related papers