A machine-learning classifier screened 2.6 million cancer research papers published between 1999 and 2024 and flagged nearly one in ten as textually similar to retracted "paper mill" publications, according to a peer-reviewed study in The BMJ.
A machine-learning filter puts a number on a long-suspected problem
Paper mills are commercial operations that write and sell fabricated or low-quality manuscripts, sometimes bundled with authorship slots, to researchers under pressure to publish. Byrne, Barnett, and colleagues built a text classifier based on BERT — a language model architecture — and trained it on 2,202 papers that had been retracted from the literature and tagged specifically as paper-mill products in the Retraction Watch database, alongside a matched set of presumed-genuine control papers.
Tested against an independent set of suspected paper-mill papers compiled separately by research-integrity investigators, the model correctly classified papers as paper-mill-like or genuine 91% of the time in internal validation and 93% in external validation. It also caught roughly 72% of papers that three earlier, unrelated studies had already flagged for problems such as unverifiable gene sequences or misidentified cell lines — papers the model had never seen during training, which the authors treat as evidence the classifier is not simply memorizing its own training set.
Applied to the full screening corpus of 2,647,471 original cancer research papers, the model flagged 261,245 of them — 9.87%, with a 95% confidence interval of 9.83% to 9.90%. The authors are explicit that a "flagged" paper only means its title and abstract share textual patterns with known paper-mill output; it is a statistical screen, not a determination that any specific paper is fraudulent.
The flagged share has climbed sharply since the early 2000s
The trend line is where the study's argument sharpens from "there's a problem" to "the problem is accelerating." In the early 2000s, flagged papers made up roughly 1% of annual cancer research output. By 2022, that share had climbed to more than 15% of the year's output (26,457 of 171,656 papers) — the study describes the rise as following a roughly exponential curve, with a slight pullback in 2023 and 2024 that the authors attribute to some mix of publisher pushback, paper mills shifting to new templates, and normal indexing lag for the most recent years rather than to a resolved problem.
The same upward pattern shows up inside the most selective part of the literature. Among the top 10% of journals by impact factor, the flagged share also rose over time, passing 10% by 2022 — even as the impact-factor threshold required to be in that top decile itself rose, from 3 in 1999 to 7 in 2021. The authors read this as undercutting the assumption that impact factor alone protects a journal from paper-mill submissions.
China and several Middle Eastern and South Asian countries show the highest flagged shares
Breaking the flagged papers down by the first author's institutional country shows a wide spread. China had both the highest flagged share (36% of its cancer papers, 177,907 of 497,672) and the largest absolute number of flagged papers. Iran followed at 20%, with Saudi Arabia, Egypt, Pakistan, and Malaysia each in the 13–16% range. The United States, by contrast, had a flagged share of just 2% — though because it publishes so much cancer research, it still produced the second-highest absolute count of flagged papers (10,511) after China.
The authors are careful about what this pattern does and doesn't support. A skew toward flagging Chinese-authored papers could, in principle, mean the model has simply learned to associate Chinese scientific writing style with fraud rather than genuine paper-mill features. But the study's own error analysis argues against that: false positives — genuine control papers wrongly flagged — were rare overall (39 of 3,375) and showed no comparable geographic skew, while false negatives — actual paper-mill papers the model missed — were disproportionately linked to Chinese-affiliated authors (accounting for 90% of misses, versus 47% of the validation set overall). If the model were simply penalizing a writing style, the bias would more plausibly show up as excess false positives, not as the model under-flagging papers from that same group.
Gastric, bone, and liver cancer research carry the heaviest concentration of flagged papers
The flagged papers are not spread evenly across cancer subfields. Gastric cancer research had the highest flagged share, at 22% (18,398 of 82,690 papers), followed by bone cancer at 21% (8,458 of 39,433) and liver cancer at 20% (26,730 of 136,719) — all well above the 9.87% corpus-wide average. Most other cancer types clustered in a 10–15% range, with breast, skin, prostate, and blood cancers sitting at the low end. By sheer volume, lung and liver cancer research produced the most flagged papers overall, because both fields publish at high volume.
The study offers a partial, hedged explanation rather than a firm mechanism: gastric and liver cancers are unusually common in China, and both fields have separately been shown to have high rates of misidentified cell lines in prior integrity research — which the authors note "may also reflect vulnerabilities exploited by paper mills when popular research topics are targeted," alongside the possibility that early templates were simply reused and adapted repeatedly within these subfields. Flagged papers were also concentrated in fundamental and preclinical cancer biology research rather than in clinical, epidemiological, or health-policy work — areas the authors suggest may be both easier to fabricate convincingly and harder for reviewers to independently verify.
None of this is being treated as a closed case by the researchers themselves. The classifier's authors put the estimated positive predictive value at roughly 0.7 assuming a true prevalence near 10%, meaning close to three in ten flagged papers could still be false positives under their own math — which is why the model is being piloted, not deployed as a verdict. Three journals from a major publisher are currently using it inside their submission systems, feeding flagged manuscripts to human editors for closer scrutiny rather than automatic rejection, and authors are not told when their paper has been flagged, specifically so paper mills cannot use the feedback to refine their templates. The Queensland University of Technology-issued summary of the findings frames the tool as "a scientific spam filter" — useful for triage, not a substitute for the manual investigation that confirms or clears any individual paper.
Comments (0)
Please sign in to join the discussion.
No comments yet.
Be the first to share your perspective on this topic.