Citation & References

Citation hygiene: what reviewers check in your reference list

Two teams measured the same thing from opposite directions and arrived at nearly the same number. Baethge and Jergas pooled 46 studies covering 32,074 quotations in the medical literature and found that 16.9% of them misstated the source they pointed at, with about half of those counted as major errors. Wakeling and colleagues skipped the outside auditors and asked 2,648 corresponding authors to judge a real citation of their own work. Those authors, the people best placed to know, called 16.6% of the citations inaccurate.

So roughly one citation in six does not say what the citing sentence claims it says. Reviewers have read the same studies, or have simply been burned enough times to behave as though they had.

Nobody audits a bibliography line by line. What a reviewer does is pull the three or four citations holding up the argument and open them, which is enough to find a problem often enough to be worth the ten minutes. The citation mistakes that turn up in a research paper fall into two families: the reference is wrong about what its source says, or it is wrong about which source it is. Both families are findable by the author in an afternoon, and almost nobody looks.

What counts as a citation mistake?

A citation mistake is any gap between what your sentence claims a source establishes and what that source, correctly identified, actually contains. That covers misreading, but it also covers a reference pointing at a paper that has been retracted or never existed. The two families behind those cases carry different costs. An identity error looks like carelessness, which is survivable and embarrassing. A quotation error is a claim your paper makes and cannot support, and a reviewer who finds one starts wondering what else was asserted from memory.

Mistake What it looks like in the manuscript How a reviewer finds it
Quotation error The cited paper does not support the sentence, or reports something narrower Opens a load-bearing citation and reads the passage
Secondary citation You cite an original study for a result you read about in a review Knows the original; it says something more limited
Wrong identity Year, volume, pages, or DOI point at a different paper Clicks the DOI, or recognizes the title
Version drift A preprint is cited after the journal version appeared with revised numbers Reads the published version
Retracted source A claim rests on a paper withdrawn two years ago Reference manager alert, or memory of the case
Fabricated reference Plausible authors, real journal, no such article The search returns nothing and the DOI fails to resolve
Clump padding Four references stacked behind one clause, none of them discussed Scans in-text citation density
Single-source clump Many references to one journal or one group, loosely related Notices the pattern while reading the list

Only the last two rows are visible from the manuscript alone. Everything above them requires opening something, which is exactly why those errors survive to review.

How often are citations wrong in published papers?

Between one in six and one in four, depending on the study, and the estimate has been stable for a decade despite everyone knowing about it.

The 2015 meta-analysis by Jergas and Baethge covered 28 studies and 7,321 references and put the total quotation error rate at 25.4%, with 11.9% of references seriously incorrect. Mogull recalculated the same literature in 2017 using a stricter denominator, errors per quotation examined rather than per article, and reported 14.5%. The 2025 update by Baethge and Jergas, now covering 46 studies, landed at 16.9% overall and 8.0% major. Its meta-regression found no improvement over time.

One number in that update deserves separate attention. Secondary quotations, where an author cites a primary source for something they encountered in a review or another paper’s introduction, accounted for 5.3% of quotations. That is the mechanism behind a large share of the rest: a claim gets copied from one paper’s framing into the next, drifting a little each time, until a sentence in your introduction attributes something to a 1998 study that the 1998 study never claimed.

These figures come from medicine, where this kind of audit is done most often. Comparable audits in other disciplines are scarce, so the working assumption has to be that yours is not the exception.

How do you check that a source says what you claim it says?

Give every load-bearing citation an evidence locus: a page, table, or figure number where the support actually sits, written down where you can see it.

The pass takes an afternoon for a typical manuscript.

  1. Extract every sentence that makes an empirical claim and carries a citation. Searching the manuscript for (20, [, or et al. catches most of them. Background sentences citing a textbook definition can be skipped.
  2. Rank them by how much the paper leans on each one. The claims that justify your design or your gap statement go at the top. Fifteen sentences is a normal count for the top of that list.
  3. For each, open the cited source and find the specific sentence, table, or figure that supports your claim. Write the locus in a scratch column next to the sentence: Fig 3, p. 1147, Table 2, column 4.
  4. Any claim where you cannot find the locus in two minutes gets one of three fixes. Rewrite the claim to match what the source says, replace the citation with one that supports it, or drop the claim.
  5. Mark every citation you have never actually read in full. If you learned the result from a review, either cite the review as a review, or read the primary and cite it for exactly what it reports.

Step five is where the 5.3% lives. It is also the step authors resist most, because the citation feels earned: you did read something, and that something told you this. Reviewers who know the primary literature are precisely the people who catch it.

Some dedicated academic tools, Fukuro among them, will flag claims in your text that carry no citation at all, and references that appear in the list but never get discussed. No tool can open a PDF and confirm that its Table 2 supports your sentence about effect sizes. That remains a human pass, and it is the one that changes sentences.

Where do fabricated references come from?

Increasingly from language models, and the rates are not marginal. When Bhattacharyya and colleagues had ChatGPT generate 30 short medical papers in 2023, 47% of the 115 references it produced were fabricated outright, 46% were real papers cited inaccurately, and 7% were both real and correct. Model quality has improved since then, and current systems with search access invent fewer sources. The failure has not disappeared, and it now arrives in a subtler form: a real paper attached to a claim it does not make, which is a quotation error wearing a valid DOI. Editors have started screening for it. Journals in several publisher groups now ask authors to declare AI assistance, and a bibliography containing one unresolvable DOI invites a check of the rest.

The mechanical version of this check is fast. Copy your reference list into Crossref’s Simple Text Query, which accepts up to 1,000 references in any style, one per line, and returns the DOI it matches to each. References that come back unmatched are your worklist. Some will be legitimately absent, since books, theses, and older articles often have no DOI. The rest are typos, invented sources, or entries whose metadata drifted far enough from the real record that a matcher cannot recognize them, which is roughly what a reviewer’s search engine will also experience.

How do you catch a retracted reference?

Switch on retraction alerts in your reference manager, because the manual check never happens on schedule. Hsiao and Schneider examined 13,252 citation contexts to retracted biomedical papers published after the retraction and found that only 5.4% acknowledged it. Authors were not defending the retracted work. They had no idea. The infrastructure improved in 2023, when Crossref acquired the Retraction Watch database and made the retraction metadata freely available through its API. That put retraction status where software can reach it. Zotero flags retracted items in your library and warns you when you try to cite one, covering entries that carry a DOI or PMID, which is around three quarters of the Retraction Watch records. Other managers have added comparable checks.

Run the check while drafting rather than at submission. A retraction caught early costs you a paragraph and an afternoon of reading replacements. The same retraction caught by reviewer two forces a revision round that has to explain how a withdrawn paper ended up carrying part of your argument.

When does a reference list look manipulated?

A list looks manipulated when its shape stops matching its argument. Reviewers are not counting self-citations, but they notice when a cluster of references to one journal or one group sits at a point in the paper where the argument does not need them.

COPE treats citation manipulation as a publication ethics matter rather than a stylistic one, covering excessive self-citation, journal-level inflation, and coordinated citation between journals. Editors have their own version of the problem. Wilhite and Fong surveyed academics about editors who ask for citations to the editor’s own journal as an implicit condition of acceptance, and found that while 86% consider the practice inappropriate, 57% said they would add superfluous citations before submitting to a journal known to coerce, and under 7% expected any author to refuse.

That gap between what authors condemn and what authors do is why reference lists carry residue. Some of it is deliberate, most of it is accumulated, and a reviewer cannot tell which from the outside. What they can see is a clause supported by four references from the same lab, or a related work section citing a target journal eleven times without engaging any of the eleven. The fix is a job test: every reference exists to support a specific sentence, and if you cannot name the sentence, the reference is decoration. Applied to a suspicious cluster, that test either produces the sentence each reference supports, or it produces a cut.

Why does any of this change a reviewer’s mind?

Because the reference list is the only part of the paper a reviewer can verify without trusting you.

They cannot rerun your experiment. They cannot audit your raw data, in most cases they never see it, and your analysis code either exists or is a promise in a data availability statement. The bibliography is different: every entry is a checkable claim about something already public, and a reviewer with a browser can test any of them in thirty seconds.

That asymmetry is why citation errors land harder than their intellectual weight suggests. A wrong page number is trivial. A wrong page number found by a reviewer who has just been told your effect held under three sensitivity analyses is a data point about how the rest of the paper was assembled, and it is the only data point of that kind they are able to collect.


References and useful policies

  1. Baethge C, Jergas H. Systematic review and meta-analysis of quotation inaccuracy in medicine. Research Integrity and Peer Review. 2025;10:13.
  2. Jergas H, Baethge C. Quotation accuracy in medical journal articles: a systematic review and meta-analysis. PeerJ. 2015;3:e1364.
  3. Mogull SA. Accuracy of cited “facts” in medical research articles: a review of study methodology and recalculation of quotation error rate. PLOS ONE. 2017;12(9):e0184727.
  4. Wakeling S, Paramita ML, Pinfield S. How do authors perceive the way their work is cited? Findings from a large-scale survey on quotation accuracy. Journal of the Association for Information Science and Technology. 2025;76(10):1396-1410.
  5. Hsiao TK, Schneider J. Continued use of retracted papers: temporal trends in citations and (lack of) awareness of retractions shown in citation contexts in biomedicine. Quantitative Science Studies. 2021;2(4):1144-1169.
  6. Bhattacharyya M, Miller VM, Bhattacharyya D, et al. High rates of fabricated and inaccurate references in ChatGPT-generated medical content. Cureus. 2023;15(5):e39238.
  7. Wilhite AW, Fong EA. Coercive citation in academic publishing. Science. 2012;335(6068):542-543.
  8. COPE. Citation manipulation (discussion document).
  9. COPE. Handling citation manipulation (position statement).
  10. ICMJE. Recommendations: preparing a manuscript for submission.
  11. Crossref. Simple Text Query.
  12. Crossref. Retraction Watch data.
  13. Retraction Watch. The Retraction Watch Database.
  14. Zotero. Retracted item notifications.