Peer Review Was Built for a Different World

Peer Review Was Built for a Different World. Who Checks the Citations Now?

July 19, 20268 min read

Peer review carries enormous weight in academic publishing.

A paper that has passed peer review has survived scrutiny by people with relevant expertise. Its methods may have been challenged. Its reasoning may have been questioned. Its contribution may have been judged against the existing literature.

But there is one assumption we should examine more carefully:

Who actually checked the references?

Not whether they were formatted correctly.

Not whether they looked plausible.

Who checked that each source actually existed, and that the authors, title, journal, year and DOI were accurate?

Recent evidence suggests this question is becoming harder to ignore.

In 2026, an audit published in The Lancet examined 97.1 million references across approximately 2.5 million biomedical papers in the PubMed Central Open Access collection. The researchers identified 4,406 fabricated references appearing across 2,810 papers. The prevalence increased sharply in recent years: from approximately one paper in 2,828 in 2023 to one in 458 in 2025 and one in 277 among papers published during the first seven weeks of 2026.

Another study, published as a preprint in July 2026, examined accepted camera-ready papers from ICLR, ICML, NeurIPS and USENIX Security. Using a deliberately conservative definition focused on nonexistent works and substantial authorship mismatches, the researchers found likely hallucinated references in the archival record. In 2025, roughly one in twenty accepted papers at both NeurIPS and USENIX Security contained at least two likely hallucinated academic-paper references under that definition.

These findings do not prove that peer review is failing as a whole.

They point to a narrower problem.

Peer review is not the same as citation verification.

Peer review is not citation verification

Peer review serves an important purpose, but that purpose is broad.

Reviewers may assess methodology, interpretation, originality, relevance and engagement with prior research. They may notice an important paper has been omitted, question whether a source is appropriate, or recognise that a citation has been misused.

Systematically confirming that every cited work exists and that its bibliographic metadata is accurate is a different task.

A manuscript can therefore receive serious intellectual scrutiny without every reference being independently verified.

The 2026 Phantom References preprint makes this distinction particularly visible. Its authors chose citations precisely because they present a more auditable problem than evaluating the truth of technical claims: a scholarly reference can be checked against external bibliographic records to establish whether a compatible work exists. Their findings suggest that even papers accepted at highly selective conferences can contain references whose scholarly identity cannot be externally verified.

This should not be interpreted as evidence that reviewers are careless.

It raises a different question:

Was citation verification ever clearly assigned to them in the first place?

For much of the history of scholarly publishing, any ambiguity around that responsibility existed under different production conditions.

Those conditions are changing.

Citation production has scaled

Before generative AI, producing references at scale generally required more human friction: searching, copying, importing, reading, or encountering citations through existing literature.

Errors still occurred.

Authors mistyped titles. Years were wrong. Names were misspelled. References were copied second-hand. Sources were cited without being read.

None of this began with AI.

What has changed is the speed and volume at which plausible references can now be produced.

Generative systems can suggest citations in seconds. They can produce bibliographies that appear academically convincing while containing combinations of real authors, plausible titles, genuine journals and incorrect or nonexistent details.

The important change is therefore not simply that AI can make mistakes.

Humans make mistakes too.

The structural change is this:

Citation production has scaled. Citation verification has not.

The Lancet audit provides one indication of what this mismatch may look like in practice. The researchers found a steep rise in fabricated references beginning during the period in which generative AI tools became widely adopted, although temporal association alone does not establish that AI caused every fabricated citation they identified.

A separate 2026 large-scale preprint examining 111 million references across 2.5 million papers and preprints also reported a sharp post-LLM-adoption increase in nonexistent references. Its authors concluded that preprint moderation and journal publication processes were detecting only a fraction of the errors identified by their audit. As a preprint, those findings should be interpreted with appropriate caution, but they reinforce the case for investigating the capacity of existing safeguards.

Whatever the precise contribution of generative AI to these trends, the operational asymmetry is difficult to escape.

A researcher, student or writer can now produce candidate references far faster than a human can manually verify them.

The speed at which citations can enter scholarly work has increased.

The capacity to check them remains dependent on human attention and workflows that were not necessarily designed for systematic verification at this scale.

And that exposes another problem.

Custom HTML/CSS/JavaScript

Collective responsibility is not operational ownership

Academic integrity is often described as a shared responsibility.

Authors have responsibilities.

Supervisors have responsibilities.

Editors have responsibilities.

Reviewers have responsibilities.

Publishers have responsibilities.

That principle is reasonable.

But shared responsibility does not automatically create a clearly owned task.

When everyone has some responsibility for citation integrity, it can become unclear who is responsible for performing the specific act of citation verification.

An author may assume that a reference suggested by a trusted tool is legitimate.

A supervisor may focus on the quality of the argument.

A reviewer may assume that basic bibliographic checks have already been completed.

An editor may assume that reviewers would notice suspicious references.

A production team may focus on formatting rather than establishing whether the source exists.

Each person may act reasonably within the boundaries of their role.

And yet the citation itself may never be independently checked.

That is the difference between collective responsibility and operational ownership.

A system can value integrity while still leave a critical task unowned.

The question is therefore not simply whether authors, editors or reviewers care about citation accuracy.

It is whether the workflow contains an identifiable point at which someone is expected to check it.

Who performs that check?

When does it happen?

What exactly is being verified?

And who is responsible if the check never occurs?

Without clear answers, responsibility can remain everywhere in principle and nowhere in practice.

What happens when verification is assumed?

Scholarly arguments are embedded in chains of reference.

Not every citation performs the same function. Some provide evidence. Others establish context, attribute an idea, acknowledge prior work or identify disagreement.

But references connect current work to the literature around it.

When an incorrect or nonexistent reference enters that chain, it can acquire the appearance of legitimacy simply by appearing inside a scholarly work.

Later authors may encounter it and assume that someone earlier in the process checked it.

They may copy it.

Reference managers may import it.

AI systems trained on, indexed against, or retrieving from the scholarly record may encounter it.

The error can move further from the point at which it originated.

The large-scale Lancet audit illustrates the persistence problem. More than 98% of the papers containing fabricated references identified by the researchers had received no publisher action at the time of their audit in February 2026, according to reporting based on the study and associated databases.

This does not mean that one incorrect citation invalidates an entire paper.

Nor does verifying a citation establish that the cited source supports the author's claim.

Citation verification and claim verification are different tasks.

But errors can persist and propagate when the existence and identity of the underlying source are assumed rather than checked.

The missing verification boundary

None of this is an argument against peer review.

Peer review remains a central quality-control mechanism in scholarly communication.

The problem is expecting it to perform a function it was never consistently designed to perform.

A reviewer evaluating a complex manuscript may already be assessing methodology, interpretation, originality and disciplinary significance.

Expecting the same reviewer to manually verify dozens or hundreds of bibliographic records may not be realistic.

The more useful question is whether scholarly workflows now need a clearer citation-verification boundary.

At some point, someone has to establish that the citation being used corresponds to a real source and that its supplied bibliographic information is accurate.

That requirement does not disappear as technology improves.

Models will improve.

Retrieval systems may reduce some errors.

Reference-management tools may add better safeguards.

Different systems will fail in different ways.

The durable principle remains:

Whatever produced the citation, somebody must verify the citation.

It does not matter whether the reference came from memory, a colleague, a research assistant, a reference manager, a language model or a retrieval system.

The origin of the citation does not remove the need for verification.

And that verification boundary needs to be defined carefully.

Checking whether a source exists and whether its bibliographic metadata is accurate does not establish that the source supports a particular claim.

It does not validate the methodology of the cited research.

It does not determine whether the paper is scientifically correct.

Those are separate questions.

Citation verification answers a narrower one:

Is the scholarly work being cited actually the work the reference says it is?

AI has transformed the economics of citation production without creating an equivalent verification boundary.

That leaves academic publishing with a practical question:

When everyone is responsible for citation integrity, who actually checks?

Until that question has a clear operational answer, peer review may continue doing exactly what it was designed to do while citation errors pass through a gap that belongs to no one.

Closing that gap does not require redefining peer review.

It requires defining citation verification as a distinct task, deciding where it belongs in the scholarly workflow, and assigning someone responsibility for ensuring that it actually happens.

Custom HTML/CSS/JavaScript
Sean Honan

Sean Honan

Sean Honan writes about citation risk, AI-generated references, and academic source verification for Citation Risk. He focuses on helping editors, thesis coaches, and writers catch fabricated, mismatched, or incomplete citations before submission.

LinkedIn logo icon
Back to Blog