AI citation hallucinations in research

Empirical evidence from 2025 indicates a significant and rapid rise in hallucinated citations in peer-reviewed literature, primarily driven by the integration of Large Language Models (LLMs) into scholarly writing. Verified Answer #1

A comprehensive audit of 2.5 million biomedical papers by Topaz et al. (2026) identified that the rate of papers containing at least one fabricated citation increased markedly between 2023 and 2025. Verified Answer #1

Specifically, the quarterly rate of fabricated citations per 10,000 papers rose from approximately 4 in 2023 to 56.9 by early 2026, representing a more than 12-fold increase (Topaz et al., 2026, https://doi.org/10.1016/S0140-6736(26)00603-3). Verified Answer #1

Large-scale audits of major scientific repositories, including arXiv, bioRxiv, SSRN, and PubMed Central, estimated that over 146,000 hallucinated citations were introduced in 2025 alone (Zhao et al., 2026, https://arxiv.org/abs/2605.07723). Verified Answer #1

Taxonomy of Citation Hallucinations: Research differentiates between two primary types of failure modes: 'wholly fabricated' references and 'real-source-but-misleading' citations. Verified Answer #1

A detailed taxonomy of 100 hallucinated citations in papers accepted at the NeurIPS 2025 conference (Ansari, 2026, https://arxiv.org/abs/2602.05930) classified errors into: 1. Verified Answer #1

Total Fabrication (66%): Citations referring to completely non-existent papers, authors, or DOIs; 2. Verified Answer #1

Partial Attribute Corruption (27%): Metadata is significantly altered; 3. Verified Answer #1

Identifier Hijacking (4%): Valid identifiers are paired with incorrect titles/authors; 4. Verified Answer #1

Placeholder Hallucination (2%); and 5. Verified Answer #1

Semantic Hallucination (1%): Real sources are cited but do not support the specific claim. Verified Answer #1

Comparison of Rates and Detection Bias: Empirical data shows that 'wholly fabricated' references (Total Fabrication) represent the dominant detected category in large-scale automated audits, often accounting for ~66% of identified hallucinations in conference settings (Ansari, 2026). Verified Answer #1

In contrast, citations that resolve to real sources but fail to support the claim (semantic hallucinations) appear as a smaller fraction (e.g., 1% in the Ansari taxonomy). Verified Answer #1

However, researchers highlight a critical detection bias: wholly fabricated references are readily identified through automated database cross-referencing (verifying if a DOI or title exists), while identifying 'real-source' inaccuracies requires complex semantic validation of whether the cited content supports the specific claim. Verified Answer #1

Consequently, while total fabrications are the most frequent detected failure, contextually inaccurate references (real but unsupported) remain an underreported, persistent challenge in AI-assisted research (Ansari, 2026; Topaz et al., 2026). Verified Answer #1