HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers
Patrik Reizinger, Wieland Brendel
cs.CR, cs.AI, cs.LG
Submitted: 2026-07-20
Comments: 59 pages, 7 figures, 40 tables. Benchmark and code: https://github.com/rpatrik96/hallmark (v1.2.0); verification tool: https://github.com/rpatrik96/bibtexupdater (v1.5.0)
Code: https://github.com/rpatrik96/hallmark
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricated references: GPTZero found 53 papers with hallucinated
Terminology
Abstract
Large language models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricated references: GPTZero found 53 papers with hallucinated citations among NeurIPS 2025's accepted set. Rule- and LLM-based verifiers are emerging, but no shared benchmark compares them and gives detailed failure diagnostics. We close that gap with HALLMARK (Hallucination benchmark): 2,526 BibTeX entries spanning 14 hallucination types, three difficulty tiers, six diagnostic sub-tests per entry, and a contamination-resistant held-out split. On it we evaluate a DOI-lookup baseline, frontier LLMs zero-shot, tool-augmented agents, and our own rule-based, co-designed verifier bibtex-updater. Across the benchmark one result is consistent: the false-positive rate, not recall, decides whether a verifier is deployable. HALLMARK makes it concrete through three failure modes: agentic lookups buy recall but inflate false positives; at a venue-realistic base rate, the order-of-magnitude spread in false-positive rates (FPRs) -- not recall -- governs whether a verifier's flags are mostly true catches or mostly noise; and most LLMs over-flag papers published past their training cutoff, where only the two latest-cutoff models hold their false-positive rate near in-distribution levels (a signal we report as descriptive, since it is confounded with possible recall of those entries). Thus FPR is the deployment bottleneck, but an undetected fabrication remains the costlier error for the scientific record.
Sources
- CheckIfExist: Detecting Citation Hallucinations in the Era of AI-Generated Content
- Compound Deception in Elite Peer Review: A Failure Mode Taxonomy of 100 Fabricated Citations at NeurIPS 2025
- The Case of the Mysterious Citations
- Evaluating Large Language Models Trained on Code
- DeepSeek-V3 Technical Report
- HalluHard: A Hard Multi-Turn Hallucination Benchmark
- RefChecker: Reference-based Fine-grained Hallucination Checker and Benchmark for Large Language Models
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation
- Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
- HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
- HalluCiteChecker: A Lightweight Toolkit for Hallucinated Citation Detection and Verification in the Era of AI Scientists
- CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era
- GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models
- Qwen3 Technical Report
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs