Is this Citation on Point?
Apurv Verma
Bloomberg
cs.DL, cs.CL
Submitted: 2026-08-12
Updated: 2026-08-14
Comments: Accepted to the 1st Workshop on AI for Law at ICML 2026
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 75/100
The gist: This paper studies proposition-level citation support verification in legal documents, focusing on whether current LLMs can detect when a legal citation points to a real case but does not support the
Terminology
Summary
This paper studies proposition-level citation support verification in legal documents, focusing on whether current LLMs can detect when a legal citation points to a real case but does not support the proposition for which it is offered. The authors distinguish this from fabrication detection, noting that the harder problem is detecting citations that point to real cases but do not support the propositions for which they are offered—a failure mode that existing evaluations of LLMs for legal use cases largely overlook.
The paper formalizes the task: Given a passage of legal text and the content of a cited document, the task is to determine whether the citation is on point. That is, does the cited authority actually support the proposition for which it is cited?
The authors develop a three-category taxonomy of citation roles—substantive, procedural, and secondary—in consultation with legal experts, and restrict evaluation to substantive citations, which assert a legal rule and offer a case as its basis.
The methodology uses controlled perturbations of real legal citations from two corpora: CLERC (court opinions, 2,000 citations) and BriefMe (legal briefs, 750 citations). Corruptions are constructed at three difficulty levels: Easy (replacing the citation with one from an entirely different legal document), Medium (replacing with a different citation from the same source document, only available for CLERC), and Hard (changing only the pinpoint page number within the same case). Human validation confirms high agreement between annotators and our heuristic labels (84–88% across both datasets).
Fourteen frontier model configurations from three families are evaluated: GPT-4o, GPT-4o-mini, GPT-4.1, GPT-4.1-mini, GPT-5, GPT-5-mini, GPT-5.4 (with and without high reasoning effort), Claude Sonnet 4, Claude Sonnet 4.6, Claude Opus 4.6, Gemini 2.5 Flash, Gemini 2.5 Pro, and Gemini 3.1 Pro. Metrics are Recall on corrupted citations and False Positive Rate on valid citations.
Main results show: Models are near-saturated on Easy examples, but recall on Hard examples falls to 36.5–60.6% on court opinions and 51.5–82.7% on briefs. No model achieves both low FPR and high recall on Hard examples.
Specifically, Models catch 93–100% of wrong-case corruptions. They catch only 37–61% of wrong-pinpoint corruptions on court opinions and 52–83% on legal briefs.
Regarding scale and reasoning: Within the OpenAI family, GPT-5 improves recall on Hard examples over GPT-4.1 on both CLERC (55.5% vs. 38.4%) and BriefMe (77.0% vs. 51.5%).
However, the metrics do not increase monotonically with model capability. GPT-5.4 without reasoning trails GPT-5 on both datasets. Claude Opus 4.6 also trails both Sonnet variants in Hard-example recall.
Extended reasoning helps: Adding high reasoning effort to GPT-5.4 increases recall on Hard examples from 48.5% to 59.8% on CLERC and from 64.0% to 82.0% on BriefMe,
but GPT-5.4 with high reasoning effort still misses 40% of pinpoint mismatches on court opinions.
Regarding document type: recall on Hard examples is 12–26 pp higher on briefs than on court opinions across all models.
The authors hypothesize that briefs are advocacy documents and often state narrower propositions, so a wrong page creates a sharper mismatch than it does in a more discursive opinion.
Error analysis reveals three recurring patterns: invented support rationales
(in roughly two-thirds of missed corruptions, the model claims the cited page expressly states
or explicitly says
text that is not there), topical matching instead of propositional verification
(models treat topical overlap as sufficient), and failure to check quoted text
(among missed Hard corruptions with verbatim quotes, the quoted language is absent from the cited page 92% of the time).
A page-grounded prompt intervention adds three verification steps: verbatim quote check, page-level verification, and distinguishing topic from support. Results show: Page-grounded prompting improves recall on Hard examples. Across all models, we see an improvement: 9.7–27.8 pp on CLERC and 6.5–35.5 pp on BriefMe.
However, across every model, the false positive rate rises by 1.0–24.7 pp across the two datasets. The additional instructions added to the prompt make the models more skeptical of all citations, not just the corrupted ones.
The paper concludes: Database lookup largely solves the fake citation problem. The harder scenario is a real case citation whose cited content does not support the proposition in the source text.
The authors state: Current models can be made more skeptical, but not selectively skeptical enough to distinguish valid from invalid citations at the proposition level.
They note that Recognizing the right legal topic and verifying support for the cited proposition are distinct capabilities, and current models conflate them.
Improvements for AI systems
Improvements to AI systems:
-
Add proposition-level citation verification as a distinct evaluation and training task. The system should be trained and benchmarked specifically on detecting real-case citations that do not support the stated proposition, not just fabricated citations. This requires a labeled dataset with three difficulty tiers (wrong-case, wrong-source, wrong-pinpoint) and a taxonomy of citation roles (substantive, procedural, secondary) to filter evaluation to substantive claims.
-
Implement a page-grounded verification module with explicit sub-checks. The system should, for each cited pinpoint, perform three verifications: (a) verbatim quote matching—check whether any quoted text in the passage appears on the cited page; (b) page-level content verification—confirm the cited page actually contains the asserted legal rule; (c) propositional support check—distinguish topical overlap from logical support, requiring the cited text to entail the proposition, not merely relate to it.
-
Decouple topical relevance from propositional entailment in the model’s reasoning. The system should be explicitly prompted to separate “this case is about the same area of law” from “this case establishes the specific rule cited.” This can be enforced via a two-stage output: first classify topic, then independently verify support, with a hard rule that topical overlap alone is insufficient for a positive verdict.
-
Add a selective skepticism mechanism with calibrated thresholds. The system should learn to increase scrutiny only for pinpoint-level mismatches and quoted-text absences, while avoiding a global rise in false positives. This requires a confidence score for each verification sub-check, and a decision rule that flags a citation as unsupported only when the page-level or quote-level check fails, not when the model is merely uncertain.
-
Incorporate a “quoted text absence” detector as a high-priority alarm. Since 92% of missed hard corruptions with quotes have the quoted language absent from the cited page, the system should automatically flag any citation where a verbatim quote in the source text cannot be located on the cited page, regardless of other topical or semantic signals.
-
Train on both court opinions and legal briefs separately, with domain-specific calibration. Because recall is 12–26 percentage points higher on briefs than opinions, the system should adjust its verification strictness based on document type—applying tighter pinpoint verification for discursive opinions where propositions are broader and mismatches are subtler.
-
Use extended reasoning only for hard examples, not uniformly. The system should dynamically allocate higher reasoning effort (e.g., multi-step verification) when it detects a pinpoint-level citation or a quoted passage, but use faster inference for easy wrong-case citations, to improve efficiency without sacrificing recall.
-
Add an error-pattern audit layer. After generating a verdict, the system should self-check for three known failure modes: (a) invented support rationales—verify that any claim like “expressly states” is backed by actual retrieved text; (b) topical matching—reject conclusions based solely on shared keywords or case subject matter; (c) quote verification—re-run the verbatim check if the model claims the quote exists.
What the improved AI system can do:
-
Detect wrong-pinpoint citations in court opinions with recall improved from 37–61% to above 80%, while keeping false positive rate below 10% on valid citations.
-
Correctly reject a citation where the cited page discusses the same legal topic but does not state the specific rule asserted, even when the case name is real and topically relevant.
-
Flag a citation as unsupported if a verbatim quote in the source text is absent from the cited page, with near-100% sensitivity on such cases.
-
Distinguish between “this case is about negligence” and “this case establishes the duty of care in this exact factual scenario,” refusing to accept the former as support for the latter.
-
Maintain high recall on easy wrong-case citations (93–100%) while selectively increasing scrutiny only for pinpoint-level mismatches, avoiding the global skepticism that currently inflates false positives.
-
Provide a transparent audit trail for each verdict, showing which sub-check (quote, page, proposition) failed, enabling human reviewers to verify the model’s reasoning.
Abstract
In 2023, a New York judge sanctioned two attorneys in Mata v. Avianca for filing a brief with hallucinated citations generated by ChatGPT. Such failures are largely caught by database lookups; the harder problem is detecting citations that point to real cases but do not support the propositions for which they are offered -- a failure mode that existing evaluations of LLMs for legal use cases largely overlook. In this paper, we study proposition-level citation support verification through controlled perturbations of real legal citations obtained from two legal corpora, either replacing the cited case or changing only the pinpoint page within the same case. We evaluate fourteen model configurations on the resulting examples. Models catch 93-100% of wrong-case corruptions. They catch only 37-61% of wrong-pinpoint corruptions on court opinions and 52-83% on legal briefs. When models fail to catch wrong-pinpoint corruptions, they accept the citation based on topical overlap rather than page-level support. Scale and extended reasoning narrow the gap but do not close it: GPT-5.4 with high reasoning effort still misses 40% of pinpoint mismatches on court opinions and 18% on briefs. Prompting the model to verify support at the cited page improves recall, but it also raises the false positive rate. Recognizing the right legal topic and verifying support for the cited proposition are distinct capabilities, and current models conflate them.
Sources
- Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models
- CLERC: A Dataset for Legal Case Retrieval and Retrieval-Augmented Analysis Generation
- LLM-as-a-Judge: Rapid Evaluation of Legal Document Recommendation for Retrieval-Augmented Generation
- Who Checks the Citations? Benchmarking Legal Hallucination Detection
- Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL Conferences
- BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs
- GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models
- From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
- CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era