Who Checks the Citations? Benchmarking Legal Hallucination Detection
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Who Checks the Citations? Benchmarking Legal Hallucination Detection".
Jane: Text A appears to be an excerpt from a technical research paper concerning AI systems for detecting legal citation hallucinations, while Text B is entirely unrelated content,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we're talking about this paper called "Who Checks the Citations? Benchmarking Legal Hallucination Detection." It looks like they're trying to figure out if AI can actually catch when those systems make up legal citations.
Jane: Yeah, exactly. The title itself says it’s about checking citations, and they want to see if we can build tools that verify what the AI spits out in legal documents.
Lu: It’s interesting because they aren't just guessing; they created a dataset of one thousand three hundred brief excerpts where they intentionally injected different kinds of citation errors into real court filings <ref:2606.21155#pg0,a dataset of 1,300 brief excerpts>.
Meng: So it’s like building a test kitchen for the AI systems to see how well they handle bad citations. How does this data help us actually measure the problem?
Tom: Well, what’s really compelling is their taxonomy of these hallucinations. They didn't just list errors; they grounded that whole system in actual court filings.
Jane: That means they know exactly what kinds of mistakes lawyers and judges are making when they use AI, which helps them build a better detection framework.
Lu: The study found that even the most advanced models, like GPT-five still struggle with certain types of errors <ref:2606.21155#pg0>. Specifically, content misrepresentation seems to be the hardest thing to spot because it can really mess with a legal decision.
Tom: That makes sense; if you change the substance of a citation rather than just the name, that’s where things get complicated for the AI. What about how many models they tested?
Meng: They benchmarked five different large language models, testing them both in ways that involve some kind of agentic system and systems that aren't. That gives them a good look at how different approaches perform under pressure.
Jane: And when they looked at the results, GPT-five actually scored quite well, hitting an eighty-four point four percent recall and a fifty-five point zero percent F1 score when it was running in that agentic setting <ref:2606.21155#pg0,84.4% recall and a 55.0% F1 score>.
Tom: But there are some real limitations they pointed out, and I think those are important for us to keep in mind. They mentioned that these verification processes can be quite demanding on computing resources.
Lu: Yeah, they found that the agentic verification process itself is resource-intensive; for example, GPT-five took an average of fifteen point three steps just to verify one excerpt <ref:2606.21155#pg0>. That’s a lot of work for a single piece of text.
Jane: And another big hurdle they pointed out is that even with all this testing, the models still find it hard to reliably spot those more subtle errors in the citation data.
Tom: It seems like there's a gap between what the models *can* do and what they *can* reliably do when it comes to these tricky mistakes. What else did they look at in terms of long-term trends?
Jane: They looked at an eight-year trend, tracking eight different ChatGPT generations from late two thousand twenty-three all the way through late two thousand twenty-five <ref:2606.21155#pg0>. And what they found was that the hallucination rates haven't actually gone down consistently.
Meng: That’s a big warning sign for us; it suggests this isn't just a temporary glitch in early AI; the problem is sticking around and even getting bigger.
Tom: It ties back into the idea that court filings are growing, and the number of citations per filing is also going up. They argue these trends won't fix themselves on their own.
Lu: Their main suggestion for moving forward seems to be focusing on two structural fixes instead of just better models. They said we need to increase how much complete legal citation data is publicly available.
Jane: And they also stressed the need for focused guidance around how we deploy AI in legal settings so that people know what to expect from these tools.
Tom: So, the core message is that automated verification isn't a complete solution on its own; it has to be part of a bigger strategy. What does this mean for the folks who actually use these systems?
Meng: It means we need better ways to access the data itself if we want anything practical to happen for judges and lawyers.
Jane: It points toward making AI literacy in this area a necessary skill, not just an optional extra when using these tools.
Tom: Well, that brings us to the end of our discussion on "Who Checks the Citations? Benchmarking Legal Hallucination Detection." We've seen that while models are getting better at catching obvious errors, the real fight is with those subtle misrepresentations and accessing clean data.
Lu: It really highlights how crucial it is to ground AI verification in real-world court filing patterns.
Jane: And we’ll be looking at how these findings shape our next set of research into reliable AI systems for legal tasks.
Meng: I'm curious to see if we can actually build those improved detection systems that they suggest, especially with that focus on agentic verification you mentioned earlier.
Tom: That's what we’ll be looking at next. We’ve covered the paper "Who Checks the Citations? Benchmarking Legal Hallucination Detection."
The paper's summary: Tom: So we’ve seen how they built this whole testing setup for citation errors using that LEPHANTOMCITE dataset, and now we need to talk about what their actual findings mean for legal AI.
Jane: Right, so the core idea is that they didn't just look at a few mistakes; they created this detailed breakdown of every type of hallucination possible in legal citations.
Lu: They’ve mapped out a whole system for spotting errors, and the most worrying part is that even the biggest models, like GPT-five still miss some of these trickier mistakes.
Meng: What kind of errors are they struggling with specifically? Is it just simple typos or something more complex?
Tom: They flagged content misrepresentation as the hardest one to catch because that’s where the AI can seriously twist a legal outcome. Plus, incorrect pincites and verbatim misquotes are still big problems they found.
Jane: It sounds like they found that if the AI changes what a citation actually *says* instead of just messing up the format, it’s much harder to detect.
Lu: And they showed how five different large language models perform when you test them in different ways, but they pointed out that even GPT-five hits limits when we push it into agentic verification settings.
Meng: Can you tell us about those resource demands? We’re talking about real-world deployment here, and if the checking process takes fifteen steps for every single excerpt, that’s a lot of work.
Tom: It is intensive, yeah. That’s a real practical hurdle they identified. But then there's this structural issue that stops them cold: the way official legal documents are paginated online isn't complete enough to let even the best agents do their job perfectly.
Jane: So it’s not just about making the AI smarter, it’s about fixing the data itself so the AI has a complete picture to work with.
Lu: They also looked at how this problem has been changing over time, tracking eight different versions of ChatGPT for years, and they found that these hallucination rates haven't dropped consistently at all.
Meng: That’s a tough look for us. It suggests the current path of just making models slightly better isn't going to solve this; the underlying problem is sticking around and getting worse with more paperwork.
Tom: Exactly. The authors aren’t suggesting that if we just train another model, we’ll be fine in a few months because they see an eight-year trend showing no improvement.
Jane: So what does this mean for the folks who actually rely on these AI tools in their day-to-day work?
Lu: It means automated checking tools by themselves aren't enough; they need to be part of a bigger strategy that addresses how we get and use legal data.
Meng: I think it points toward making access to complete legal citation data a top priority for any system trying to be reliable in this area.
Tom: And for anyone using these systems, the implication is that we need better guidance on how to deploy AI responsibly in the legal world because the tools aren't ready just yet.
Jane: So, it’s not about waiting for a perfect model; it’s about fixing the environment around those models instead.
The paper's improvements: Tom: We’ve seen how they set up the initial testing, and now we need to talk about what they actually propose to fix these citation hallucinations.
Jane: So, instead of just saying "the AI is bad," they suggest three major ways to make the whole system more reliable for checking citations.
Lu: The first big suggestion is using agentic systems like BOED, which means the AI doesn't just look at one thing and stop; it has to perform sequential verification steps, updating its understanding after every action.
Meng: Sequential verification sounds like a lot of computation. How does that specific agentic process actually help it catch more errors than a simple check?
Tom: It helps because the agent can do a thorough search and explore other sources if the first lookup doesn't work out, which is better for catching those tricky things they mentioned earlier.
Jane: They also want to move away from just using one fixed list of errors. They propose a dynamic taxonomy that adapts based on actual court filings, maybe even adding labels for stylistic choices to catch subtler misrepresentations.
Lu: That makes sense, because a citation can be technically correct but stylistically misleading, and their current system doesn't account for that nuance well enough.
Meng: So it’s about making the error detection smarter by feeding it more context from real legal documents instead of relying on a static checklist.
Tom: And they want to use the results from those agent trajectories to diagnose *why* a model fails, like figuring out if GPT-five spends too much time looking at local opinions versus actually verifying the core text.
Jane: That’s really interesting because it moves us past just knowing *what* went wrong to understanding *how* the AI's internal process is flawed.
Lu: If we can do that, we could start building better tools that don't just flag errors but actually learn from their own mistakes across different model architectures.
Meng: That’s a solid engineering goal; diagnosing failure modes gives us actionable data for developers instead of just knowing the final score.
Tom: And they’re still pushing for better data access, which is a huge structural change they want to see happen before these detection tools become fully useful in the long run.
Jane: So what this suggests is that we need a two-pronged approach: smarter AI verification methods paired with real improvements in how we can get clean legal data.
Conclusion: Tom: So we’re wrapping up our deep dive into "Who Checks the Citations? Benchmarking Legal Hallucination Detection." Essentially, they showed that just having bigger models isn't enough to solve this citation problem.
Jane: That's right. The whole paper boils down to this: if we want reliable legal AI, we have to treat citation checking as a mandatory first step, not an optional afterthought.
Lu: I think the real big picture here is that it forces us to rethink how we build trust in any system that uses language for high-stakes tasks.
Meng: It suggests that practical deployment depends more on fixing the data infrastructure than just tweaking the model weights themselves.
Lalam: From my side, this work reinforces my vision by showing how even seemingly small errors in foundational data can create massive inconsistencies in the final output we generate.
Tom: Exactly. They’re calling for better access to complete legal citation records as a necessary structural fix for the whole system.
Jane: And they need more responsible deployment guidance so people know exactly what level of accuracy to expect from these tools right now.
Lu: It opens up a lot of avenues for future research, especially if we can build those improved agentic verification systems they mentioned wanting to develop.
Meng: I’m looking forward to seeing how the engineers take that framework and turn it into something that runs reliably on a production server without crippling resource demands.
Lalam: If we can solve these underlying issues, it could really improve the culture of legal tech by making the output genuinely dependable for everyone involved.
Patty Liu, Dominik Stammbach, Peter Henderson
Princeton University
cs.CL
Submitted: 2026-06-19
Updated: 2026-10-04
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 87/100
The gist: Text A appears to be an excerpt from a technical research paper concerning AI systems for detecting legal citation hallucinations, while Text B is entirely unrelated content, appearing to be excerpts
Key concepts
- LEPHANTOMCITE benchmarking dataset
- This is a custom collection of 1,300 legal excerpts intentionally filled with various citation errors. It was created to provide a real-world testing ground for new AI systems. By using this specific data, researchers can accurately measure how well different detection tools perform when faced with common citation mistakes found in actual court documents.
- Taxonomy of legal citation hallucinations
- This is a detailed classification system for different types of errors that occur when AI generates citations. The study categorized these errors, such as incorrect pincites or verbatim misquotes, to understand exactly what kinds of mistakes models struggle with most. This helps researchers pinpoint where the AI verification tools need the most improvement.
- Agentic operational settings
- This refers to how an AI system performs a task when it operates autonomously and takes multiple steps to verify information. In this study, agentic settings were tested alongside non-agentic ones. The research showed that while these advanced methods offer high recall, they require significant computational resources and time for each check.
- Content misrepresentation
- This is identified as the most difficult type of citation error for AI to detect accurately. It means the AI generates a citation that looks correct but fundamentally misrepresents the actual legal content or context of the source material. This type of error poses the greatest risk because it can significantly distort legal outcomes.
Terminology
Summary
Text A appears to be an excerpt from a technical research paper concerning AI systems for detecting legal citation hallucinations, while Text B is entirely unrelated content, appearing to be excerpts from a legal case or trial transcript discussing physical ailments and medical history.
My task is to synthesize these two disparate pieces of information into one long, detailed summary of the paper Who Checks the Citations? Benchmarking Legal Hallucination Detection,
as described in Text A.
Here is the comprehensive summary:
Detailed Research Summary: Benchmarking Legal Hallucination Detection
This research study, titled Who Checks the Citations?
, investigates the efficacy of AI-based systems in mitigating legal citation hallucinations by developing automated detection mechanisms. The core objective is to establish a robust framework for building and auditing reliable tools capable of verifying legal citations, recognizing that such verification is a critical, first-class problem in responsible AI deployment within the legal domain.
Methodology and Dataset:
The authors introduced a novel methodology centered around creating an evaluation framework grounded in actual court filings. To facilitate this, they developed the LEPHANTOMCITE benchmarking dataset, which comprises 1,300 brief excerpts intentionally injected with various types of citation errors. This dataset serves as the empirical foundation for testing the proposed detection systems.
Taxonomy and Error Analysis:
A key contribution of the study is the proposal of a comprehensive taxonomy of legal citation hallucinations, which is meticulously grounded in real court filings. The analysis revealed that while advanced models like GPT-5 achieve high recall across all hallucination types, they still struggle significantly with subtle error categories. Specifically, the most challenging errors to detect remain content misrepresentation, as this type of hallucination has the highest potential to distort legal outcomes. Other problematic categories identified include incorrect pincites and verbatim misquotes.
Model Benchmarking and Performance:
The study benchmarks five different Large Language Models (LLMs) across both agentic and non-agentic operational settings. The results indicate that the latest iterations, exemplified by GPT-5, demonstrate superior performance—achieving an 84.4% recall and a 55.0% F1 score within an agentic framework. However, the research highlights significant limitations:
-
Resource Intensity: Agentic verification processes are inherently resource-intensive; for instance, GPT-5 averages 15.3 steps per excerpt during its verification process.
-
Subtle Error Limitations: Despite advancements, all tested models exhibit difficulty in reliably detecting the more nuanced and subtle error categories within the citation data.
-
Data Accessibility Barriers: A structural barrier limiting current automated verification capabilities is the incompleteness of official pagination within publicly available legal repositories, which restricts what even the most sophisticated agents can effectively verify.
Longitudinal Evidence and Trend Analysis:
The study provides crucial longitudinal empirical evidence demonstrating that legal citation hallucination rates are not a temporary artifact of early LLMs. Over an eight-year period, tracking eight different ChatGPT generations from late 2023 through late 2025, the authors found that hallucination rates have not consistently declined. Furthermore, the burden on the judiciary is escalating along two compounding dimensions: an increasing volume of court filings and a steady rise in citations per filing. The authors strongly argue that these trends will not self-correct as models continue to improve, necessitating a fundamental shift in focus toward verification.
Conclusion and Recommendations:
The overarching conclusion of the work is that automated verification tools alone are insufficient without addressing underlying systemic issues. The study concludes by proposing structural paths forward:
-
Improving Data Accessibility: Increasing public access to complete legal citation data would substantially reduce the current burden on courts and litigants.
-
Targeted AI Literacy: Guidance focused on responsible deployment of AI in legal contexts is necessary to ensure reliability.
Ultimately, the research asserts that responsible deployment of AI in legal contexts requires treating citation verification as a first-class problem. The LEPHANTOMCITE benchmark serves as a vital foundation for future efforts aimed at building, benchmarking, and auditing reliable legal citation checking tools needed by courts and litigants alike.
Improvements for AI systems
- Bold header: Agentic Verification for Citation Accuracy
This improvement involves leveraging agentic systems like BOED to perform sequential, information-dependent verification, as case citation verification is inherently sequential and information-dependent.
This allows the system to maintain an explicit, language-based belief state that is updated after each action,
which helps in verifying citations
by performing a thorough search
and exploring alternative sources when initial lookups fail.
- Bold header: Enhanced Error Analysis and Failure Mode Identification
The system can be improved to diagnose specific failure modes by analyzing agent trajectories, such as identifying that GPT-5 devotes 37.3% of its actions to SEARCH LOCAL OPINION,
which suggests stronger models continue into the opinion text to verify quotes and holdings thoroughly.
This enables developers to understand why certain hallucinations persist across different model architectures.
- Bold header: Context-Aware Hallucination Taxonomy
Instead of relying solely on a fixed taxonomy, the system should dynamically adapt by using a taxonomy grounded in actual court filings
and incorporating optional ground-truth labels to include stylistic choices,
which helps in identifying subtle misrepresentations that are currently challenging for models.
Abstract
Attorneys, judges, and pro se filers increasingly use AI to draft legal documents, yet these tools frequently fabricate citations. Despite predictions that newer models would hallucinate less or that court sanctions would deter negligent filers, we found over 1,000 filings containing fabricated citations---with this number growing year-over-year. This study evaluates whether AI-based systems can mitigate these errors by automatically detecting hallucinations. We propose a taxonomy of legal citation hallucinations grounded in actual court filings and introduce a dataset of 1,300 brief excerpts containing injected errors. Benchmarking five models in agentic and non-agentic settings as well as Claude Code reveals that while the latest iterations perform better---GPT-5 achieves 84.4% recall and a 55.0% F1 score in an agentic framework---all models struggle with subtle error categories. Agentic verification remains resource-intensive, with GPT-5 averaging 15.3 steps per excerpt. Furthermore, restricted information access limits the efficacy of even the best agents. This gap creates policy concerns, as it disadvantages both AI systems and litigants who lack subscriptions to commercial legal databases. Together, our dataset, tools, and policy recommendations provide a foundation for building and auditing reliable legal citation checking tools.
Sources
- PURR: Efficiently Editing Language Model Hallucinations by Denoising Language Model Corruptions
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models
- How Ready are Pre-trained Abstractive Models and LLMs for Legal Case Judgement Summarization?
- HalluHard: A Hard Multi-Turn Hallucination Benchmark
- Agentic Forecasting with Structured Linguistic Beliefs
- gpt-oss-120b & gpt-oss-20b Model Card
- olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
- OpenAI GPT-5 System Card
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering