Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models
summary
The gist
This paper evaluates whether large language models (LLMs) can automate compliance with *The Bluebook: A Uniform System of Citation*, a 500+ page manual of complex legal citation rules.
In short
The episode reviews a paper testing if Large Language Models can automate legal procedure using Bluebook citations. Testing five AI models on 866 tasks, accuracy ranged from 69% to 74%. The hosts conclude that AI is not ready for legal work because the models rely on memorization rather than applying rules, making them unreliable for complex procedural tasks.
Key concepts
- The Bluebook
- A massive manual detailing how to format citations in legal documents. It serves as a proxy for broader legal procedure because it is highly structured, intricate, and demands strict adherence to specific rules.
- Legal Procedure
- The set of rules regarding how things are filed and handled within the law. The paper uses the Bluebook's structure to test if AI can handle these complex procedural requirements beyond simple citation formatting.
- Memorization vs. Rule Application
- The study found that models performed better on real cases than synthetic ones, indicating they were merely regurgitating training data rather than applying the actual rules of the Bluebook in a practical, legal sense.
- In-Context Learning
- The test where a long document containing the Bluebook rules was provided to AI models. While some improvement was seen (a six-point increase), this learning process struggled to synthesize and apply complex rules across multiple sections.
Terminology used across episodes
This episode discusses
The paper
Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models · Read on arXiv
Matthew Dahl, Eric Martínez
Yale Law School
Legal practice requires careful adherence to procedural rules. In the United States, few are more complex than those found in The Bluebook: A Uniform System of Citation. Compliance with this system's 500+ pages of byzantine formatting instructions is the raison d'etre of thousands of student law review editors and the bete noire of lawyers everywhere. To evaluate whether large language models (LLMs) are able to adhere to the procedures of such a complicated system, we construct an original dataset of 866 Bluebook tasks and test flagship LLMs from OpenAI, Anthropic, Google, Meta, and DeepSeek. We show (1) that these models produce fully compliant Bluebook citations only 69%-74% of the time and (2) that in-context learning on the Bluebook's underlying system of rules raises accuracy only to 77%. These results caution against using off-the-shelf LLMs to automate aspects of the law where fidelity to procedure is paramount.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models".
Jane: The paper was written by Matthew Dahl and Eric Martínez from Yale Law School.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. Today we're looking at a paper that's got a title that just makes you smile — "Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models." And I gotta say, as someone who spent way too many hours in law school wrestling with that little blue book, this title hits close to home.
Jane: Oh, Tom, I remember those days. The Bluebook is this massive manual — over five hundred pages — that tells lawyers exactly how to format citations in legal documents. Every comma, every period, every abbreviation has a rule. And for decades, law students and paralegals have been the ones stuck making sure every single citation follows those rules perfectly.
Tom: Right, and it's brutal. I mean, we're talking about rules for when to italicize a comma. A comma, Jane. Not even the word — the actual punctuation mark. So when this paper from Matthew Dahl at Yale Law School came out asking whether large language models can handle this stuff, I was genuinely curious.
Jane: And the answer, at least from the title, seems to be a cautious "not yet." But here's what I love about this paper — it's not just testing whether AI can memorize citation formats. It's using the Bluebook as a stand-in for something much bigger: legal procedure itself.
Tom: That's a great point. Legal procedure is all these rules about how you file things, when you file them, what format they need to be in. And if AI can't handle something as structured as the Bluebook, how can we trust it with more complex procedural rules?
Jane: Exactly. And the paper makes this really clever connection. They argue that legal citation is actually a form of legal procedure because it has three key properties — it's universal, meaning it applies to all legal arguments regardless of the subject; it's intricate, with all these layers of exceptions; and it's inflexible, demanding strict adherence.
Tom: So they built a dataset of eight hundred sixty-six Bluebook tasks and tested five major AI models — GPT-four point one, Claude three point five Sonnet, Gemini two point five Flash, Llama three point one 405B, and DeepSeek V3. And the results? Well, let's just say the Bluebook isn't going anywhere just yet.
Jane: Right, the best model only hit seventy-four percent accuracy. But before we get into those numbers, I want to flag something that really struck me. The paper's author is at Yale Law School, not a computer science department. That's a sign that legal scholars themselves are starting to seriously evaluate these tools.
Tom: And that's exactly the kind of interdisciplinary work we need. So stick around — next we're going to break down what those models actually got right and where they completely fell apart.
Summary: Tom: So Jane, we've established that the models aren't perfect. But let's get into the specifics, because the breakdown is actually fascinating. The paper found that across all five models, accuracy ranged from sixty-nine percent to seventy-four percent in what they call a zero-shot setting — meaning the models were given no examples, just asked to produce Bluebook-compliant citations.
Jane: And that's already pretty concerning for anyone thinking about automating legal work. But here's the really interesting part — the models weren't equally bad at everything. They were actually quite good at some tasks and terrible at others.
Tom: Right, let's talk about that. For case law citations — you know, citing court decisions — the models did pretty well on the core components. They got the volume, reporter, and page combinations right ninety-six percent of the time. And court and date abbreviations? ninety percent accuracy. That's the bread and butter of legal citation.
Jane: But then you look at the other tasks and it falls apart. State statutes? Only thirty-six percent accuracy. Electronic statutes — those are laws accessed through databases like LexisNexis — just ten percent. And secondary sources like books and law review articles? twenty-nine percent. Those are terrible numbers.
Tom: And what's wild is that the paper includes this memorization analysis that really makes you question even the good results. They took real cases and then created synthetic versions with fake page numbers. And guess what? The models did better on the real cases than the fake ones.
Jane: That means the models weren't actually applying the Bluebook rules — they were just regurgitating citations they'd memorized from their training data. Only two models, Gemini and DeepSeek, didn't show that memorization behavior.
Tom: Which is a huge red flag. Because in real legal practice, you're constantly citing cases that the model might not have seen before. And if it can't apply the rules to new cases, it's not really following procedure — it's just pattern matching.
Jane: And the errors weren't trivial either. The paper shows examples where models completely invented citations, confused the parties in a case, or misidentified the court. One model even cited a case that doesn't exist. In legal practice, that's not just an inconvenience — it could get you sanctioned by a judge.
Tom: Let's bring in Lu and Meng on this, because I think they'll have some strong opinions about what this means for AI development.
Lu: Thanks, Tom. I think this paper is really important because it shows that rule-following is a distinct capability from general language understanding. These models can write beautiful prose about the law, but when it comes to executing precise, multi-step procedures, they fall short. That's a fundamental limitation, not just a tuning issue.
Meng: I'd push back slightly on that, Lu. The memorization finding suggests that some of the failures come from how the models were trained, not necessarily from an inherent ceiling. But you're right that the procedural reasoning itself is the bottleneck. From an engineering standpoint, you'd need to build a system that explicitly separates the rule-following from the language generation.
Jane: That's a really useful distinction. So the models aren't just dumb — they're applying the wrong kind of intelligence. They're using memory when they should be using rules. And that's a much harder problem to solve.
Tom: And it's a problem that matters beyond just citations. Next, we're going to talk about whether giving the models the actual rulebook helps — spoiler alert, it doesn't help as much as you'd hope.
Improvements: Tom: So here's the natural question — if the models don't know the Bluebook rules, can we just give them the rules? That's exactly what the paper tested next, and Jane, I think you'll appreciate the setup.
Jane: Oh, I love this part. They took the Indigo Book — that's a public-domain version of the Bluebook — and fed it to Gemini two point five Flash, which has this massive one-million-token context window. The rules run about ninety thousand tokens, so it fits comfortably. And they told the model, "Here are the rules, now follow them."
Tom: And the results? Accuracy went from seventy-one percent in the zero-shot setting to seventy-seven percent with the rules provided. So a six-point improvement. That's something, but it's nowhere near the level you'd need for real legal work.
Meng: I have to say, that's actually a pretty damning result for the whole "long context" narrative. These companies are advertising million-token context windows, but if the model can't effectively use ninety thousand tokens of rules to improve its performance, what's the point?
Lu: It's worse than that, Meng. The paper points out that this is part of a growing literature showing that long-context models struggle with what they call "long in-context learning." They can retrieve facts from long documents, but they can't synthesize and apply complex rules from them. It's like giving someone a legal textbook and expecting them to pass the bar exam after one reading.
Jane: And that's the real-world implication here. Because the Bluebook is actually on the shorter end of legal procedure documents. The Federal Rules of Civil Procedure are about seventy-four thousand tokens. The Federal Rules of Bankruptcy Procedure are ninety-four thousand tokens. These are comparable lengths, and if the model can't handle the Bluebook, it's not going to handle those either.
Tom: But here's what I find interesting — the improvement wasn't uniform. The paper shows that in-context learning helped a lot on some tasks. Statutes, secondary sources, signals — those all improved noticeably. But other tasks barely moved.
Lu: Right, and that tells us something about the nature of the tasks. The tasks that improved are ones where the rules are relatively self-contained. The tasks that didn't improve — like case names and court abbreviations — are ones where the rules require applying judgment across multiple sections of the manual. The model can't hold all those cross-references in mind simultaneously.
Meng: So the practical takeaway for anyone building legal tech is that you can't just stuff the rulebook into the prompt and expect magic. You'd need to build a system that retrieves the specific relevant rules for each citation task, not one that tries to process everything at once.
Jane: That's a really concrete engineering insight. And it suggests that the future of legal AI isn't about bigger context windows — it's about smarter architectures that know which rules to apply when.
Tom: And that's a much harder problem. But before we wrap up, I want to get Lalam's take on what this means for the broader picture of AI in professional settings.
Conclusion: Tom: Alright, we've covered the results, the methodology, the memorization problem, and the in-context learning limitations. Let's bring it all together for "Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models."
Jane: The bottom line — and I know we're not supposed to say that, but it fits here — is that these models are not ready to replace human Bluebook compliance. The best model hit seventy-four percent accuracy, and even with the rules provided, it only reached seventy-seven percent. For legal work, where a single citation error can get a filing rejected, that's just not good enough.
Tom: And it's not just about the Bluebook. The paper makes a compelling case that the Bluebook is a proxy for legal procedure more broadly. If AI can't handle this relatively contained rule system, we should be very skeptical about claims that AI can handle the full complexity of legal practice.
Lu: I think the deeper lesson is about the difference between knowing and doing. These models know a lot about the law — they can discuss legal concepts fluently. But procedural compliance is about doing, not knowing. It's about executing a series of precise steps under constraints. And that's a capability that's still very much underdeveloped.
Meng: From a practical standpoint, this paper gives engineers a clear roadmap of what needs to improve. We need models that can apply rules to novel inputs without memorization. We need architectures that can effectively use long rule documents. And we need evaluation benchmarks like this one to measure progress.
Lalam: If I may add — what excites me about this paper is that it's a step toward making AI systems more trustworthy in high-stakes domains. The legal profession is conservative for good reason. Showing where AI fails is just as valuable as showing where it succeeds. This kind of rigorous evaluation builds the foundation for eventual adoption, even if that adoption is years away.
Jane: That's a great note to end on. The paper isn't saying "never" — it's saying "not yet, and here's exactly why." That's the kind of honest assessment the field needs.
Tom: So, bye-bye to the Bluebook? Not quite. But this paper gives us a clear picture of what needs to happen before we can say that for real. Thanks for joining us, and we'll see you next time with another paper.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language