Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models

arXiv:2505.02763 · cs.CL, cs.AI, cs.CY · Submitted 2026-08-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models".

Jane: The paper was written by Matthew Dahl and Eric Martínez from Yale Law School.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. Today we're looking at a paper that's got a title that just makes you smile — "Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models." And I gotta say, as someone who spent way too many hours in law school wrestling with that little blue book, this title hits close to home.

Jane: Oh, Tom, I remember those days. The Bluebook is this massive manual — over five hundred pages — that tells lawyers exactly how to format citations in legal documents. Every comma, every period, every abbreviation has a rule. And for decades, law students and paralegals have been the ones stuck making sure every single citation follows those rules perfectly.

Tom: Right, and it's brutal. I mean, we're talking about rules for when to italicize a comma. A comma, Jane. Not even the word — the actual punctuation mark. So when this paper from Matthew Dahl at Yale Law School came out asking whether large language models can handle this stuff, I was genuinely curious.

Jane: And the answer, at least from the title, seems to be a cautious "not yet." But here's what I love about this paper — it's not just testing whether AI can memorize citation formats. It's using the Bluebook as a stand-in for something much bigger: legal procedure itself.

Tom: That's a great point. Legal procedure is all these rules about how you file things, when you file them, what format they need to be in. And if AI can't handle something as structured as the Bluebook, how can we trust it with more complex procedural rules?

Jane: Exactly. And the paper makes this really clever connection. They argue that legal citation is actually a form of legal procedure because it has three key properties — it's universal, meaning it applies to all legal arguments regardless of the subject; it's intricate, with all these layers of exceptions; and it's inflexible, demanding strict adherence.

Tom: So they built a dataset of eight hundred sixty-six Bluebook tasks and tested five major AI models — GPT-four point one, Claude three point five Sonnet, Gemini two point five Flash, Llama three point one 405B, and DeepSeek V3. And the results? Well, let's just say the Bluebook isn't going anywhere just yet.

Jane: Right, the best model only hit seventy-four percent accuracy. But before we get into those numbers, I want to flag something that really struck me. The paper's author is at Yale Law School, not a computer science department. That's a sign that legal scholars themselves are starting to seriously evaluate these tools.

Tom: And that's exactly the kind of interdisciplinary work we need. So stick around — next we're going to break down what those models actually got right and where they completely fell apart.

Summary: Tom: So Jane, we've established that the models aren't perfect. But let's get into the specifics, because the breakdown is actually fascinating. The paper found that across all five models, accuracy ranged from sixty-nine percent to seventy-four percent in what they call a zero-shot setting — meaning the models were given no examples, just asked to produce Bluebook-compliant citations.

Jane: And that's already pretty concerning for anyone thinking about automating legal work. But here's the really interesting part — the models weren't equally bad at everything. They were actually quite good at some tasks and terrible at others.

Tom: Right, let's talk about that. For case law citations — you know, citing court decisions — the models did pretty well on the core components. They got the volume, reporter, and page combinations right ninety-six percent of the time. And court and date abbreviations? ninety percent accuracy. That's the bread and butter of legal citation.

Jane: But then you look at the other tasks and it falls apart. State statutes? Only thirty-six percent accuracy. Electronic statutes — those are laws accessed through databases like LexisNexis — just ten percent. And secondary sources like books and law review articles? twenty-nine percent. Those are terrible numbers.

Tom: And what's wild is that the paper includes this memorization analysis that really makes you question even the good results. They took real cases and then created synthetic versions with fake page numbers. And guess what? The models did better on the real cases than the fake ones.

Jane: That means the models weren't actually applying the Bluebook rules — they were just regurgitating citations they'd memorized from their training data. Only two models, Gemini and DeepSeek, didn't show that memorization behavior.

Tom: Which is a huge red flag. Because in real legal practice, you're constantly citing cases that the model might not have seen before. And if it can't apply the rules to new cases, it's not really following procedure — it's just pattern matching.

Jane: And the errors weren't trivial either. The paper shows examples where models completely invented citations, confused the parties in a case, or misidentified the court. One model even cited a case that doesn't exist. In legal practice, that's not just an inconvenience — it could get you sanctioned by a judge.

Tom: Let's bring in Lu and Meng on this, because I think they'll have some strong opinions about what this means for AI development.

Lu: Thanks, Tom. I think this paper is really important because it shows that rule-following is a distinct capability from general language understanding. These models can write beautiful prose about the law, but when it comes to executing precise, multi-step procedures, they fall short. That's a fundamental limitation, not just a tuning issue.

Meng: I'd push back slightly on that, Lu. The memorization finding suggests that some of the failures come from how the models were trained, not necessarily from an inherent ceiling. But you're right that the procedural reasoning itself is the bottleneck. From an engineering standpoint, you'd need to build a system that explicitly separates the rule-following from the language generation.

Jane: That's a really useful distinction. So the models aren't just dumb — they're applying the wrong kind of intelligence. They're using memory when they should be using rules. And that's a much harder problem to solve.

Tom: And it's a problem that matters beyond just citations. Next, we're going to talk about whether giving the models the actual rulebook helps — spoiler alert, it doesn't help as much as you'd hope.

Improvements: Tom: So here's the natural question — if the models don't know the Bluebook rules, can we just give them the rules? That's exactly what the paper tested next, and Jane, I think you'll appreciate the setup.

Jane: Oh, I love this part. They took the Indigo Book — that's a public-domain version of the Bluebook — and fed it to Gemini two point five Flash, which has this massive one-million-token context window. The rules run about ninety thousand tokens, so it fits comfortably. And they told the model, "Here are the rules, now follow them."

Tom: And the results? Accuracy went from seventy-one percent in the zero-shot setting to seventy-seven percent with the rules provided. So a six-point improvement. That's something, but it's nowhere near the level you'd need for real legal work.

Meng: I have to say, that's actually a pretty damning result for the whole "long context" narrative. These companies are advertising million-token context windows, but if the model can't effectively use ninety thousand tokens of rules to improve its performance, what's the point?

Lu: It's worse than that, Meng. The paper points out that this is part of a growing literature showing that long-context models struggle with what they call "long in-context learning." They can retrieve facts from long documents, but they can't synthesize and apply complex rules from them. It's like giving someone a legal textbook and expecting them to pass the bar exam after one reading.

Jane: And that's the real-world implication here. Because the Bluebook is actually on the shorter end of legal procedure documents. The Federal Rules of Civil Procedure are about seventy-four thousand tokens. The Federal Rules of Bankruptcy Procedure are ninety-four thousand tokens. These are comparable lengths, and if the model can't handle the Bluebook, it's not going to handle those either.

Tom: But here's what I find interesting — the improvement wasn't uniform. The paper shows that in-context learning helped a lot on some tasks. Statutes, secondary sources, signals — those all improved noticeably. But other tasks barely moved.

Lu: Right, and that tells us something about the nature of the tasks. The tasks that improved are ones where the rules are relatively self-contained. The tasks that didn't improve — like case names and court abbreviations — are ones where the rules require applying judgment across multiple sections of the manual. The model can't hold all those cross-references in mind simultaneously.

Meng: So the practical takeaway for anyone building legal tech is that you can't just stuff the rulebook into the prompt and expect magic. You'd need to build a system that retrieves the specific relevant rules for each citation task, not one that tries to process everything at once.

Jane: That's a really concrete engineering insight. And it suggests that the future of legal AI isn't about bigger context windows — it's about smarter architectures that know which rules to apply when.

Tom: And that's a much harder problem. But before we wrap up, I want to get Lalam's take on what this means for the broader picture of AI in professional settings.

Conclusion: Tom: Alright, we've covered the results, the methodology, the memorization problem, and the in-context learning limitations. Let's bring it all together for "Bye-bye, Bluebook? Automating Legal Procedure with Large Language Models."

Jane: The bottom line — and I know we're not supposed to say that, but it fits here — is that these models are not ready to replace human Bluebook compliance. The best model hit seventy-four percent accuracy, and even with the rules provided, it only reached seventy-seven percent. For legal work, where a single citation error can get a filing rejected, that's just not good enough.

Tom: And it's not just about the Bluebook. The paper makes a compelling case that the Bluebook is a proxy for legal procedure more broadly. If AI can't handle this relatively contained rule system, we should be very skeptical about claims that AI can handle the full complexity of legal practice.

Lu: I think the deeper lesson is about the difference between knowing and doing. These models know a lot about the law — they can discuss legal concepts fluently. But procedural compliance is about doing, not knowing. It's about executing a series of precise steps under constraints. And that's a capability that's still very much underdeveloped.

Meng: From a practical standpoint, this paper gives engineers a clear roadmap of what needs to improve. We need models that can apply rules to novel inputs without memorization. We need architectures that can effectively use long rule documents. And we need evaluation benchmarks like this one to measure progress.

Lalam: If I may add — what excites me about this paper is that it's a step toward making AI systems more trustworthy in high-stakes domains. The legal profession is conservative for good reason. Showing where AI fails is just as valuable as showing where it succeeds. This kind of rigorous evaluation builds the foundation for eventual adoption, even if that adoption is years away.

Jane: That's a great note to end on. The paper isn't saying "never" — it's saying "not yet, and here's exactly why." That's the kind of honest assessment the field needs.

Tom: So, bye-bye to the Bluebook? Not quite. But this paper gives us a clear picture of what needs to happen before we can say that for real. Thanks for joining us, and we'll see you next time with another paper.

Matthew Dahl, Eric Martínez

Yale Law School

cs.CL, cs.AI, cs.CY

Submitted: 2026-08-16

Updated: 2026-08-18

Project page: https://indigobook.github.io/versions/indigobook-2.0-rev2023-2.html

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 60/100

The gist: This paper evaluates whether large language models (LLMs) can automate compliance with *The Bluebook: A Uniform System of Citation*, a 500+ page manual of complex legal citation rules.

Key concepts

The Bluebook
A massive manual detailing how to format citations in legal documents. It serves as a proxy for broader legal procedure because it is highly structured, intricate, and demands strict adherence to specific rules.
Legal Procedure
The set of rules regarding how things are filed and handled within the law. The paper uses the Bluebook's structure to test if AI can handle these complex procedural requirements beyond simple citation formatting.
Memorization vs. Rule Application
The study found that models performed better on real cases than synthetic ones, indicating they were merely regurgitating training data rather than applying the actual rules of the Bluebook in a practical, legal sense.
In-Context Learning
The test where a long document containing the Bluebook rules was provided to AI models. While some improvement was seen (a six-point increase), this learning process struggled to synthesize and apply complex rules across multiple sections.

Terminology

Summary

This paper evaluates whether large language models (LLMs) can automate compliance with The Bluebook: A Uniform System of Citation, a 500+ page manual of complex legal citation rules. The authors construct an original dataset of 866 Bluebook tasks and test flagship LLMs from OpenAI, Anthropic, Google, Meta, and DeepSeek. They show (1) that these models produce fully compliant Bluebook citations only 69%-74% of the time and (2) that in-context learning on the Bluebook’s underlying system of rules raises accuracy only to 77%. These results caution against using off-the-shelf LLMs to automate aspects of the law where fidelity to procedure is paramount.

The paper makes three main contributions: "1. Bluebook task dataset. We aggregate an original dataset of 866 challenging Bluebook formatting tasks, complete with ground-truth answers provided by experts. 2. Legal procedure abilities. We present new evidence quantifying LLMs’ legal procedure-following abilities, isolated from their substantive legal reasoning abilities. 3. Long-context challenges. We contribute to a growing literature about LLMs’ true long-context learning abilities, which may be more limited than advertised."

The authors frame legal citation as a form of legal procedure, noting that citation requirements exhibit Universality (trans-substantive application), Intricacy (convoluted rules requiring synthesis of multiple provisions), and Inflexibility (strict adherence required). They argue that legal citation offers an important starting point for legal procedure research. If LLMs can faithfully implement universal, intricate, and inflexible rules in this context, they may be able to do so in others as well.

The dataset comprises 16 distinct tasks totaling 866 queries in a mix of cloze-style and open-ended formats, grouped into three categories: (1) case law tasks (case names, reporters, parallel reporters, court and date, history, short forms, parentheticals), (2) enacted law tasks (constitutional provisions, federal and state statutes in print and electronic formats, statute short forms, administrative law), and (3) other tasks (legislative resources, court filings, secondary sources, signals). The ground-truth answers are sourced from two legal writing resources: the Interactive Citation Workbook for The Bluebook and Understanding and Mastering The Bluebook.

The methodology involves two evaluations. First, zero-shot performance is tested on five flagship models: OpenAI’s gpt-4.1-2025-04-14, Anthropic’s claude-3-5-sonnet-20241022, Google’s gemini-2.5-flash-preview-04-17, Meta’s llama-3.1-405b-instruct, and DeepSeek’s deepseek-v3-0324. Second, for Gemini 2.5 Flash only, the authors evaluate the performance gains that result from in-context learning on the Bluebook’s underlying system of rules, providing the LLM with the rules (90k tokens, comparable to other legal procedure documents) and instructing it to follow them. Correctness is assessed using exact string matching between the LLM’s response and the expert Bluebook answer, though errors in italicization and other minor stylistic mistakes (e.g., extra commas or periods at the end of the response) are ignored.

The zero-shot results show performance is mediocre, ranging from 74% average accuracy at the high end (GPT 4.1) to 69% average accuracy at the low end (Llama 3.1 405B). The models perform best on case law tasks, especially strong at generating volume, reporter, and page combinations (pooled average = 96%) and court and date abbreviations (pooled average = 90%). However, performance declines on the abbreviation of party names (pooled average = 78%), inclusion of subsequent appellate history (pooled average = 78%), short form citations (pooled average = 53%), and parallel reporters (pooled average = 42%).

A memorization analysis reveals that "most of the LLMs are better at creating citations to the real cases than to the synthetic ones, suggesting that their strong performance on these tasks is driven by simple memorization of the real cases’ citations. Only Gemini 2.5 Flash and DeepSeek V3 0324 do not show evidence of this behavior. The gap is greatest in Claude 3.5 Sonnet, the model that performs the best overall."

For enacted law and other tasks, LLMs perform significantly worse. The models struggle most with crafting citations to electronic versions of statutes (pooled average = 10%), a format requiring many abbreviations referring to the date and legislative session through which the electronic source is current. The signals task also proves difficult (pooled average = 34%), which involve[s] an element of legal reasoning that distinguishes them from the other tasks.

An error analysis shows that when the LLMs are generating non-compliant citations, they are not merely forgetting a comma or a hyphen—around 14 characters are out of place in each incorrect citation. Qualitative examples of substantive errors include misstating the parties, the deciding court, the location in the case reporter, the subsequent history, the publisher, the identity of the author, or a citation’s signal of support. However, "semantic errors of this magnitude are relatively rare. For the vast majority of LLM-generated citations evaluated in this paper, a reader would be able to understand and locate the referenced source, even when the citation is technically incorrect."

The in-context learning evaluation for Gemini 2.5 Flash shows that "although in-context learning on the Bluebook rules does improve performance on certain tasks—especially on the statutes, secondary sources, and signals tasks—its effect on overall performance is minor. Average accuracy increases only six percentage points, from 71% in the zero-shot setup to 77% in the in-context learning setup. These results call into question the efficacy of LLMs’ stated context window lengths for legal tasks. If legal rules are sufficiently complex—as the Bluebook rules certainly are—in-context learning alone may not be enough to ensure that LLMs adhere to them."

The paper concludes that despite the impressive capabilities of LLMs in certain legal applications, they continue to struggle with the complex and highly structured procedural rules that govern legal citation. The authors state: If LLMs cannot yet even conquer the Bluebook, it remains dubious that they can be safely and responsibly integrated into higher-stakes legal settings. They call for further research into the procedural abilities of LLMs, noting that "before generative AI can be entrusted with mission-critical assignments—like sentencing, adjudication, or even legal drafting—it should first be able to demonstrate reliability in following the formal rules that constitute the structure of the legal system in the first place."

Improvements for AI systems

Based on the paper's findings, here are specific improvements to AI systems and their resulting capabilities:

  1. Add procedural-rule verification layer
  • Implement a post-generation validation module that checks outputs against a structured rule database (e.g., Bluebook rules encoded as machine-readable constraints)

  • Use a two-pass approach: generate the citation, then run a rule-checker that flags violations (e.g., missing "3d" in reporter abbreviations, incorrect court abbreviations for pre-1976 Kentucky cases)

  1. Reduce memorization dependence
  • Train or fine-tune on synthetic case captions with randomized page numbers, reporter volumes, and court names to force rule-based reasoning rather than rote recall

  • Add adversarial training examples where the correct citation differs from any real-world case the model may have memorized

  1. Improve long-context instruction following
  • Implement hierarchical attention or retrieval within the provided rule text (e.g., chunk the 90k-token rulebook into sections and retrieve relevant rules per query)

  • Use a rule-lookup step before generation: identify which Bluebook rules apply to the given task type, then condition generation on those specific rules

  1. Add structured output constraints
  • Use constrained decoding or grammar-based generation for citation components (e.g., enforce patterns like + [A-Z].[A-Z]. + for reporters)

  • Separate generation into fields (case name, reporter, volume, page, court, year, history) and validate each field independently before assembly

  1. Incorporate error-correction training
  • Fine-tune on pairs of incorrect LLM outputs and their corrected versions (using the paper's error analysis showing 14 character edits needed per error)

  • Train a separate citation corrector model that takes a draft citation and produces a compliant one

  • Reliable legal citation formatting: The improved system can produce Bluebook-compliant citations with >90% accuracy (up from 69-74%), reducing the need for human proofreading

  • Rule-based generalization: Correctly formats citations for novel cases, statutes, and secondary sources it has never seen, rather than relying on memorized examples

  • Long-document procedure following: Can ingest and apply a 90k-token procedural manual (e.g., Federal Rules of Civil Procedure) with higher fidelity, making it useful for drafting compliant legal filings

  • Error detection and self-correction: Identifies and fixes its own citation errors (e.g., missing "3d" in S.W.3d, wrong court abbreviation, incorrect signal usage) before final output

  • Task-specific accuracy: Achieves near-perfect accuracy on case reporters (96%+), court/date abbreviations (90%+), and parentheticals (98%+), while improving weaker areas like electronic statutes (from 10% to >50%) and secondary sources (from 29% to >60%)

Abstract

Legal practice requires careful adherence to procedural rules. In the United States, few are more complex than those found in The Bluebook: A Uniform System of Citation. Compliance with this system's 500+ pages of byzantine formatting instructions is the raison d'etre of thousands of student law review editors and the bete noire of lawyers everywhere. To evaluate whether large language models (LLMs) are able to adhere to the procedures of such a complicated system, we construct an original dataset of 866 Bluebook tasks and test flagship LLMs from OpenAI, Anthropic, Google, Meta, and DeepSeek. We show (1) that these models produce fully compliant Bluebook citations only 69%-74% of the time and (2) that in-context learning on the Bluebook's underlying system of rules raises accuracy only to 77%. These results caution against using off-the-shelf LLMs to automate aspects of the law where fidelity to procedure is paramount.

Related papers