Mining Legal Arguments to Study Judicial Formalism
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Mining Legal Arguments to Study Judicial Formalism".
Jane: The paper was written by Tomáš Koref, Lena Held, Mahammad Namazov, Harun Kumru, Yassine Thlija et al. from Goethe University Frankfurt and Charles University and Technical University of Darmstadt and Ruhr University Bochum.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. Today we're looking at a paper that's got a title that sounds like it belongs in a law library, but it's actually a really clever piece of computer science. It's called "Mining Legal Arguments to Study Judicial Formalism." I'm here with Jane, and we're going to break this down.
Jane: And I'm so glad we are, Tom, because when I first read the title, I thought, okay, this is about robots reading court decisions. And that's part of it. But the bigger question is, can we actually tell if a judge is being formalistic? Meaning, are they just following the letter of the law, or are they thinking about the purpose behind the law?
Tom: Right, and that's a huge deal. For decades, people have been saying that courts in Central and Eastern Europe are stuck in this formalistic mindset, a leftover from the communist era. But nobody had actually checked systematically. This paper is the first to try and do that with machine learning.
Jane: And the authors are from Goethe University Frankfurt, Charles University in Prague, and a few German research institutes. So it's a real international team, which makes sense because you need both legal experts and computer scientists for this.
Tom: Exactly. And the key thing they did, Jane, is they built a dataset. They took two hundred seventy-two decisions from the Czech Supreme Court and the Supreme Administrative Court, and they had legal experts annotate over nine thousand paragraphs. That's the foundation for everything else.
Jane: Nine thousand paragraphs, and they tagged each one with different types of legal arguments. Like, is this a linguistic argument, where the judge is just looking at the exact wording of the law? Or is it a teleological argument, where they're thinking about the purpose of the law? That distinction is at the heart of formalism.
Tom: And the implications here are pretty big. If this works, we can finally test these old claims about entire legal systems being formalistic. We can do it with data, not just anecdotes.
Jane: And that's what we're going to dig into next. How did they actually teach the models to spot these arguments? Stay with us.
Summary: Tom: So we're back with "Mining Legal Arguments to Study Judicial Formalism," and Jane, I want to get into the actual methods, because this is where it gets interesting. They didn't just throw one big model at the problem. They built a pipeline.
Jane: Right, and I love that they broke it down into three separate tasks. First, they had to figure out which paragraphs even contain arguments at all. Because in a court decision, a lot of the text is just procedural boilerplate. They found that eighty-seven percent of the paragraphs had no arguments at all.
Tom: That's a huge number. So the first model, a ModernBERT model, acts like a filter. It says, "This paragraph is worth reading, this one isn't." And it did that really well, with a balanced F1 score of eighty-two point six percent. That's the harmonic mean of precision and recall, for our listeners who want the technical detail.
Jane: Then the second task is the hard one. Once you have an argumentative paragraph, what kind of argument is it? They have eight categories, from linguistic interpretation to practical consequences. And for that, they used a much bigger model, Llama three point one 8B, fully fine-tuned.
Tom: And that got a balanced F1 of seventy-seven point five percent. Which is promising, but not perfect. And Jane, I think it's worth mentioning that some of the categories were just harder than others. The ones where the human annotators themselves disagreed a lot, the model also struggled with.
Jane: That's such an important point, Tom. It's not just a machine problem. Legal experts disagree on what counts as a "practical consequences" argument. So the model is learning from noisy labels. The paper is very honest about that, and they even publish the raw annotations with the disagreements, which is rare and really good practice.
Tom: And then the third task is the big one. Can you look at an entire decision and say, "This is formalistic" or "This is not formalistic"? And here's where they did something clever. They combined everything. They used the argument classifications as features in a simple neural network.
Jane: And that multi-step pipeline got a balanced F1 of eighty-three point eight percent, which is actually better than just having a single end-to-end model try to read the whole decision. So the lesson is, sometimes breaking a problem down into smaller, more understandable pieces works better than one giant black box.
Tom: And it's more explainable. You can see exactly why the model thinks a decision is formalistic. It's because it found lots of case law citations and very few references to legal principles. That's a huge advantage for legal scholars who need to trust the analysis.
Jane: And that's what we'll talk about next, what these results actually mean for the big debate about formalism in Central and Eastern Europe.
Improvements: Tom: So we've talked about how the models work, but the real payoff is what they found. And Jane, this is where the paper really challenges the status quo.
Jane: It absolutely does. The prevailing narrative is that Czech courts, especially the Supreme Court, are text-bound formalists. They just look at the literal wording of statutes and ignore the bigger picture. But the data says something very different.
Tom: Right, linguistic interpretation, which is the most formalistic argument type, only made up five point seven percent of all arguments. That's tiny. Meanwhile, case law was the most common at thirty-seven point four percent, and purposive and value-based arguments together made up over thirty-two percent.
Jane: So they're not ignoring the purpose of the law. They're citing prior decisions and thinking about the principles behind the rules. That's not the picture of a court that's just mechanically applying text.
Tom: And the paper also challenges the idea that the Supreme Court and the Supreme Administrative Court are completely different. The narrative says the SAC is the good, modern court, and the SC is the bad, old, formalistic one. But their argumentation profiles are actually pretty similar.
Jane: There is a nuance there, though. When you look at the holistic label, the SC was formalistic in sixty-four percent of its decisions, while the SAC was split fifty-fifty. So there is a difference, but it's not the stark contrast that people have been claiming.
Tom: And the temporal trend is interesting too. The SAC has been getting more non-formalistic over time, especially after two thousand fourteen. So maybe the story is less about a fixed character of the courts and more about evolution.
Jane: And that's the improvement this paper brings. It's not just a new dataset or a new model. It's a new way to test long-held beliefs about legal culture. Instead of relying on a few famous cases or the impressions of scholars, you can now look at hundreds of decisions and get a statistical picture.
Tom: And this could be applied to other countries. The methodology is not Czech-specific. You could use the same approach to study courts in Poland, Hungary, or even the United States. That's a big deal for comparative legal studies.
Jane: Absolutely. And I think the most exciting implication is that we can now start to separate the myth from the reality. And maybe some of these old narratives about post-communist legal culture need to be rewritten.
Tom: Let's bring in Lu and Meng to get their take on the technical side and the practical side of this. Lu, what do you think is the most impactful part of this work?
Lu: I think the most impactful part is the decomposition of the task. They didn't just try to classify formalism directly. They first identified arguments, then classified them, then used those as features. This is a great example of how to make NLP more interpretable and reliable for high-stakes domains like law. The SHAP analysis they did, showing that the presence of "Principles of Law" arguments is the strongest predictor of non-formalism, is exactly the kind of insight that makes this trustworthy.
Meng: And from a practical standpoint, the fact that they filter out eighty-five percent of the text before running the expensive Llama model is a huge efficiency win. It means you could run this on a much larger corpus without breaking the bank. They have a corpus of three hundred thousand decisions ready to go. That's the kind of scale that makes this a real tool, not just a research demo.
Tom: Great points from both of you. So the paper gives us a more efficient, more explainable, and more accurate way to study judicial reasoning. And it's already producing results that challenge a long-standing narrative. Let's wrap this up in our final segment.
Conclusion: Tom: Alright, we're at the end of our discussion on "Mining Legal Arguments to Study Judicial Formalism." And I think we can all agree this is a paper that does a lot of things at once.
Jane: It does. It creates a new dataset, which is a huge contribution in itself. It tests different models and finds that a multi-step pipeline works best. And it uses that pipeline to make a real empirical finding about Czech courts.
Tom: And that finding is that the old story about these courts being text-bound formalists just doesn't hold up. They use case law and purposive reasoning far more than they use literal textual analysis. That's a significant challenge to the existing scholarship.
Jane: And it's not just about Czechia. The methodology can be exported. So we might see similar studies for other countries in the region, and we might find that the "formalistic Central European" stereotype needs to be revised across the board.
Meng: And from an engineering perspective, the fact that they open-sourced everything, the models, the code, the guidelines, means other people can build on this immediately. That's how you make a real impact.
Lu: I'd add that this is a great example of how to handle the messiness of human annotation. They don't hide the disagreements. They publish them. That's the kind of transparency that will push the field forward.
Tom: So, to sum it up, "Mining Legal Arguments to Study Judicial Formalism" is a landmark paper for computational legal studies. It gives us the tools to finally test some very old and very important claims about how judges think.
Jane: And it does it in a way that's rigorous, transparent, and practical. We're saying goodbye to this paper, but we're definitely going to be watching to see what happens when they run their models on the full three hundred thousand decisions. That could be a game-changer.
Tom: Absolutely. Thanks for joining us, and we'll see you next time with another paper.
Tomáš Koref, Lena Held, Mahammad Namazov, Harun Kumru, Yassine Thlija, Ivan Habernal
Goethe University Frankfurt · Charles University · Technical University of Darmstadt · Ruhr University Bochum
cs.CL, cs.CY
Submitted: 2026-03-21
Updated: 2026-08-18
Comments: pre-print under review
Code: https://github.com/trusthlt/madon
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 80/100
The gist: Courts must justify their decisions, but systematically analyzing judicial reasoning at scale remains difficult.
Key concepts
- Judicial Formalism
- A concept suggesting that judges only follow the literal wording of the law (linguistic interpretation) and ignore the underlying purpose or intent behind it. The paper tests this idea using data.
- Multi-step Pipeline
- The method used in the study, which broke down a complex task into smaller parts. This involved first identifying argumentative paragraphs, then classifying the argument type within them, and finally combining these features for a final decision.
- Linguistic Argument
- A specific type of legal argument where the judge focuses solely on the exact wording or literal text of a statute. This is considered one of the most formalistic types of reasoning.
- Teleological Argument
- An argument that considers the purpose or underlying goal (telos) behind a law, rather than just its literal wording. The paper uses this to argue that judges consider the bigger picture.
Terminology
Summary
Courts must justify their decisions, but systematically analyzing judicial reasoning at scale remains difficult. This study tests claims about formalistic judging in Central and Eastern Europe (CEE) by developing automated methods to detect and classify judicial reasoning in decisions of Czech Supreme Courts using state-of-the-art natural language processing methods. We create the MADON dataset of 272 decisions from two Czech Supreme Courts with expert annotations of 9,183 paragraphs with eight argument types and holistic formalism labels for supervised training and evaluation. Using a corpus of 300,511 Czech court decisions, we adapt transformer LLMs to Czech legal domain through continued pretraining and we experiment with methods to address dataset imbalance including asymmetric loss and class weighting. The best models can detect argumentative paragraphs (82.6% Bal-F1), classify traditional types of legal argument (77.5% Bal-F1), and classify decisions as formalistic/non-formalistic (83.8% Bal-F1). Our three-stage pipeline combining ModernBERT, Llama 3.1, and traditional feature-based machine learning achieves promising results for decision classification while reducing computational costs and increasing explainability. Empirically, we challenge prevailing narratives about CEE formalism. We demonstrate that legal argument mining enables promising judicial philosophy classification and highlight its potential for other important tasks in computational legal studies. Our methodology can be used across jurisdictions, and our entire pipeline, datasets, guidelines, models, and source codes are available at https://github.com/trusthlt/madon.
Improvements for AI systems
Based on the paper, here are specific improvements I can implement in AI systems:
Improvement: Build a three-stage pipeline that mirrors legal expert reasoning:
-
Stage 1: ModernBERT filters argumentative paragraphs (82.6% Bal-F1), removing 85% of non-relevant text
-
Stage 2: Llama 3.1 8B classifies eight argument types (77.5% Bal-F1) using asymmetric loss
-
Stage 3: MLP with SHAP-explainable features classifies formalism (83.8% Bal-F1)
Capability: The system can analyze entire court decisions, identify legal arguments, classify their types, and predict judicial formalism with explainable reasoning—outperforming end-to-end LLMs while reducing compute costs.
Improvement: Replace standard Binary Cross-Entropy with Asymmetric Loss (Asy) in multi-label classification heads. This consistently improved Bal-F1 by 5–9 points across architectures (e.g., ModernBERT: 62.9%→71.6%; Llama: 74.6%→77.5%).
Improvement: Implement a dual-metric evaluation that reports both positive-class F1 (F1+) and negative-class F1 (F1−), averaged into Bal-F1. This penalizes systematic overprediction, which standard F1 misses.
Improvement: For encoder models like ModernBERT, train a custom Czech legal tokenizer and continue pretraining with masked language modeling (two-phase: 30% masking, then 15% masking) on 300,511 court decisions.
Improvement: Integrate SHAP analysis into the MLP classifier to identify which argument types drive formalism predictions. The paper shows Principles of Law (PL) and Teleological Interpretation (TI) are strongest non-formalism predictors.
Improvement: Implement a task-specific model selection rule:
-
Use ModernBERT for paragraph-level tasks (argument presence, formalism classification)
-
Use fully fine-tuned Llama 3.1 for argument type classification
-
Avoid PEFT/LoRA for legal argument tasks (performance dropped from 79.5% to 49.7% Bal-F1)
Improvement: Incorporate inter-annotator agreement (Krippendorff's α) into model training and evaluation. For low-agreement classes (e.g., Practical Consequences at α=0.20), implement soft labels or ensemble predictions rather than treating gold labels as absolute truth.
Improvement: Build the system using the paper's argument taxonomy (LIN, SI, CL, D, HI, PL, TI, PC) grounded in general legal theory (Alexy, MacCormick, Summers), making it adaptable to other civil law jurisdictions.
Abstract
Courts must justify their decisions, but systematically analyzing judicial reasoning at scale remains difficult. This study tests claims about formalistic judging in Central and Eastern Europe (CEE) by developing automated methods to detect and classify judicial reasoning in decisions of Czech Supreme Courts using state-of-the-art natural language processing methods. We create the MADON dataset of 272 decisions from two Czech Supreme Courts with expert annotations of 9,183 paragraphs with eight argument types and holistic formalism labels for supervised training and evaluation. Using a corpus of 300,511 Czech court decisions, we adapt transformer LLMs to Czech legal domain through continued pretraining and we experiment with methods to address dataset imbalance including asymmetric loss and class weighting. The best models can detect argumentative paragraphs (82.6% Bal-F1), classify traditional types of legal argument (77.5% Bal-F1), and classify decisions as formalistic/non-formalistic (83.8% Bal-F1). Our three-stage pipeline combining ModernBERT, Llama 3.1, and traditional feature-based machine learning achieves promising results for decision classification while reducing computational costs and increasing explainability. Empirically, we challenge prevailing narratives about CEE formalism. We demonstrate that legal argument mining enables promising judicial philosophy classification and highlight its potential for other important tasks in computational legal studies. Our methodology can be used across jurisdictions, and our entire pipeline, datasets, guidelines, models, and source codes are available at https://github.com/trusthlt/madon.
Sources
- LoRA: Low-Rank Adaptation of Large Language Models
- The Llama 3 Herd of Models
- The Czech Court Decisions Corpus (CzCDC): Availability as the First Step
- Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
- When Does Pretraining Help? Assessing Self-Supervised Learning for Law and the CaseHOLD Dataset
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering