LegalPincite: Multi-level Legal Information Retrieval Dataset

arXiv:2608.03756 · cs.IR, cs.CL · Submitted 2026-08-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "LegalPincite: Multi-level Legal Information Retrieval Dataset".

Jane: This paper introduces LegalPincite, a large-scale legal Information Retrieval (IR) dataset constructed from Court of Justice of the European Union (CJEU) judgments, designed to address critical limitations in existing datasets.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Alright, so the paper is titled "LegalPincite: Multi-level Legal Information Retrieval Dataset." It’s essentially presenting a new resource that aims to fix those citation problems we talked about earlier by building something big from Court of Justice of the European Union judgments.

Jane: Right, and the authors are Theresia Veronika Rampisela, Henrik Palmer Olsen, and Giovanni Colavizza from the University of Copenhagen. They are clearly deep in this legal information retrieval space.

Lu: The title itself tells you a lot about their approach; they aren't just making another dataset; they’re focusing on multi-level retrieval support, which means testing how well systems work at different granularities.

Meng: Multi-level support implies they are testing things from the broad case level down to the specific paragraph level, which gives us a much better picture of system performance across various complexity settings.

Lalam: It suggests that whatever AI we build using this data can be more nuanced in its understanding of legal documents because it has to handle these different levels of detail simultaneously.

The paper's summary: Tom: Now, let's look at the actual summary of "LegalPincite: Multi-level Legal Information Retrieval Dataset." Essentially, they describe constructing this dataset by merging an unpublished collection with a published one that already had some relevance annotations.

Jane: They explain that they had to resolve several tricky issues during construction, like dealing with language ambiguity using tools called langdetect, and fixing segmentation errors by downloading the HTML or XHTML versions of the judgments.

Lu: What’s crucial is how they handled data leakage in the queries; they used heuristics and named entity recognition to strip out citation info like case titles or paragraph numbers from the queries themselves.

Meng: That part about masking query information is key for building robust AI because we can't afford to train models that accidentally rely on identifiers that don't exist during inference.

Lalam: And they didn't just use clean data; they included all paragraphs, even those that were neither citing nor cited, which makes the corpus much more comprehensive than what was previously available in many setups.

The paper's improvements: Tom: Moving on to the improvements they detail in "LegalPincite: Multi-level Legal Information Retrieval Dataset," they highlight how their construction addresses those limitations we discussed earlier, focusing on masking citation information and including all paragraphs from the corpus.

Jane: They also mention a specific way they aggregated relevance annotations across two different human experts, using a rule that checks if both experts agree on the presence or expansion of a candidate paragraph in the citing paragraph.

Lu: This multi-expert validation method for relevance is smart because it attempts to create a more consistent ground truth when relying on human input, even though they admit that those raw human annotations weren't directly usable in their initial state.

Meng: So, this means the resulting dataset isn't just a collection of text; it’s been heavily curated to ensure the relevance labels are as accurate as possible for training purposes.

Lalam: It also addresses data freshness by updating the data in May two thousand twenty-six with new judgments from January two thousand twenty-one through December two thousand twenty-five which is important because legal interpretations evolve over time.

Conclusion: Tom: So, to wrap up on "LegalPincite: Multi-level Legal Information Retrieval Dataset," the authors conclude that this resource is a large-scale collection for finding relevant legal citations at various query-document levels. They emphasize its contribution to the advancement of legal IR research and suggest avenues for future work, like citation link prediction and cross-lingual retrieval.

Jane: They are really setting up this dataset to be useful not just for finding citations, but also for bigger tasks like predicting citation links or even doing legal textual entailment.

Lu: It feels like the implication is that we can move beyond simple matching and start building AI that understands the relationships between different paragraphs within a case structure more deeply than before.

Meng: From an engineering standpoint, having these various query-document levels lets us test if our retrieval mechanisms are actually working well at every possible scale of precision.

Lalam: I think the real impact here is that this structured data will help improve the culture of legal AI by providing a reliable foundation for systems that can analyze and synthesize complex legal arguments with much more accuracy.

University of Copenhagen

cs.IR, cs.CL

Submitted: 2026-08-04

Updated: 2026-09-26

Comments: Accepted for publication at the 8th Natural Legal Language Processing Workshop (NLLP 2026), co-located with EMNLP 2026

Code: https://github.com/coastalcph/paragraph_network

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 77/100

The gist: This paper introduces LegalPincite, a large-scale legal Information Retrieval (IR) dataset constructed from Court of Justice of the European Union (CJEU) judgments, designed to address critical

Key concepts

Multi-level Retrieval Support
This refers to testing search engines at three distinct scales: finding an entire case based on a query, finding a specific paragraph within a case, and finding one paragraph based on another paragraph. This complexity tests the system's ability to handle different levels of specificity in legal searches.
Data Leakage Mitigation
The researchers actively removed sensitive citation details like case titles and paragraph numbers from queries to prevent the model from cheating by using direct identifiers. They also used expert consensus rules for relevance labeling, ensuring the data is realistic for real-world testing.
Relevance Annotation Aggregation
Relevance labels are determined by two human experts answering specific questions about a candidate paragraph's relationship to the citing text. A paragraph is marked relevant only if both experts agree on certain criteria related to its presence or expansion in the original source.

Terminology

Summary

This paper introduces LegalPincite, a large-scale legal Information Retrieval (IR) dataset constructed from Court of Justice of the European Union (CJEU) judgments, designed to address critical limitations in existing datasets. It provides multi-level retrieval support—case-to-case, paragraph-to-case, and paragraph-to-paragraph—and includes ground truth citations with partial human expert validation. This resource is significant because it moves beyond oversimplified settings by masking citation information in queries and including all paragraphs from the corpus, thereby supporting the rigorous development and evaluation of legal IR methods at multiple query-document levels.

Dataset Construction

LegalPincite merges two existing datasets: an unpublished dataset containing all paragraphs (citing, cited, and neither) from CJEU judgments up to July 2024, and a published dataset consisting of 110,601 pairs of citing and cited paragraphs with relevance annotations. The construction involved resolving four categories of issues: language ambiguity (using langdetect), missing text (recovering missing text from the all-paragraph dataset), segmentation errors (downloading HTML/XHTML versions to correct parsing errors), and non-reusable human annotations. Furthermore, data was updated in May 2026 by scraping new judgments and extracting 41,547 unique citations using the Cellar API.

Data Preprocessing and Leakage Mitigation

The dataset undergoes several cleaning steps to ensure quality. For query masking, the authors employ a strategy to mitigate data leakage by applying heuristics and named entity recognition to remove citation-related information, such as case title, case number, paragraph number, or involved parties. The final format is structured into CSV files compatible with pyterrier. Relevance annotations are aggregated across two experts using a specific rule: a candidate paragraph is relevant if both experts answer 'Yes' to at least one of two questions regarding the rule's presence or expansion in the citing paragraph.

Data Splitting and Statistics

The cleaned data is split into training, development, and testing sets based on document year (pre-2018 for train queries, 2018–2021 for dev queries, and 2022–2025 for test queries). A crucial aspect of the splitting strategy is preventing temporal leakage: For dev and test, the ground-truth (relevant) citations and candidate cases/paragraphs are restricted to those with a document year preceding all queries in the split. Statistical analysis shows that while lexical overlap and semantic similarity are generally low (median < 0.5), indicating that relevance cannot be determined via simple matching alone, making the dataset a challenging and realistic legal IR benchmark.

Experimental Evaluation

The paper evaluates retrieval baselines—including TF-IDF, BM25, LMIR, and DPH—across three query-document levels: case-to-case (case), paragraph-to-case (par-case), and paragraph-to-paragraph (par-par). Results indicate that effectiveness does not differ much between the two splits, but significantly differs across the retrieval levels. For instance, LMIR performs best for case-to-case retrieval, while BM25 scores the best for paragraph-to-paragraph retrieval. The authors also analyze data leakage by comparing performance under three settings: original (unmasked), without paragraph IDs (w/o par ID), and without citation information (w/o citation), concluding that simply removing paragraph IDs to mitigate leakage is insufficient, as it may still inflate performance.

Conclusion and Future Scope

LegalPincite is presented as a large-scale and multi-level legal IR test collection in English that can be used for finding relevant legal citations at various query-document levels. Beyond legal IR, the paper suggests potential extensions for tasks such as citation link prediction, cross-lingual retrieval (given its multilingual origin), legal textual entailment, and evaluation of Large Language Model legal reasoning. The dataset is released publicly under a CC-BY license on Hugging Face Hub to enhance discoverability and interoperability. The authors acknowledge limitations, such as the potential for feedback loops in ground truth annotations.

Improvements for AI systems

As a fastidious and diligent AI researcher, I have analyzed the LegalPincite: Multi-level Legal Information Retrieval Dataset paper. The core contribution is a large-scale, high-fidelity legal dataset designed to address critical limitations in existing legal Information Retrieval (IR) systems, specifically regarding paragraph-level citation retrieval (pincites).

Here are the specific improvements and capabilities this dataset enables for AI systems:


  1. The improved system can perform multi-level retrieval tasks with high precision, moving beyond simple case-to-case or paragraph-to-case matching.

  2. The system can execute highly nuanced queries such as:

Ease a query to find the specific paragraph within a case that cites it (paragraph-to-paragraph retrieval).

  1. The system can be trained and evaluated on realistic legal scenarios by mitigating data leakage from queries, which typically leak citation information (like case numbers or paragraph numbers).

  2. The improved system can handle noisy, real-world query text effectively because the dataset includes a mechanism for masking citation-related entities using pre-trained Legal Named Entity Recognition models and regular expressions.

  3. The system can be robust against corpus simplification by being trained on a full set of paragraphs (citing, cited, and non-citing/non-cited), ensuring the retrieval pool is realistic rather than artificially constrained to only citing or cited text.

  4. The system can utilize human expert validation for ground truth labeling, allowing researchers to build highly accurate relevance judgments even when relying on limited expert annotation sets (e.g., evaluating performance on subsets with EUR-Lex vs. human annotations).

  5. The system can be optimized for specific retrieval levels based on the query type:

Ease the selection of the optimal IR baseline by understanding that different techniques excel at different tasks (e.g., LMIR for case-to-case, BM25 for paragraph-to-paragraph).

  1. The improved system can be extended into specialized legal AI applications beyond retrieval, including:

Ease developing Citation Link Prediction systems to predict which paragraphs will cite others.

Ease building Cross-lingual Retrieval systems by leveraging the multilingual nature of the dataset.

Ease creating Legal Textual Entailment models by extracting and comparing relevant portions of cited paragraphs to answer complex legal questions.

Ease powering Legal Retrieval Augmented Generation (RAG) systems for preliminary ruling analysis by pairing initial questions with their grounded, paragraph-level evidence.

  1. The system can be evaluated using comprehensive IR metrics (HR@k, NDCG@k, MAP, MRR), allowing for rigorous comparison of retrieval methods across different query-document levels and temporal splits (e.g., pre-2018 vs. 2022-2025).

Abstract

A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While legal practice often requires pinpoint citations (pincites) to specific case paragraphs, most existing public legal IR datasets lack paragraph-level citation annotations. Yet, publicly available datasets with such information contain data leakage in the query text and exclude paragraphs that are neither citing nor cited from the corpora, creating an unrealistic and oversimplified retrieval setting, potentially leading to inflated performance. To address these limitations, we contribute a large-scale legal IR dataset constructed from Court of Justice of the European Union (CJEU) judgments. The dataset contains: (i) masked case/paragraph queries, with removed citation information; (ii) a corpus that includes all paragraphs; and (iii) case- and paragraph-level ground truth citations, with partial human expert validation. Our dataset supports both the development and rigorous evaluation of legal IR methods, at multiple query-document levels (case-to-case, paragraph-to-case, and paragraph-to-paragraph retrieval). Link to dataset and code: https://huggingface.co/datasets/theresiavr/legalpincite

Sources

Related papers