Annotated Surrogate Retrieval for Polish Statutory Law
cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: 16 pages, 5 figures, 5 tables. Code and data: https://github.com/OryCore/Research
Code: https://github.com/OryCore/Research
License: http://creativecommons.org/licenses/by/4.0/
The gist: We present a family of retrieval methods for Polish statutory law built on document surrogates: language-model annotations attached to statutory articles at index time.
Terminology
Abstract
We present a family of retrieval methods for Polish statutory law built on document surrogates: language-model annotations attached to statutory articles at index time. Three designs occupy different points on the cost-quality frontier. ASCR is a surrogate cascade with reranking; ASCR-H fuses a dense list into that cascade; and DTF replaces both language-model stages with three lexical and dense retrievers, weighted reciprocal rank fusion, and a deterministic re-scoring prior, using no model call before generation. We evaluate all three against fourteen lexical, dense, fused and ablated baselines plus four controls, on 300 questions from the 2024 and 2025 Polish bar and legal counsel entrance examinations (264 with their reference article in the corpus), over 82,508 articles from 1,133 acts. On paired McNemar tests, ASCR-H places the reference provision at rank one significantly more often than every other non-oracle configuration except one of its own ablations (eighteen of twenty comparisons significant in its favour at p < 0.005), reaching 72.3% against 61.7% for BM25 and 52.3% for dense retrieval. The advantage is concentrated at the head and does not survive depth: it is significant at cutoffs of one and five, disappears by ten, and by twenty DTF leads on point estimate (86.0% versus 84.5%) at one ninth the latency and less than half the cost. Ablation attributes 27.6 points of rank-one accuracy to the reranking stage alone. We further report that the ranking advantage does not extend to citation accuracy, where DTF matches the oracle ceiling, and three negative results on lemmatisation, pseudo-relevance feedback and query rewriting. Surrogate annotation covers 27.0% of the corpus but every reference provision in the benchmark, an asymmetry we disclose and discuss. Benchmark, per-question outputs and paired significance tests are publicly available.
Sources
- Legal RAG Bench: an end-to-end benchmark for legal RAG
- Assessing generalization capability of text ranking models in Polish
- PIRB: A Comprehensive Benchmark of Polish Dense and Hybrid Text Retrieval Methods
- LLM-as-a-Judge is Bad, Based on AI Attempting the Exam Qualifying for the Member of the Polish National Board of Appeal
- Are LLM-Based Retrievers Worth Their Cost? An Empirical Study of Efficiency, Robustness, and Reasoning Overhead
- LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain
- Legal-DC: Benchmarking Retrieval-Augmented Generation for Legal Documents
- Silver Retriever: Advancing Neural Passage Retrieval for Polish Question Answering
- Finding the Law: Enhancing Statutory Article Retrieval via Graph Neural Networks
- BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
- Fine-Tuning LLaMA for Multi-Stage Text Retrieval
- Multilingual E5 Text Embeddings: A Technical Report
- CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law
- Passage Re-ranking with BERT
- Document Expansion by Query Prediction
- Multi-Legal-Bench: When the Answer Is in the Input. Label Leakage in Legal Benchmarks Built from Court Registries
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering