Citation-Closure Retrieval and Per-Rule Attribution for Real-World Regulatory Compliance Question Answering
cs.AI
Submitted: 2026-05-28
Updated: 2026-08-31
Comments: Findings of EMNLP 2026
Code: https://github.com/yeongjoonJu/RefWalk
License: http://creativecommons.org/licenses/by/4.0/
The gist: Deploying Large Language Models (LLMs) for regulatory compliance demands rigorous traceability via comprehensive citations across multi-tiered authority structures.
Terminology
Abstract
Deploying Large Language Models (LLMs) for regulatory compliance demands rigorous traceability via comprehensive citations across multi-tiered authority structures. Unlike traditional multi-hop or legal QA, this task requires structured procedural lookups and evidence-set closure rather than entity resolution or case-law reasoning. Existing RAG systems struggle here due to flattened citation edges, fragmented retrieval expansions, and fragile post-hoc attribution. We formalize Regulatory Compliance QA with RegOps-Bench, a novel benchmark featuring an Operational Knowledge Graph derived from complex national R&D regulations. To address these bottlenecks, we propose RefWalk, a unified framework driven by a shared topic anchor. RefWalk traverses cross-document citations, fuses multi-view candidates via max-based aggregation, and enforces per-rule attribution to explicitly map claims to sources. We establish a strong baseline with substantial improvements in retrieval recall and citation accuracy. Finally, a contrastive evaluation on a U.S. health compliance dataset (HIPAA) reveals that existing systems exhibit saturation on flat-structure rules, underscoring the need for RegOps-Bench. Our code is available at https://github.com/yeongjoonJu/RefWalk.
Sources
- SustainableQA: A Comprehensive Question Answering Dataset for Corporate Sustainability and EU Taxonomy Reporting
- Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules
- From RAG to Memory: Non-Parametric Continual Learning for Large Language Models
- RAG-Fusion: a New Take on Retrieval-Augmented Generation
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- PIKE-RAG: sPecIalized KnowledgE and Rationale Augmented Generation
- LawThinker: A Deep Research Legal Agent in Dynamic Environments
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection