Improving Argument Saliency Coverage in Small LLMs for Long Legal Opinion Summarization via Sequence-Level Distillation

arXiv:2608.29884 · cs.CL · Submitted 2026-08-30 · Read on arXiv

cs.CL

Submitted: 2026-08-30

Updated: 2026-08-30

Comments: EMNLP 2026 Main Conference

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 77/100

The gist: The paper details methods for "Improving Argument Saliency Coverage in Small LLMs for Long Legal Opinion Summarization via Sequence-Level Distillation." Summary of Findings and Experimental Results:

Terminology

Summary

The paper details methods for Improving Argument Saliency Coverage in Small LLMs for Long Legal Opinion Summarization via Sequence-Level Distillation.

Summary of Findings and Experimental Results:

The research evaluates the performance of various systems on summarizing a legal opinion, specifically Ezurike v. Ezurike (d_2006nssc73.txt). The evaluation uses Atomic fact coverage, where the verifier assesses whether a summary passage supports, misses, or contradicts specific facts derived from the source text.

Comparison of Teacher Models (Table 7):

The study compares three systems based on their ability to cover atomic facts:

  1. (a) Qwen3-0.6B (baseline): This non-tuned model performed poorly, achieving an ARCScore of 0.250. On the 13 atomic facts, the baseline model showed significant gaps in coverage, with a record of 2 Supported (of 13) – Missing (of 13) – Contradicted (of 13).

  2. (b) + GPT-5-mini teacher: Tuning the model with a GPT-5-mini teacher improved performance, raising the ARCScore to 0.450. The coverage statistics were improved compared to the baseline, but still showed gaps in fact representation.

  3. (c) + Qwen3-14B teacher: Utilizing a Qwen3-14B teacher resulted in the highest performance, achieving an ARCScore of 0.700. This demonstrates that distillation from a more powerful teacher model significantly enhances the small LLM's ability to capture argument saliency and fact coverage.

Few-Shot Distillation Performance (Table 8):

The paper also investigates Few-shot distillation across models and inference strategies, specifically using summary-only distillation from Qwen3-14B, nonthinking inference. This section demonstrates the impact of increasing sample size on performance, using two random seeds (seed 42 and seed 123) for two different models (Qwen3-1.7B and Qwen3-4B).

The trend observed is consistent across seeds:

  • Qwen3-4B: The ARC scores show a clear progression as the number of samples increases, moving from.631 with 0 shots to.720 with 10 samples, improving further to.756 with 100 samples, and stabilizing at.744 with 1000 samples.

  • General Trend: The overall conclusion drawn from this data is that gains are largely realized by 10 documents and saturate by 100.

Improvements for AI systems

The research presented demonstrates significant progress in knowledge distillation and few-shot learning for complex summarization tasks. However, given the high stakes (costing millions of dollars), several critical improvements must be implemented to transition these findings from academic benchmarks to robust, production-ready AI systems.

Here are the specific improvements I would make and what the resulting enhanced AI system would be capable of:


The Improvement:

Instead of relying solely on atomic fact verification (, –,) at the sentence level, the system must incorporate a multi-granularity verifier. This verifier would operate in three stages:

  1. Atomic Level (Current): Verifying single claims against source passages.

  2. Conceptual/Thematic Level: Identifying high-level rhetorical roles (e.g., Primary Conflict, Key Finding, Causation Chain) and verifying if the summary adequately covers these concepts, even if they are spread across multiple facts.

  3. Structural Coherence Level: Assessing the logical flow and narrative completeness of the summary against a structured knowledge graph derived from the source document (e.g., identifying all necessary actors, dates, and relationships).

Improved System Capability:

The system would move beyond merely reporting what was covered to guaranteeing comprehensive coverage. It could flag summaries that are factually correct but structurally incomplete (e.g., The summary mentioned the Petitioner's custody rights but failed to mention the critical reason why, which was the Respondent’s financial instability). This is crucial for legal or medical summarization where missing context is catastrophic.

  • Domain Experts: One teacher fine-tuned exclusively on legal texts, another on medical records, and a third on technical specifications.

  • Strategy Weights: The system would learn to assign weights to different inference strategies (e.g., giving higher weight to Summary Thought if the source is highly argumentative, and higher weight to Thinking if the source is purely descriptive).

  1. Extraction: Extract all entities, actions, and time markers (e.g., "A happened before B," "C caused D").

  2. Graph Construction: Build a directed graph where nodes are entities/events and edges are weighted relationships (causality, temporal sequence).

  3. Guided Generation: The final summary LLM would be prompted not just with the text, but with the structured graph, forcing it to generate narratives that respect established causality and timeline constraints.

Abstract

We show that sequence-level distillation from a capable long-context teacher model is a simple, annotation-free, and data-efficient strategy for improving argument saliency coverage in long legal opinion summarization, where small LLMs often struggle to retain the most salient argumentative content. Across student model sizes, distillation consistently surpasses tuning on expert-written summaries in our legal-opinion setting. We further demonstrate that most gains are achieved with as few as 10 training summaries, highlighting the strong data efficiency of teacher-generated supervision. Finally, we find that summary distillation is sufficient for improvements: reasoning-chain distillation remains competitive with summary-only distillation, but provides marginal benefit when combined with summary supervision.

Sources

Related papers