Retrieved-Span Training for Efficient Query-Focused Meeting Summarization on QMSum
cs.CL
Submitted: 2026-08-18
Updated: 2026-08-18
Comments: 24 pages, 4 figures
Code: https://github.com/ErtasAI/qmsum-retrieved-span-training
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: QMSum provides no scorer, making query-focused meeting summarization results difficult to compare.
Terminology
Abstract
QMSum provides no scorer, making query-focused meeting summarization results difficult to compare. We rescore or generate 15 systems under one implementation. Through a common inference port, a released 406M Fusion-in-Decoder specialist loses 6.30 ROUGE-1 when moved from capped long input to 2,000-word retrieved spans. Fine-tuning it on this span regime recovers the loss. On test it scores 36.33 ROUGE-1 versus 35.41 for our 1.2B system; the meeting-cluster 95% interval for the difference is [-0.27, +2.22], so QMSum does not statistically separate them. The smaller system uses about one-third as many total parameters and less than half the peak inference memory. Within the fixed 1.2B base, span-regime fine-tuning adds 5.29 [+4.02, +6.56], while replacing the first 4,500 transcript words with 2,000 retrieved words adds 1.55 on test and 0.29 on validation. Separately, under one concise prompt and reference-overlap scorer, a released 406M specialist exceeds five proprietary hosted models by at least 6.2 ROUGE-1, but output length and absent human or factuality evaluation limit this ordering. Conclusions are limited to QMSum and automatic metrics.
Sources
- Re-evaluating Evaluation in Text Summarization
- Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification
- With Little Power Comes Great Responsibility
- QLoRA: Efficient Finetuning of Quantized LLMs
- Small or Large? Zero-Shot or Finetuned? Guiding Language Model Choice for Specialized Applications in Healthcare
- LoRA: Low-Rank Adaptation of Large Language Models
- Neural Text Summarization: A Critical Evaluation
- Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach
- Learning to Rank Utterances for Query-Focused Meeting Summarization
- DYLE: Dynamic Latent Extraction for Abstractive Long-Input Summarization
- Socratic Pretraining: Question-Driven Pretraining for Controllable Summarization
- QontSum: On Contrasting Salient Content for Query-focused Summarization
- Learning to Rank Salient Content for Query-focused Summarization
- Exploring Neural Models for Query-Focused Summarization
- Retrieval meets Long Context Large Language Models
- BERTScore: Evaluating Text Generation with BERT
- Summ^N: A Multi-Stage Summarization Framework for Long Input Dialogues and Documents
- QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization
- DialogLM: Pre-trained Model for Long Dialogue Understanding and Summarization
- A Hierarchical Network for Abstractive Meeting Summarization with Cross-Domain Pretraining
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering