Retrieved-Span Training for Efficient Query-Focused Meeting Summarization on QMSum

arXiv:2609.25028 · cs.CL · Submitted 2026-08-18 · Read on arXiv

cs.CL

Submitted: 2026-08-18

Updated: 2026-08-18

Comments: 24 pages, 4 figures

Code: https://github.com/ErtasAI/qmsum-retrieved-span-training

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

The gist: QMSum provides no scorer, making query-focused meeting summarization results difficult to compare.

Terminology

Abstract

QMSum provides no scorer, making query-focused meeting summarization results difficult to compare. We rescore or generate 15 systems under one implementation. Through a common inference port, a released 406M Fusion-in-Decoder specialist loses 6.30 ROUGE-1 when moved from capped long input to 2,000-word retrieved spans. Fine-tuning it on this span regime recovers the loss. On test it scores 36.33 ROUGE-1 versus 35.41 for our 1.2B system; the meeting-cluster 95% interval for the difference is [-0.27, +2.22], so QMSum does not statistically separate them. The smaller system uses about one-third as many total parameters and less than half the peak inference memory. Within the fixed 1.2B base, span-regime fine-tuning adds 5.29 [+4.02, +6.56], while replacing the first 4,500 transcript words with 2,000 retrieved words adds 1.55 on test and 0.29 on validation. Separately, under one concise prompt and reference-overlap scorer, a released 406M specialist exceeds five proprietary hosted models by at least 6.2 ROUGE-1, but output length and absent human or factuality evaluation limit this ordering. Conclusions are limited to QMSum and automatic metrics.

Sources

Related papers