VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation
Yejin Jeon, Marie Maltais, Virginia Ceccatelli, Min Ma, David Ifeoluwa Adelani
Mila - Quebec AI Institute · McGill University · Google DeepMind · Canada CIFAR AI Chair
cs.SD, cs.CL
Submitted: 2026-08-11
Updated: 2026-08-12
Code: https://github.com/Lightning-AI/torchmetrics
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The paper introduces VoxSumm, a multilingual corpus and benchmark for joint speech summarization and translation (JSumT).
Terminology
Summary
The paper introduces VoxSumm, a multilingual corpus and benchmark for joint speech summarization and translation (JSumT). The authors formalize the JSumT task, which requires generating a succinct, faithful target-language summary directly from a long spoken document in a source language. The dataset comprises 10,045 BBC article-summary pairs across 24 languages, encompassing approximately 703 hours of speech data.
The paper states: "We address this methodological gap by formalizing joint speech summarization and translation (JSumT): the generation of a succinct, faithful target-language summary directly from a long spoken document in a source language. We additionally introduce VOX S UMM, the first multilingual and cross-lingual benchmark for this task, comprising 10,045 BBC article-summary pairs across 24 languages and encompassing approximately 703 hours of speech data."
The dataset construction pipeline consists of three stages: collecting cross-lingual article-summary pairs, synthesizing both source articles and target summaries into speech, and validating the generated speech through automatic and human evaluation. The textual foundation is derived from CrossSum, and speech is synthesized using OmniVoice with a fixed reference speaker. Quality control measures include discarding languages with CER above 25, and validation through NISQA scores (average 4.39) and human evaluations (average rating 4.07).
The paper evaluates three representative speech-capable large language models: Gemma4-12B, Qwen3-Omni (30B), and Gemini3.1-Pro, under three prompting settings: zero-shot, five-shot, and chain-of-thought. The evaluation uses BERTScore-F1 and xCOMET-XL metrics, with human evaluations conducted across 10 languages.
Key findings include: Gemini3.1-Pro attains the strongest average BERTScore of 0.703, followed by Qwen3-Omni (0.657) and Gemma4-12B (0.617).
The paper also finds that FS prompting proves particularly effective for Gemini3.1-Pro, reaching 0.727 and outperforming its ZS and CoT variants by approximately 0.036.
The paper reports on the effect of task direction: "this reordering produces only a minimal drop in the overall average performance across all models, but the effect is asymmetric across language directions: English→XX is affected substantially more (-0.081) than XX→English (-0.012). The authors attribute this to
instruction-following failures specific to the translate-first ordering."
Regarding language direction, the paper states: generating an English summary from non-English speech (XX→Eng) almost always outperforms generating a non-English summary from English speech (Eng→XX) across all models and both task orderings.
The paper concludes: "We have introduced JSumT, a task requiring models to compress the salient content of a long spoken document into a summary rendered in a target language, and present VOX S UMM, the first multilingual and cross-lingual benchmark for long-form speech summarization spanning 24 languages. Our evaluation reveals substantial variation: Gemini3.1-Pro performs best, few-shot prompting benefits stronger models, and generating English summaries from non-English speech is typically better than the reverse. Moreover, translating before summarizing amplifies instruction-following failures, making summarization-then-translate the more robust pipeline."
The paper also discusses limitations, noting that VOX S UMM is derived from the CrossSum dataset, whose cross-lingual pairs are identified automatically using semantic similarity between summaries
and that the speech in VOX S UMM is generated from professionally written news text using a multilingual TTS system rather than collected from naturally occurring broadcasts.
Improvements for AI systems
Improvement 1: Robust Cross-Lingual Summarization Pipeline with Direction-Aware Task Ordering
-
What to implement: Modify the inference pipeline of speech-capable LLMs to default to a summarize-then-translate ordering (source speech → source summary → target summary) rather than translate-first. Add a lightweight classifier that detects the language direction (e.g., English→X vs. X→English) and, for English→X, automatically switches to translate-first only if the model’s confidence in instruction-following is high (e.g., via logit entropy).
-
What the improved AI system can do: It will avoid the-0.081 average performance drop observed for English→XX under translate-first ordering. It will produce more faithful and succinct summaries for all 24 languages, especially for low-resource target languages, by leveraging the more robust pipeline.
Improvement 2: Few-Shot Prompting with Dynamic Example Selection for Stronger Models
-
What to implement: For models with >20B parameters (e.g., Gemini-class), automatically enable few-shot prompting with 5 examples dynamically retrieved from the VoxSumm training set based on source-language similarity and target-language pair. For smaller models (e.g., <15B), keep zero-shot to avoid context-length degradation. Add a meta-prompt that instructs the model to first generate a compressed source-language summary, then translate, using the few-shot examples as structural templates.
-
What the improved AI system can do: It will replicate the +0.036 BERTScore gain seen for Gemini3.1-Pro with few-shot prompting, while preventing smaller models from suffering context-window overflow. It will generate higher-quality summaries for users in multilingual settings (e.g., EU parliamentary or news consumption) where source and target languages vary.
Improvement 3: Asymmetric Quality Control for TTS-Based Speech Summarization
-
What to implement: Integrate a post-generation validation module that computes CER (character error rate) and NISQA scores on the synthesized speech input. If CER > 25% or NISQA < 4.0, the system automatically re-synthesizes the speech with a different reference speaker or falls back to a text-only summarization pipeline. Additionally, for English→XX directions, apply a stricter CER threshold (e.g., 20%) because the paper shows higher sensitivity to instruction-following failures in that direction.
-
What the improved AI system can do: It will reduce the risk of summarization errors caused by noisy or unnatural TTS input, particularly for non-English target languages. It will maintain high output fidelity (targeting NISQA ≥ 4.39 and human rating ≥ 4.07) even when deployed on user-provided audio that is not professionally recorded.
Improvement 4: Direction-Aware Evaluation and Calibration for Multilingual Summaries
-
What to implement: Add a calibration layer that adjusts the model’s decoding temperature and repetition penalty based on the language direction. For XX→English, use a lower temperature (e.g., 0.3) to exploit the model’s stronger performance; for English→XX, use a higher temperature (e.g., 0.7) to encourage lexical diversity in the target language. Also, compute BERTScore and xCOMET-XL separately for each direction and report them as separate metrics to avoid masking asymmetric failures.
-
What the improved AI system can do: It will produce more consistent summary quality across all 24 languages, avoiding the “English-centric” bias. It will provide users with direction-specific confidence scores, enabling them to trust the system more for non-English outputs or to request a re-summarization when the confidence is low.
Improvement 5: Instruction-Following Failure Detector for Translate-First Ordering
-
What to implement: Train a small binary classifier (e.g., a fine-tuned XLM-R) on the VoxSumm validation set to detect when a model has failed to follow the translate-first instruction (e.g., output is in the wrong language or contains untranslated source text). During inference, if this detector fires, the system automatically re-runs the model with summarize-then-translate ordering.
-
What the improved AI system can do: It will automatically recover from the instruction-following failures that the paper identifies as the cause of the-0.081 drop in English→XX. This makes the system more reliable for real-world use where users may not specify the task ordering, and it reduces the need for manual intervention or re-prompting.
Abstract
As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, long-document summarization research remains text-centric, whereas multilingual speech research has largely prioritized translation, preserving source content rather than compressing it. We address this methodological gap by formalizing joint speech summarization and translation (JSumT): the generation of a succinct, faithful target-language summary directly from a long spoken document in a source language. We additionally introduce VoxSumm, the first multilingual and cross-lingual benchmark for this task, comprising 10,045 BBC article-summary pairs across 24 languages and encompassing approximately 703 hours of speech data. Our evaluation of representative speech-language models reveals pronounced variation across models and generation settings: Gemini3.1-Pro demonstrates the greatest consistency, summarization into English generally surpasses generation into non-English target languages, and translating an entire document before summarization compounds instruction-following failures. Through the release of VoxSumm, we establish a foundation for developing and evaluating multilingual systems capable of jointly interpreting, compressing, and translating long-form speech.
Sources
- Seamless: Multilingual Expressive and Streaming Speech Translation
- PolyVoice: Language Models for Speech to Speech Translation
- Omnilingual MT: Machine Translation for 1,600 Languages
- Understanding LLM Reasoning for Abstractive Summarization
- BERTScore: Evaluating Text Generation with BERT
- OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment