SEAR: Segment-Evidence-Aware Routing for Weak-to-Strong Multilingual Speech MCQ

arXiv:2609.11355 · cs.CL, cs.SD · Submitted 2026-09-10 · Read on arXiv

cs.CL, cs.SD

Submitted: 2026-09-10

Updated: 2026-09-10

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

The gist: This paper describes our system for Task 2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge.

Terminology

Abstract

This paper describes our system for Task 2 of the second Multilingual Conversational Speech Language Model (MLC-SLM) Challenge. We adapt Qwen3-Omni-30B-A3B-Instruct with a segment-evidence-aware data and post-training pipeline. A language model converts timestamped ASR into coherent event spans, which are expanded by a boundary margin and cropped from the original recording. We then synthesize complementary semantic MCQs with Qwen3.6-27B and acoustic MCQs with Gemini 3.1 Flash-Lite, followed by structural, grounding, answer-consistency, and target-model trainability checks, yielding 359,825 verified MCQs across 21 language and accent variants. A text-only probe partitions the data into weak, text-answerable items used for supervised fine-tuning and strong, audio-dependent items used for reinforcement learning with Group Sequence Policy Optimization (GSPO), stabilized by debiased advantages, sequence-level importance correction, and dynamic filtering. Our system obtains 90.92% accuracy on the final official evaluation set.

Related papers