Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs".
Jane: Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities, focusing on English-Mandarin code-switching as a case study.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, to start off, let’s talk about who wrote this; we have Quang Trung, Cheng Yi Lewis Won, Minh Duc Pham, Yingxu He, Shuo Sun, and Ai Ti Aw all coming from the Institute for Infocomm Research at I2R in Singapore.
Jane: It's impressive that a team from such an established institute is tackling something so nuanced like code-switching speech recognition.
Lu: I think what’s really cool about the authors is how they are focusing on English-Mandarin code-switching specifically, which shows a deep dive into one of the most complex language mixing scenarios.
Meng: From an engineering standpoint, focusing on a specific pair like English and Mandarin helps them define exactly where the failure points are instead of trying to solve everything at once across dozens of languages.
Lalam: And they propose using Direct Preference Optimization, which is a clever way to teach the AI what the *right* way to handle that mix actually looks in practice.
The paper's summary: Tom: Okay, moving on to the main summary of "Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs," it boils down to this: they found three specific ways these large models mess up code-switching.
Jane: Those three modes are language omission, where the model just drops one language entirely, translation instead of transcription, and hallucination where it makes stuff up.
Lu: That’s a clear breakdown because it shows that the failures aren't random noise; they follow patterns based on how the AI interprets those mixed inputs.
Meng: And then they build their solution by creating preference pairs: ground truth transcriptions as the good examples, and flawed transcriptions mimicking those three failure modes as the bad ones.
Lalam: This whole DPO approach is brilliant because it directly trains the model to prefer preserving that actual mixed-language composition instead of defaulting to a single language translation.
The paper's improvements: Tom: Now, let’s talk about the actual technical improvements they found in applying this method, specifically how it shifts the models' behavior after DPO training.
Jane: The core improvement is that the AI learns to actually keep the mixed-language pattern intact instead of just translating everything away when it hears a code-switched utterance.
Lu: They showed that this alignment process isn't just theoretical; they achieved consistent behavioral shifts across three different Audio LLM architectures they tested.
Meng: The results are pretty striking, showing relative MER reductions up to eighty-nine point six percent in the in-distribution benchmarks and even twenty point zero percent on the out-of-distribution ones for some models.
Lalam: That level of reduction is huge because it means we’re seeing a massive correction in how the model generates these complex transcriptions, which is a major win for usability.
Conclusion: Tom: So, to wrap things up with "Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs," the big implication is that DPO is a viable mechanism to align models that already have strong multilingual capabilities.
Jane: It shows we don't necessarily need to start training from scratch on massive amounts of new code-switched data; we can nudge existing powerful AI toward the right behavior with targeted preference signals.
Lu: This work opens up possibilities for creating much more robust and culturally aware AI systems that can genuinely understand and process natural, mixed speech patterns in real-world conversations.
Meng: Practically speaking, this means ASR systems will be far more reliable when deployed in diverse environments where people naturally switch between languages without warning.
Lalam: I think the biggest impact is on how we interact with AI; it moves us closer to having a system that accurately reflects the rich, mixed reality of human communication.
Tom: Incredible stuff! We’ve seen how DPO tackles those three failure modes and drives those incredible relative MER reductions across the board. That’s a massive leap forward for multimodal audio processing.
Jane: It certainly is, Tom; this research on Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs really shows that alignment techniques can unlock real capabilities in these complex areas.
Lu: It’s a fascinating direction, and I think seeing how they used Global Translation versus Partial Translation strategies to mimic the failures is a very insightful design choice.
Meng: From my side, the practical application is huge; imagine voice assistants that don't drop half of what you say when you switch languages mid-sentence.
Lalam: And for culture, this means AI can finally capture the nuance of how people actually communicate across language boundaries instead of just seeing it as separate words.
Institute for Infocomm Research (I2R), A⋆ STAR, Singapore · Nanyang Technological University, Singapore
cs.CL, cs.SD
Submitted: 2026-05-13
Updated: 2026-10-01
Importance score: 82/100
The gist: Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities, focusing on English-Mandarin code-switching as a
Key concepts
- Direct Preference Optimization (DPO)
- A technique proposed by the authors that uses preference pairs—ground truth transcriptions as good examples and flawed transcriptions as bad ones—to directly train the AI model. This teaches the model what the correct way to handle mixed speech actually looks in practice.
- Code-Switching Speech Recognition
- The specific area of study where Audio LLMs show systematic failures when transcribing speech that mixes two languages, specifically English and Mandarin. The paper focuses on fixing these complex language mixing scenarios.
- Language Omission
- One of the three ways the large models mess up code-switching. This failure mode occurs when the model simply drops one of the languages entirely from the transcription during processing.
- Relative MER Reduction
- A measure used to show improvement in performance. The researchers achieved relative MER reductions up to eighty-nine point six percent on in-distribution benchmarks and twenty point zero percent on out-of-distribution ones, indicating massive correction.
Terminology
Summary
Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities, focusing on English-Mandarin code-switching as a case study. The paper identifies three failure modes: "(1) language omission, where the model outputs only one language while dropping the other; (2) translation-instead-of-transcription, where the model translates mixed-language content into a single language rather than preserving the original; and (3) hallucination, where the model generates repeated or fabricated content."
The authors hypothesize that Audio LLMs possess a latent ability to produce correct code-switching transcriptions, and this behavior can be elicited through Direct Preference Optimization (DPO). To test this hypothesis, they construct DPO training pairs by pairing ground-truth code-switching transcriptions (chosen) with synthetically generated flawed transcriptions (rejected) that mimic the observed failure modes. They employ two complementary strategies targeting translation-based failure modes: "Global Translation (80%): This strategy translates all content from one language to the other (all Chinese→English or all English→Chinese), thereby mimicking translation-instead-of-transcription failures. Partial Translation (20%): In contrast, this strategy translates only specific short spans within the utterance, mimicking partial language omission where isolated segments are incorrectly rendered in the wrong language."
They construct DPO training pairs from two complementary datasets: "CS-Dialogue [24]: This contains spontaneous Mandarin English code-switching dialogues from 200 speakers, where each utterance is tagged as English-only (EN), Chinese-only (CN), or code-switched (MIX). From this data, we construct segments in two ways: first, by grouping consecutive MIX-tagged utterances containing natural intra-sentential codeswitching; and second, by concatenating EN and CN utterances from the same conversation to create inter-sentential code-switching. EMILIA [25]: To complement CS-Dialogue with additional scale and diversity, we create synthetic code-switching samples by concatenating English and Chinese clips from the EMILIA corpus. Each segment is constructed by randomly sampling clips from both languages and concatenating them, producing inter-sentential code-switching audio at scale."
They apply DPO to align Audio LLM output behavior for codeswitching transcription, optimizing the policy πθ to increase the likelihood of generating yc (ground-truth transcription) while decreasing the likelihood of generating yr (flawed transcription), directly without explicit reward modeling
using the LDPO objective function. They train three Audio LLMs—MERaLiON-2-3B [7], Phi-4-multimodalinstruct [6], and Qwen2-Audio-7B-Instruct [2]—on approximately 100K preference pairs (∼570 hours) derived from natural and synthetic code-switching data.
The results demonstrate consistent improvements across all models and benchmarks. MERaLiON-2-3B shows modest SEAME improvements (0.7–2.0%), which reflects that this model already incorporates extensive code-switching data in its supervised training, leaving limited room for further gains.
However, Phi-4-multimodalinstruct, on the other hand, shows dramatic improvement: MER drops from 70.98% to 7.38% on EMILIA (89.6% relative reduction).
Furthermore, Qwen2-Audio-7B-Instruct demonstrates substantial improvement on SEAME dev man (20.0% relative), alongside consistent gains on in-distribution benchmarks (5.9–19.3%).
Qualitative analysis confirms that DPO effectively corrects all three failure modes, showing that DPO shifts generation toward the desired behavior: models become more likely to preserve the mixed-language pattern and produce more stable transcriptions.
The authors conclude that this work shows "among the first demonstrations that Direct Preference Optimization can elicit correct code-switching transcription behavior from multilingual Audio LLMs, and it offers a viable mechanism for aligning models that already possess multilingual capability. They report relative MER reductions of
up to 89.6% (in-distribution) and 20.0% (out-of-distribution)."
The paper's contributions are summarized as:
• We identify three systematic failure modes in English-Mandarin code-switching transcription exhibited by state-of-the-art multilingual Audio LLMs.
• We propose a DPO approach to construct preference pairs that contrast correct code-switching transcriptions with failure-mimicking alternatives.
"• We demonstrate consistent improvements across three Audio LLM architectures on English-Mandarin benchmarks, achieving relative MER reductions of up to 20.0% (out-ofdistribution) and 89.6% (in-distribution).
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on this research, along with what those improved systems can achieve:
-
Improve multilingual Automatic Speech Recognition (ASR) performance for code-switching scenarios by implementing Direct Preference Optimization (DPO).
-
Develop a mechanism to systematically address three core failure modes in English-Mandarin code-switching transcription: language omission, translation instead of transcription, and hallucination.
-
Enable Audio LLMs to preserve the original mixed-language composition rather than translating all content into a single language when prompted for transcription.
-
Improve the robustness of ASR systems against catastrophic errors (like repetition or fabrication) during complex code-switched utterances by aligning generation patterns toward ground-truth transcriptions.
-
Create a lightweight, preference-based alignment layer that can be applied to existing multilingual Audio LLMs (such as MERaLiON, Qwen2-Audio, or Phi-4 Multimodal) to elicit correct code-switching transcription behavior without requiring extensive retraining on new code-switched data.
-
Enhance the generalization of ASR models across different language pairs and speech contexts by developing a transferrable DPO framework that leverages synthetic preference pairs generated via existing large multilingual corpora (like EMILIA).
The improved AI systems can specifically:
-
Accurately transcribe spontaneous English-Mandarin code-switching dialogues, maintaining both languages within the transcription.
-
Produce verbatim transcriptions of mixed utterances, eliminating the tendency to translate entire segments into a single language.
-
Generate stable and factually accurate transcriptions that avoid repetitive or fabricated content when faced with complex linguistic mixing.
-
Perform reliably on out-of-distribution code-switching benchmarks (like SEAME dev man/sge) by correcting latent behavioral biases identified during the DPO alignment phase.
Sources
- Qwen2-Audio Technical Report
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- Qwen2.5-Omni Technical Report
- Qwen3-Omni Technical Report
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Qwen3 Technical Report
- CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering