Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

summary

Video file (mp4)

The gist

Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities, focusing on English-Mandarin code-switching as a

In short

The episode discusses a paper on using Direct Preference Optimization (DPO) to improve English-Mandarin code-switching speech recognition in Audio LLMs. The research addresses three failure modes: language omission, translation instead of transcription, and hallucination. DPO trains the model to prefer preserving the actual mixed-language composition, leading to significant reductions in error rates across different models.

Key concepts

Direct Preference Optimization (DPO)
A technique proposed by the authors that uses preference pairs—ground truth transcriptions as good examples and flawed transcriptions as bad ones—to directly train the AI model. This teaches the model what the correct way to handle mixed speech actually looks in practice.
Code-Switching Speech Recognition
The specific area of study where Audio LLMs show systematic failures when transcribing speech that mixes two languages, specifically English and Mandarin. The paper focuses on fixing these complex language mixing scenarios.
Language Omission
One of the three ways the large models mess up code-switching. This failure mode occurs when the model simply drops one of the languages entirely from the transcription during processing.
Relative MER Reduction
A measure used to show improvement in performance. The researchers achieved relative MER reductions up to eighty-nine point six percent on in-distribution benchmarks and twenty point zero percent on out-of-distribution ones, indicating massive correction.

Terminology used across episodes

This episode discusses

The paper

Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs · Read on arXiv

Institute for Infocomm Research (I2R), A⋆ STAR, Singapore · Nanyang Technological University, Singapore

Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities. Focusing on English-Mandarin, we identify three failure modes: language omission, translation-instead-of-transcription, and hallucination. We apply Direct Preference Optimization (DPO) to align models, constructing preference pairs in which chosen responses preserve mixed-language content while rejected responses mimic failure patterns. Training three Audio LLMs on 100K pairs (570 hours), we observe consistent behavioral shifts: models learn to preserve language composition rather than translating when prompted for transcription. This alignment yields MER reductions up to 89.6% (in-distribution) and 20.0% (out-of-distribution). Our findings suggest DPO can effectively elicit correct code-switching transcription behavior from multilingual Audio LLMs.

DOI: 10.21437/Interspeech.2026-110

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs".

Jane: Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities, focusing on English-Mandarin code-switching as a case study.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, to start off, let’s talk about who wrote this; we have Quang Trung, Cheng Yi Lewis Won, Minh Duc Pham, Yingxu He, Shuo Sun, and Ai Ti Aw all coming from the Institute for Infocomm Research at I2R in Singapore.

Jane: It's impressive that a team from such an established institute is tackling something so nuanced like code-switching speech recognition.

Lu: I think what’s really cool about the authors is how they are focusing on English-Mandarin code-switching specifically, which shows a deep dive into one of the most complex language mixing scenarios.

Meng: From an engineering standpoint, focusing on a specific pair like English and Mandarin helps them define exactly where the failure points are instead of trying to solve everything at once across dozens of languages.

Lalam: And they propose using Direct Preference Optimization, which is a clever way to teach the AI what the *right* way to handle that mix actually looks in practice.

The paper's summary: Tom: Okay, moving on to the main summary of "Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs," it boils down to this: they found three specific ways these large models mess up code-switching.

Jane: Those three modes are language omission, where the model just drops one language entirely, translation instead of transcription, and hallucination where it makes stuff up.

Lu: That’s a clear breakdown because it shows that the failures aren't random noise; they follow patterns based on how the AI interprets those mixed inputs.

Meng: And then they build their solution by creating preference pairs: ground truth transcriptions as the good examples, and flawed transcriptions mimicking those three failure modes as the bad ones.

Lalam: This whole DPO approach is brilliant because it directly trains the model to prefer preserving that actual mixed-language composition instead of defaulting to a single language translation.

The paper's improvements: Tom: Now, let’s talk about the actual technical improvements they found in applying this method, specifically how it shifts the models' behavior after DPO training.

Jane: The core improvement is that the AI learns to actually keep the mixed-language pattern intact instead of just translating everything away when it hears a code-switched utterance.

Lu: They showed that this alignment process isn't just theoretical; they achieved consistent behavioral shifts across three different Audio LLM architectures they tested.

Meng: The results are pretty striking, showing relative MER reductions up to eighty-nine point six percent in the in-distribution benchmarks and even twenty point zero percent on the out-of-distribution ones for some models.

Lalam: That level of reduction is huge because it means we’re seeing a massive correction in how the model generates these complex transcriptions, which is a major win for usability.

Conclusion: Tom: So, to wrap things up with "Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs," the big implication is that DPO is a viable mechanism to align models that already have strong multilingual capabilities.

Jane: It shows we don't necessarily need to start training from scratch on massive amounts of new code-switched data; we can nudge existing powerful AI toward the right behavior with targeted preference signals.

Lu: This work opens up possibilities for creating much more robust and culturally aware AI systems that can genuinely understand and process natural, mixed speech patterns in real-world conversations.

Meng: Practically speaking, this means ASR systems will be far more reliable when deployed in diverse environments where people naturally switch between languages without warning.

Lalam: I think the biggest impact is on how we interact with AI; it moves us closer to having a system that accurately reflects the rich, mixed reality of human communication.

Tom: Incredible stuff! We’ve seen how DPO tackles those three failure modes and drives those incredible relative MER reductions across the board. That’s a massive leap forward for multimodal audio processing.

Jane: It certainly is, Tom; this research on Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs really shows that alignment techniques can unlock real capabilities in these complex areas.

Lu: It’s a fascinating direction, and I think seeing how they used Global Translation versus Partial Translation strategies to mimic the failures is a very insightful design choice.

Meng: From my side, the practical application is huge; imagine voice assistants that don't drop half of what you say when you switch languages mid-sentence.

Lalam: And for culture, this means AI can finally capture the nuance of how people actually communicate across language boundaries instead of just seeing it as separate words.

More episodes

← Home