Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios

summary

Video file (mp4)

The gist

This paper proposes a Speech-to-Text Translation (S2TT) approach that integrates phoneme representations into a Chain-of-Thought (CoT) framework to improve translation in low-resource and

In short

The episode discusses a paper using phonemes to enhance speech-to-text translation, especially for low-resource languages. The hosts explain how adding an intermediate phoneme step (Chain-of-Thought) allows the model to transfer understanding across language families, even when little or no labeled speech data exists.

Key concepts

Phoneme
The smallest unit of sound in a language, like the 'k' in 'cat.' The episode discusses using phonemes as a universal, language-agnostic bridge to help speech translation models understand and transfer knowledge across different languages.
Low-Resource/Zero-Resource Scenarios
Situations involving languages that lack large amounts of labeled audio data for training. The paper addresses this by showing that using phonemes can enable translation for languages the model has never heard before.
Chain-of-Thought (CoT)
A technique where a complex task, like translation, is broken down into intermediate steps. Instead of direct audio-to-text conversion, the model first recognizes phonemes, then transcribes them, and finally translates them.

Terminology used across episodes

This episode discusses

The paper

Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios · Read on arXiv

Gerard I. Gállego, Oriol Pareras, Martí Cortada Garcia, Lucas Takanori, Javier Hernando

Barcelona Supercomputing Center · Universitat Politècnica de Catalunya

DOI: 10.21437/Interspeech.2025-1954

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios".

Jane: The paper was written by Gerard I. Gállego, Oriol Pareras, Martí Cortada Garcia, Lucas Takanori and Javier Hernando from Barcelona Supercomputing Center and Universitat Politècnica de Catalunya.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a mouthful of a title — "Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios." Jane, what caught your eye first?

Jane: Honestly, Tom, the word "phoneme" in the title is what hooked me. A phoneme is basically the smallest unit of sound in a language — like the "k" sound in "cat" or the "sh" sound in "ship." The idea that they're using those as a bridge for translation is clever.

Tom: And the authors are from the Barcelona Supercomputing Center and Universitat Politècnica de Catalunya. Gerard Gállego, Oriol Pareras, Martí Cortada Garcia, Lucas Takanori, Javier Hernando. They're building on something called Salamandra, which is a multilingual language model.

Jane: Right, and that's the key. Most speech translation systems need tons of labeled audio data for every language you want to translate. But this team is asking — what if you could use phonemes to help languages that have almost no speech data at all?

Tom: Exactly. They're tackling what they call zero-resource scenarios. That means languages where you have zero labeled speech data for training. And they're showing that phonemes can step in and help.

Jane: And the "CoT" in the title — that's Chain-of-Thought. It's a technique where you break a task into intermediate steps. Instead of just translating speech directly, the model first recognizes phonemes, then transcribes them, then translates. Like showing your work in a math problem.

Tom: So instead of one giant leap from audio to English text, you get three smaller steps. And each step is easier to learn, especially when data is scarce.

Jane: Precisely. And the authors are claiming this helps in low-resource settings. But there's a trade-off — it slightly hurts performance when you have plenty of data. We'll get into that in a bit.

Tom: I'm curious about the zero-resource part. They're saying you can translate speech from a language the model has never heard before?

Jane: That's the bold claim, Tom. And they tested it by holding out three languages entirely — Dutch, Italian, and Polish. The model never saw any speech from those languages during training. And it still managed to translate them.

Tom: That's wild. So phonemes act like a universal key that unlocks speech understanding across languages?

Jane: Something like that. Because phonemes are language-agnostic — they're sounds, not words. If the model learns to recognize sounds in Spanish and German, it can generalize to Italian, even if it's never heard Italian speech.

Tom: But wait — Italian and Spanish are both Romance languages. Would it work for, say, Japanese?

Jane: That's the big question. The paper shows it works within language families — Germanic, Romance, Slavic. But the authors admit that cross-lingual transfer depends on how close the languages are. Polish, for instance, got much lower scores than Italian because there wasn't much Slavic speech data to learn from.

Tom: So the bridge only works if you have similar languages to build it from. Still, that's a huge step for languages that are currently left out of speech tech.

Jane: Absolutely. And we're just scratching the surface. Next up, we'll dig into the actual method — how they train this thing and why the phoneme step makes such a difference.

Tom: Stay with us, folks. This one's got legs.

Summary of the Paper: Jane: So, Tom, we've established that this paper — "Speech-to-Text Translation with Phoneme-Augmented CoT" — is about using phonemes as a stepping stone for translation. Now let's talk about what they actually built.

Tom: Right. They took Salamandra, which is a multilingual text-only LLM, and extended it to handle speech. They added a speech encoder that converts audio into discrete tokens, kind of like how text gets tokenized into words or subwords.

Jane: And they also added phoneme tokens to the vocabulary. So the model can now see three types of input — regular text, speech tokens, and phonemes. That's the foundation.

Tom: Then they trained it in three stages. This is the curriculum learning part. Stage one is just getting the new speech and phoneme embeddings to fit into the existing model. They freeze the backbone and only train the embedding layers.

Jane: That's smart. It's like teaching someone the alphabet before you ask them to read a sentence. The new representations need to align with what the model already knows.

Tom: Stage two is multitask training. The model learns phoneme recognition, phoneme-to-grapheme conversion, grapheme-to-phoneme, ASR, and text-to-text translation. But no speech-to-text translation yet.

Jane: So they're building all the component skills first. Then stage three is where they put it together — full speech-to-text translation with the chain-of-thought format. The model hears audio, outputs phonemes, then a transcription, then the English translation.

Tom: And they used about eight thousand five hundred hours of speech data for ASR and four hundred fifty-five hours for translation. The languages are mostly European — Catalan, German, Spanish, Russian, Swedish, Slovenian, plus the zero-resource ones.

Jane: What I love about this approach is that they generate phonemes synthetically from text. So even if you don't have audio for a language, you can still train the model to recognize its phonemes from written transcriptions.

Tom: That's the trick, right? You can get text data for almost any language. But audio data is expensive and hard to collect. Phonemes let you bridge that gap.

Jane: And they compared against two baselines. One is direct translation — no intermediate steps. The other is chain-of-thought without phonemes — just transcribe then translate. Their phoneme version beat both in low-resource settings.

Tom: But here's the catch — in high-resource languages like Catalan, German, and Spanish, the phoneme version actually did worse than the plain chain-of-thought. About one point eight BLEU points worse on average.

Jane: That's the trade-off we mentioned. When you have plenty of data, the extra phoneme step just adds a potential source of errors. But when data is scarce, it's a lifesaver.

Tom: And they found a fix for that too. They added something called Phoneme Data Augmentation — randomly corrupting phoneme sequences during training so the model doesn't become too dependent on them. That narrowed the gap in high-resource settings.

Jane: Plus they came up with a dual prompting strategy that lets you skip phonemes at inference time. And that actually gave the best results overall — even better than always using phonemes.

Tom: So the phonemes are most useful during training, but you don't necessarily need them when you're actually translating. That's a really practical insight.

Jane: It is. And it suggests the phonemes are helping the model learn a better representation of speech, even if they're not needed for every translation. Next, we'll look at those improvements in more detail.

Tom: The fixes are where it gets really interesting. Stick around.

Improvements Suggested: Jane: Welcome back. We're still on "Speech-to-Text Translation with Phoneme-Augmented CoT," and Tom, we left off with the two clever fixes they introduced. Let's break those down.

Tom: Yeah, let's do it. First up is Phoneme Data Augmentation — PDA for short. The problem they noticed was that the model could become too reliant on the phoneme step. If the phoneme recognition was wrong, the whole translation would go off the rails.

Jane: So their fix was to randomly corrupt the phoneme sequences during training. They delete some phonemes, mask others, substitute them, even insert random ones. And they shift the spaces between words in the phoneme sequence.

Tom: Why would you deliberately corrupt the data? That seems counterintuitive.

Jane: It's a regularization trick. By forcing the model to handle imperfect phonemes, you make it more robust. It learns to rely on the audio signal as well, not just the phoneme output. And it worked — they got about two point one BLEU points improvement in low-resource settings compared to the plain phoneme approach.

Tom: And in high-resource settings, the gap narrowed from one point eight points down to one point three. So it's still a small hit, but much better.

Jane: Exactly. Then they went further with the Dual Prompting Strategy — DPS. The idea is to let the model work both with and without phonemes at inference time.

Tom: So during training, they mixed it up. twenty percent of samples had no phoneme step at all. The rest was mostly PDA-augmented data. And at inference, you could choose whether to generate phonemes or skip straight to transcription.

Jane: And here's the surprising part — when they skipped phonemes at inference, they got the best results across the board. In mid and low-resource languages, they gained over four point five BLEU points compared to the baseline chain-of-thought without phonemes.

Tom: So the phonemes are like training wheels. They help the model learn to ride the bike, but once it's learned, you can take them off and go faster.

Jane: That's exactly the right analogy. The model learns a better representation of speech because it's been trained to produce phonemes. But at inference, skipping them avoids any error propagation from a bad phoneme prediction.

Tom: And in zero-resource languages — Dutch, Italian, Polish — the gains were even more dramatic. Italian got seven point nine BLEU with the dual strategy, compared to two point nine with direct translation. That's almost triple.

Jane: Italian benefits from having nearby high-resource Romance languages like Catalan and Spanish. Polish, on the other hand, only got four point seven BLEU because there wasn't much Slavic speech data to learn from.

Tom: So the method works, but it's still limited by language proximity. You can't conjure understanding out of thin air.

Jane: Right. But the fact that they got any translation at all for languages with zero speech data is remarkable. And the authors suggest future work could focus on reducing error propagation further and handling accent variability.

Tom: Accent variability — that's a good point. Phonemes can sound different depending on who's speaking and where they're from. The model needs to be robust to that.

Jane: And they also mention that the phoneme step helps bridge the gap between speech and written forms. That's a fundamental insight that could apply beyond just translation.

Tom: So what's the big picture here? Next segment, we'll wrap up and talk about what this means for the field.

Jane: And for the real world. Because this could make speech translation accessible to hundreds of languages that are currently left out.

Conclusion: Tom: Alright, we're wrapping up our discussion of "Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios." Jane, give us the final take.

Jane: The core idea is simple but powerful. By adding phoneme recognition as an intermediate step in chain-of-thought translation, you can transfer speech understanding across languages. That means you can translate from languages that have no labeled speech data at all.

Tom: And the dual prompting strategy is the cherry on top. You train with phonemes, but you can skip them at inference. That gives you the best of both worlds — the cross-lingual benefits during training, and the speed and reliability of direct translation at inference.

Jane: The trade-off is real, though. In high-resource languages, you still lose a little performance. But the gains in low-resource and zero-resource settings far outweigh that cost, especially if you care about language diversity.

Tom: And the authors are clear that this isn't a magic bullet. Cross-lingual transfer depends on how close the languages are. Polish struggled because there wasn't enough Slavic data. But for language families with decent coverage, this works remarkably well.

Jane: I think the biggest impact is on accessibility. There are thousands of languages with little to no speech data. This approach could make speech translation available for many of them, using phonemes as a universal bridge.

Tom: And it's not just translation. The idea of using phonemes as an intermediate representation could help with other speech tasks — like speech recognition for low-resource languages, or even speech-to-speech translation.

Jane: Absolutely. And the fact that they built this on an open-source model — Salamandra — means other researchers can build on their work. That's how the field moves forward.

Tom: Well said, Jane. We've covered the architecture, the training curriculum, the phoneme augmentation, and the dual prompting strategy. A solid piece of research with real-world potential.

Jane: And a big thank you to the authors — Gerard Gállego, Oriol Pareras, and the whole team at Barcelona Supercomputing Center. Great work.

Tom: That's it for this paper, folks. Join us next time when we'll be looking at another fresh arXiv submission. Until then, keep listening, keep learning.

Jane: And keep talking to your machines. They're starting to understand you better.

More episodes

← Home