Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios".
Jane: The paper was written by Gerard I. Gállego, Oriol Pareras, Martí Cortada Garcia, Lucas Takanori and Javier Hernando from Barcelona Supercomputing Center and Universitat Politècnica de Catalunya.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's got a mouthful of a title — "Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios." Jane, what caught your eye first?
Jane: Honestly, Tom, the word "phoneme" in the title is what hooked me. A phoneme is basically the smallest unit of sound in a language — like the "k" sound in "cat" or the "sh" sound in "ship." The idea that they're using those as a bridge for translation is clever.
Tom: And the authors are from the Barcelona Supercomputing Center and Universitat Politècnica de Catalunya. Gerard Gállego, Oriol Pareras, Martí Cortada Garcia, Lucas Takanori, Javier Hernando. They're building on something called Salamandra, which is a multilingual language model.
Jane: Right, and that's the key. Most speech translation systems need tons of labeled audio data for every language you want to translate. But this team is asking — what if you could use phonemes to help languages that have almost no speech data at all?
Tom: Exactly. They're tackling what they call zero-resource scenarios. That means languages where you have zero labeled speech data for training. And they're showing that phonemes can step in and help.
Jane: And the "CoT" in the title — that's Chain-of-Thought. It's a technique where you break a task into intermediate steps. Instead of just translating speech directly, the model first recognizes phonemes, then transcribes them, then translates. Like showing your work in a math problem.
Tom: So instead of one giant leap from audio to English text, you get three smaller steps. And each step is easier to learn, especially when data is scarce.
Jane: Precisely. And the authors are claiming this helps in low-resource settings. But there's a trade-off — it slightly hurts performance when you have plenty of data. We'll get into that in a bit.
Tom: I'm curious about the zero-resource part. They're saying you can translate speech from a language the model has never heard before?
Jane: That's the bold claim, Tom. And they tested it by holding out three languages entirely — Dutch, Italian, and Polish. The model never saw any speech from those languages during training. And it still managed to translate them.
Tom: That's wild. So phonemes act like a universal key that unlocks speech understanding across languages?
Jane: Something like that. Because phonemes are language-agnostic — they're sounds, not words. If the model learns to recognize sounds in Spanish and German, it can generalize to Italian, even if it's never heard Italian speech.
Tom: But wait — Italian and Spanish are both Romance languages. Would it work for, say, Japanese?
Jane: That's the big question. The paper shows it works within language families — Germanic, Romance, Slavic. But the authors admit that cross-lingual transfer depends on how close the languages are. Polish, for instance, got much lower scores than Italian because there wasn't much Slavic speech data to learn from.
Tom: So the bridge only works if you have similar languages to build it from. Still, that's a huge step for languages that are currently left out of speech tech.
Jane: Absolutely. And we're just scratching the surface. Next up, we'll dig into the actual method — how they train this thing and why the phoneme step makes such a difference.
Tom: Stay with us, folks. This one's got legs.
Summary of the Paper: Jane: So, Tom, we've established that this paper — "Speech-to-Text Translation with Phoneme-Augmented CoT" — is about using phonemes as a stepping stone for translation. Now let's talk about what they actually built.
Tom: Right. They took Salamandra, which is a multilingual text-only LLM, and extended it to handle speech. They added a speech encoder that converts audio into discrete tokens, kind of like how text gets tokenized into words or subwords.
Jane: And they also added phoneme tokens to the vocabulary. So the model can now see three types of input — regular text, speech tokens, and phonemes. That's the foundation.
Tom: Then they trained it in three stages. This is the curriculum learning part. Stage one is just getting the new speech and phoneme embeddings to fit into the existing model. They freeze the backbone and only train the embedding layers.
Jane: That's smart. It's like teaching someone the alphabet before you ask them to read a sentence. The new representations need to align with what the model already knows.
Tom: Stage two is multitask training. The model learns phoneme recognition, phoneme-to-grapheme conversion, grapheme-to-phoneme, ASR, and text-to-text translation. But no speech-to-text translation yet.
Jane: So they're building all the component skills first. Then stage three is where they put it together — full speech-to-text translation with the chain-of-thought format. The model hears audio, outputs phonemes, then a transcription, then the English translation.
Tom: And they used about eight thousand five hundred hours of speech data for ASR and four hundred fifty-five hours for translation. The languages are mostly European — Catalan, German, Spanish, Russian, Swedish, Slovenian, plus the zero-resource ones.
Jane: What I love about this approach is that they generate phonemes synthetically from text. So even if you don't have audio for a language, you can still train the model to recognize its phonemes from written transcriptions.
Tom: That's the trick, right? You can get text data for almost any language. But audio data is expensive and hard to collect. Phonemes let you bridge that gap.
Jane: And they compared against two baselines. One is direct translation — no intermediate steps. The other is chain-of-thought without phonemes — just transcribe then translate. Their phoneme version beat both in low-resource settings.
Tom: But here's the catch — in high-resource languages like Catalan, German, and Spanish, the phoneme version actually did worse than the plain chain-of-thought. About one point eight BLEU points worse on average.
Jane: That's the trade-off we mentioned. When you have plenty of data, the extra phoneme step just adds a potential source of errors. But when data is scarce, it's a lifesaver.
Tom: And they found a fix for that too. They added something called Phoneme Data Augmentation — randomly corrupting phoneme sequences during training so the model doesn't become too dependent on them. That narrowed the gap in high-resource settings.
Jane: Plus they came up with a dual prompting strategy that lets you skip phonemes at inference time. And that actually gave the best results overall — even better than always using phonemes.
Tom: So the phonemes are most useful during training, but you don't necessarily need them when you're actually translating. That's a really practical insight.
Jane: It is. And it suggests the phonemes are helping the model learn a better representation of speech, even if they're not needed for every translation. Next, we'll look at those improvements in more detail.
Tom: The fixes are where it gets really interesting. Stick around.
Improvements Suggested: Jane: Welcome back. We're still on "Speech-to-Text Translation with Phoneme-Augmented CoT," and Tom, we left off with the two clever fixes they introduced. Let's break those down.
Tom: Yeah, let's do it. First up is Phoneme Data Augmentation — PDA for short. The problem they noticed was that the model could become too reliant on the phoneme step. If the phoneme recognition was wrong, the whole translation would go off the rails.
Jane: So their fix was to randomly corrupt the phoneme sequences during training. They delete some phonemes, mask others, substitute them, even insert random ones. And they shift the spaces between words in the phoneme sequence.
Tom: Why would you deliberately corrupt the data? That seems counterintuitive.
Jane: It's a regularization trick. By forcing the model to handle imperfect phonemes, you make it more robust. It learns to rely on the audio signal as well, not just the phoneme output. And it worked — they got about two point one BLEU points improvement in low-resource settings compared to the plain phoneme approach.
Tom: And in high-resource settings, the gap narrowed from one point eight points down to one point three. So it's still a small hit, but much better.
Jane: Exactly. Then they went further with the Dual Prompting Strategy — DPS. The idea is to let the model work both with and without phonemes at inference time.
Tom: So during training, they mixed it up. twenty percent of samples had no phoneme step at all. The rest was mostly PDA-augmented data. And at inference, you could choose whether to generate phonemes or skip straight to transcription.
Jane: And here's the surprising part — when they skipped phonemes at inference, they got the best results across the board. In mid and low-resource languages, they gained over four point five BLEU points compared to the baseline chain-of-thought without phonemes.
Tom: So the phonemes are like training wheels. They help the model learn to ride the bike, but once it's learned, you can take them off and go faster.
Jane: That's exactly the right analogy. The model learns a better representation of speech because it's been trained to produce phonemes. But at inference, skipping them avoids any error propagation from a bad phoneme prediction.
Tom: And in zero-resource languages — Dutch, Italian, Polish — the gains were even more dramatic. Italian got seven point nine BLEU with the dual strategy, compared to two point nine with direct translation. That's almost triple.
Jane: Italian benefits from having nearby high-resource Romance languages like Catalan and Spanish. Polish, on the other hand, only got four point seven BLEU because there wasn't much Slavic speech data to learn from.
Tom: So the method works, but it's still limited by language proximity. You can't conjure understanding out of thin air.
Jane: Right. But the fact that they got any translation at all for languages with zero speech data is remarkable. And the authors suggest future work could focus on reducing error propagation further and handling accent variability.
Tom: Accent variability — that's a good point. Phonemes can sound different depending on who's speaking and where they're from. The model needs to be robust to that.
Jane: And they also mention that the phoneme step helps bridge the gap between speech and written forms. That's a fundamental insight that could apply beyond just translation.
Tom: So what's the big picture here? Next segment, we'll wrap up and talk about what this means for the field.
Jane: And for the real world. Because this could make speech translation accessible to hundreds of languages that are currently left out.
Conclusion: Tom: Alright, we're wrapping up our discussion of "Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios." Jane, give us the final take.
Jane: The core idea is simple but powerful. By adding phoneme recognition as an intermediate step in chain-of-thought translation, you can transfer speech understanding across languages. That means you can translate from languages that have no labeled speech data at all.
Tom: And the dual prompting strategy is the cherry on top. You train with phonemes, but you can skip them at inference. That gives you the best of both worlds — the cross-lingual benefits during training, and the speed and reliability of direct translation at inference.
Jane: The trade-off is real, though. In high-resource languages, you still lose a little performance. But the gains in low-resource and zero-resource settings far outweigh that cost, especially if you care about language diversity.
Tom: And the authors are clear that this isn't a magic bullet. Cross-lingual transfer depends on how close the languages are. Polish struggled because there wasn't enough Slavic data. But for language families with decent coverage, this works remarkably well.
Jane: I think the biggest impact is on accessibility. There are thousands of languages with little to no speech data. This approach could make speech translation available for many of them, using phonemes as a universal bridge.
Tom: And it's not just translation. The idea of using phonemes as an intermediate representation could help with other speech tasks — like speech recognition for low-resource languages, or even speech-to-speech translation.
Jane: Absolutely. And the fact that they built this on an open-source model — Salamandra — means other researchers can build on their work. That's how the field moves forward.
Tom: Well said, Jane. We've covered the architecture, the training curriculum, the phoneme augmentation, and the dual prompting strategy. A solid piece of research with real-world potential.
Jane: And a big thank you to the authors — Gerard Gállego, Oriol Pareras, and the whole team at Barcelona Supercomputing Center. Great work.
Tom: That's it for this paper, folks. Join us next time when we'll be looking at another fresh arXiv submission. Until then, keep listening, keep learning.
Jane: And keep talking to your machines. They're starting to understand you better.
Gerard I. Gállego, Oriol Pareras, Martí Cortada Garcia, Lucas Takanori, Javier Hernando
Barcelona Supercomputing Center · Universitat Politècnica de Catalunya
cs.CL, cs.SD, eess.AS
Submitted: 2025-05-30
Updated: 2026-08-18
Comments: Accepted at Interspeech 2025
DOI: 10.21437/Interspeech.2025-1954
Code: https://github.com/espeak-ng/espeak-ng
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 83/100
The gist: This paper proposes a Speech-to-Text Translation (S2TT) approach that integrates phoneme representations into a Chain-of-Thought (CoT) framework to improve translation in low-resource and
Key concepts
- Phoneme
- The smallest unit of sound in a language, like the 'k' in 'cat.' The episode discusses using phonemes as a universal, language-agnostic bridge to help speech translation models understand and transfer knowledge across different languages.
- Low-Resource/Zero-Resource Scenarios
- Situations involving languages that lack large amounts of labeled audio data for training. The paper addresses this by showing that using phonemes can enable translation for languages the model has never heard before.
- Chain-of-Thought (CoT)
- A technique where a complex task, like translation, is broken down into intermediate steps. Instead of direct audio-to-text conversion, the model first recognizes phonemes, then transcribes them, and finally translates them.
Terminology
Summary
This paper proposes a Speech-to-Text Translation (S2TT) approach that integrates phoneme representations into a Chain-of-Thought (CoT) framework to improve translation in low-resource and zero-resource settings. By introducing phoneme recognition as an intermediate step, the authors enhance cross-lingual transfer, enabling translation even for languages with no labeled speech data. The system, named S ALAMANDRA-ST, is built on an open-source multilingual LLM (Salamandra 2B), which is extended to process speech and phonemes. Training follows a curriculum learning strategy that progressively introduces more complex tasks across three stages: first, aligning new speech and phoneme embeddings with the frozen backbone; second, multitask training on Phoneme Recognition (PR), Phoneme-to-Grapheme (P2G), Grapheme-to-Phoneme (G2P), Automatic Speech Recognition (ASR), and Text-to-Text Translation (T2TT); and third, fine-tuning on S2TT with a phoneme-augmented CoT (S2TT-CoT-Ph), where the model first predicts phonemes, then maps them to a transcription, and finally translates to English. The authors also propose two refinements: Phoneme Data Augmentation (PDA), which randomly deletes, masks, or substitutes phoneme spans to reduce error propagation, and Dual Prompting Strategy (DPS), which trains the model with 20% of samples excluding the phoneme step, allowing flexible inference with or without phonemes. Experiments are conducted on FLEURS, translating from nine European source languages (Catalan, German, Spanish, Russian, Swedish, Slovenian, Italian, Dutch, Polish) into English, with training data from Common Voice 17.0, VoxPopuli, NLLB, and CoVoST 2. Results show that CoT consistently outperforms direct translation, with average gains over 4.5 BLEU. Phoneme-augmented CoT improves performance in mid-/low-resource and zero-resource settings (average gain of 0.7 BLEU over CoT without phonemes) but degrades high-resource performance by an average of 1.8 BLEU. PDA amplifies benefits, yielding average gains of 2.1 BLEU in mid-/low-resource and 2.3 BLEU in zero-resource settings over CoT without phonemes, while reducing high-resource degradation to 1.3 BLEU. DPS yields the strongest results, especially when decoding without phonemes (DPS†), with average gains above 4.5 BLEU over CoT without phonemes in mid-/low-resource and zero-resource settings, and nearly closing the high-resource gap (only 0.5 BLEU degradation on average). Cross-lingual transfer depends on language proximity, with Italian benefiting from nearby high-resource Catalan and Spanish, while Polish achieves only 4.7 BLEU due to limited data in related Slavic languages. The authors conclude that phoneme-based CoT is a promising step toward making S2TT more accessible across diverse languages, with future research focusing on reducing error propagation in multi-step inference and handling accent variability.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system:
Implementation: Add a three-step inference pipeline to any speech-to-text translation system:
-
Step 1: Generate phoneme sequence from audio
-
Step 2: Convert phonemes to grapheme transcription
-
Step 3: Translate transcription to target language
Capability gained: The system can now translate speech from languages with zero labeled speech data, as long as text corpora exist for those languages.
Implementation: Train the model in this exact order:
-
Stage 1: Freeze backbone, train only new speech/phoneme embeddings on next-token prediction (1 epoch, LR=7e-5, batch=256)
-
Stage 2: Unfreeze all, train on PR, P2G, G2P, ASR, T2TT (2 epochs, LR=4e-5, batch=512)
-
Stage 3: Fine-tune on S2TT-CoT-Ph plus ASR-CoT and P2TT-CoT (1 epoch, LR=1e-5, batch=512, seq len=2048)
Implementation: During stage 3 training, apply this augmentation to 75% of samples:
-
Randomly delete, mask, or substitute phoneme spans
-
Insert random phonemes
-
Shift spaces between phonemes
-
Keep 25% of samples unmodified
Implementation: Train with a mixed data distribution:
-
20% of samples: exclude phoneme step entirely
-
75% of samples: use PDA-augmented phoneme step
-
5% of samples: original unmodified phoneme step
Implementation: Replace free-form CoT generation with:
-
Beam-search multinomial sampling, temperature=0.2, top-p=0.95, top-k=50
-
For phoneme generation specifically: reduce top-k to 10
-
Append task-specific prompts between generation stages
-
Early stopping at each stage
Implementation: Use eSpeak to synthesize phoneme sequences from text transcriptions, then train the model to map audio → phonemes → graphemes → translation, where phonemes serve as a universal intermediate representation.
Implementation: Use this specific data mix:
-
ASR: Common Voice 17.0 + VoxPopuli (8,500 hours total)
-
S2TT: CoVoST 2 (455 hours)
-
T2TT: NLLB subsampled by LASER3 alignment scores (375k samples for stage 2, next 375k for stage 3)
-
Exclude Dutch, Italian, Polish from all speech training to test zero-resource
Implementation: Use paired bootstrap resampling with p-value < 0.05 for both BLEU and COMET metrics, and report results separately for high-resource, mid/low-resource, and zero-resource language groups.
Sources
- GPT-4 Technical Report
- AudioPaLM: A Large Language Model That Can Speak and Listen
- On decoder-only architecture for speech-to-text and large language model integration
- Speech Translation with Large Language Models: An Industrial Practice
- Chain-of-Thought Prompting for Speech Translation
- Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
- Salamandra Technical Report
- Liger Kernel: Efficient Triton Kernels for LLM Training
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering