Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
summary
The gist
The paper investigates the scaling behavior of two prompting strategies for End-to-End (E2E) Speech-to-Text Translation (S2TT) using LLM-based models: Chain-of-Thought (CoT) prompting, where the
In short
The episode discusses a paper comparing two methods for speech-to-text translation: Direct prompting and Chain-of-Thought (CoT). The hosts conclude that while CoT is popular, Direct prompting scales better with more data, showing potential for richer translations and lower computational costs.
Key concepts
- Speech-to-Text Translation
- This process involves taking spoken audio in one language (e.g., Spanish) and having a system output the translated text in another language (e.g., English).
- Direct Prompting
- A method where a model hears the audio and immediately translates it into text, performing the translation in a single step without an intermediate transcription.
- Chain-of-Thought (CoT)
- An approach where the model first transcribes speech into text, and then uses that resulting text to perform the actual translation. It is a two-step process.
- Pseudo-labeling
- A technique used by the hosts to create training data. They took an existing audio corpus (Common Voice) and translated its transcriptions into multiple languages using an LLM.
Terminology used across episodes
This episode discusses
- Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting? · Paper Radio
- AudioPaLM: A Large Language Model That Can Speak and Listen
- Speech Translation with Large Language Models: An Industrial Practice
- Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
- Deep Speech: Scaling up end-to-end speech recognition
The paper
Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting? · Read on arXiv
Oriol Pareras, Gerard I. Gállego, Federico Costa, Cristina España-Bonet, Javier Hernando
Barcelona Supercomputing Center · Universitat Politècnica de Catalunya · DFKI GmbH · Saarland Informatics Campus
DOI: 10.1109/ICASSP55912.2026.11464161
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?".
Jane: The paper was written by Oriol Pareras, Gerard I. Gállego, Federico Costa, Cristina España-Bonet and Javier Hernando from Barcelona Supercomputing Center and Universitat Politècnica de Catalunya and DFKI GmbH and Saarland Informatics Campus.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everybody. Today we’re digging into a paper that’s been making the rounds in the speech translation world, and it’s called “Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?” Jane, I have to say, that title alone got me excited.
Jane: Oh, absolutely, Tom. And for our listeners who might not be deep in the weeds here, let’s break down what that title actually means. Speech-to-text translation is when you hear someone speaking in, say, Spanish, and the system outputs English text. The big question in the field right now is how to build those systems.
Tom: Right, and the paper is comparing two ways of doing it. One is called Direct prompting, where the model just hears the audio and immediately translates it. The other is Chain-of-Thought, or CoT, where the model first transcribes the speech into text, and then translates that text. It’s like doing two steps instead of one.
Jane: And that CoT approach has been really popular lately because it lets you use tons of existing data. You have speech-to-text data and text-to-text translation data, and you can kind of stack them together. So the assumption has been that CoT is just better. But this paper from Barcelona Supercomputing Center and UPC is questioning that.
Tom: Questioning it with actual experiments, which is the best part. They built these speech LLMs, trained them with both approaches, and then scaled up the amount of translation data they fed them. And what they found is that Direct prompting gets better and better as you add more data, while CoT kind of plateaus and even degrades.
Jane: It’s a really important finding because the field has been moving toward CoT as the default. The authors are basically saying, hold on, maybe we’ve been optimizing for the wrong thing. If we keep building bigger datasets, Direct might actually win.
Tom: And that’s the hook for me, Jane. The title asks a question, and the answer is a cautious yes. Direct prompting scales better. We’re going to get into the nitty-gritty of how they tested that in a moment.
Jane: Stay with us, because this has big implications for how we build speech translation systems in the future.
Summary: Tom: So we’ve set the stage. Let’s get into the actual summary of “Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?” Jane, walk us through what they actually did.
Jane: Gladly. So they took a multilingual LLM called Salamandra, which is fine-tuned for translation, and they hooked it up to a speech encoder called mHuBERT. That encoder turns audio into discrete tokens that the LLM can understand. Then they trained it on a bunch of different data types.
Tom: And here’s the clever part. They didn’t have enough real speech-to-text translation data, so they made their own. They took a huge ASR corpus, Common Voice, which has audio with transcriptions, and they translated those transcriptions into six European languages using the same LLM. That’s the pseudo-labeling pipeline.
Jane: Exactly. They filtered the results using quality estimation and language identification to make sure the translations were actually good and in the right language. Then they trained models with both Direct and CoT prompting, adding increasing amounts of this pseudo-labeled data. They went from twenty percent up to one hundred percent.
Tom: And the results, Jane? I was honestly surprised. The CoT baseline starts off better, no question. At zero added data, CoT beats Direct by about five BLEU points and seven xCOMET points. But then you start adding pseudo-labeled data, and Direct just keeps climbing.
Jane: While CoT peaks at twenty percent of the data and then starts falling. The paper shows that CoT’s transcription quality actually gets worse as you add more translation data. The WER, that’s the word error rate for transcription, goes up by as much as thirteen percent relative to baseline.
Tom: That’s a huge deal. It means the intermediate transcription step, which is supposed to be the strength of CoT, is actually becoming a bottleneck. The model is getting confused trying to do both tasks at once.
Jane: And the authors tested a variant where they didn’t train the transcription step, to see if that helped. It did help a bit, but CoT still degraded. Direct, on the other hand, kept its ASR performance stable while improving translation the whole way through.
Tom: So the summary is that Direct is more robust to scaling. It’s not the best at small data sizes, but it’s the one that benefits from more data. And that’s a really important insight for anyone building these systems.
Jane: It really is. And it makes you wonder, what’s the ceiling for Direct? We’ll talk about what this means for the field next.
Improvements: Tom: Alright, we’ve covered the summary. Now let’s talk about what this paper suggests we should do differently. And I want to bring in Lu and Meng for this one, because there are real engineering and research implications here.
Jane: Good idea. So the paper’s core suggestion is that we should reconsider Direct prompting as a primary strategy, especially as we scale up data. The authors point out that CoT forces the model to produce a transcription before translating, which constrains it to a fixed alignment between speech and text.
Lu: That’s the key insight for me, Jane. By forcing that intermediate transcription, you’re essentially making the model commit to a lossy representation. Speech carries prosody, emotion, emphasis, all sorts of paralinguistic information that doesn’t show up in a plain text transcript. CoT throws that away by design.
Tom: So you’re saying Direct could actually produce richer translations that capture things like sarcasm or urgency?
Lu: Exactly, Tom. The potential is there. But the paper is honest that this won’t happen automatically. Current training data is mostly just transcriptions and translations, so the model never learns to use those extra signals. We’d need datasets annotated specifically for speech-related cues.
Meng: And from an engineering standpoint, there’s a very practical benefit here too. Direct prompting requires roughly half the computational resources of CoT. You’re generating one sequence instead of two. That’s a massive saving in inference time and cost.
Jane: That’s a great point, Meng. And it’s not just about cost. It’s also about simplicity. Direct is easier to implement, easier to debug, and now we have evidence that it scales better. That’s a compelling combination.
Lu: I’d push back slightly on the simplicity point, Jane. The paper uses a two-stage training process with a frozen speech encoder and discrete tokens. That’s not trivial. But the prompting itself is simpler, yes.
Tom: Fair enough. So the improvements the paper suggests are really about shifting our focus. Instead of assuming CoT is always better, we should be building bigger and better direct S2TT datasets. And we should be thinking about what Direct can do that CoT fundamentally can’t.
Meng: And the numbers back that up. Direct at one hundred percent of the pseudo-labeled data nearly matches CoT’s peak for some language pairs, like Catalan to other languages. That’s with less compute and a simpler pipeline.
Jane: So the path forward is clear. More data, more direct training, and maybe a whole new class of translation quality that captures what speech actually sounds like.
Tom: I love that. Let’s wrap this up and look at the big picture.
Conclusion: Tom: Alright, we’ve reached the end of our discussion on “Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?” Jane, give us the final takeaway.
Jane: The takeaway is that Direct prompting, while starting behind CoT, scales much more consistently. The paper shows that as you add pseudo-labeled data, Direct keeps improving while CoT degrades, largely because the transcription step becomes a bottleneck.
Tom: And that’s a big deal because the field has been moving toward CoT as the default approach. This paper says, wait, let’s look at the data. Direct is cheaper, simpler, and has more headroom.
Lu: And I’ll add that the future is about exploiting what speech actually carries. Direct models aren’t forced through a text bottleneck, so they could produce translations that reflect tone, emphasis, and emotion. That’s a genuinely new capability.
Meng: From a practical side, the compute savings alone make this worth paying attention to. Half the resources for a model that scales better is a strong argument.
Lalam: If I may, there’s a cultural dimension here too. Translation that captures prosody could change how we localize media, preserve dialects, and convey the emotional weight of a speaker’s words across languages. That’s not just a technical improvement, it’s a richer form of human communication.
Jane: Beautifully put, Lalam. So we’re saying goodbye to this paper with a new question to carry forward. Can Direct prompting eventually surpass CoT entirely? The authors think it might, and the data suggests it’s worth trying.
Tom: And that’s where we leave it. Thanks to Lu, Meng, and Lalam for joining us. We’ll be back with the next paper soon. Until then, keep listening, keep learning, and keep questioning the default assumptions.
Jane: See you all next time.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language