Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?".
Jane: The paper was written by Oriol Pareras, Gerard I. Gállego, Federico Costa, Cristina España-Bonet and Javier Hernando from Barcelona Supercomputing Center and Universitat Politècnica de Catalunya and DFKI GmbH and Saarland Informatics Campus.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everybody. Today we’re digging into a paper that’s been making the rounds in the speech translation world, and it’s called “Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?” Jane, I have to say, that title alone got me excited.
Jane: Oh, absolutely, Tom. And for our listeners who might not be deep in the weeds here, let’s break down what that title actually means. Speech-to-text translation is when you hear someone speaking in, say, Spanish, and the system outputs English text. The big question in the field right now is how to build those systems.
Tom: Right, and the paper is comparing two ways of doing it. One is called Direct prompting, where the model just hears the audio and immediately translates it. The other is Chain-of-Thought, or CoT, where the model first transcribes the speech into text, and then translates that text. It’s like doing two steps instead of one.
Jane: And that CoT approach has been really popular lately because it lets you use tons of existing data. You have speech-to-text data and text-to-text translation data, and you can kind of stack them together. So the assumption has been that CoT is just better. But this paper from Barcelona Supercomputing Center and UPC is questioning that.
Tom: Questioning it with actual experiments, which is the best part. They built these speech LLMs, trained them with both approaches, and then scaled up the amount of translation data they fed them. And what they found is that Direct prompting gets better and better as you add more data, while CoT kind of plateaus and even degrades.
Jane: It’s a really important finding because the field has been moving toward CoT as the default. The authors are basically saying, hold on, maybe we’ve been optimizing for the wrong thing. If we keep building bigger datasets, Direct might actually win.
Tom: And that’s the hook for me, Jane. The title asks a question, and the answer is a cautious yes. Direct prompting scales better. We’re going to get into the nitty-gritty of how they tested that in a moment.
Jane: Stay with us, because this has big implications for how we build speech translation systems in the future.
Summary: Tom: So we’ve set the stage. Let’s get into the actual summary of “Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?” Jane, walk us through what they actually did.
Jane: Gladly. So they took a multilingual LLM called Salamandra, which is fine-tuned for translation, and they hooked it up to a speech encoder called mHuBERT. That encoder turns audio into discrete tokens that the LLM can understand. Then they trained it on a bunch of different data types.
Tom: And here’s the clever part. They didn’t have enough real speech-to-text translation data, so they made their own. They took a huge ASR corpus, Common Voice, which has audio with transcriptions, and they translated those transcriptions into six European languages using the same LLM. That’s the pseudo-labeling pipeline.
Jane: Exactly. They filtered the results using quality estimation and language identification to make sure the translations were actually good and in the right language. Then they trained models with both Direct and CoT prompting, adding increasing amounts of this pseudo-labeled data. They went from twenty percent up to one hundred percent.
Tom: And the results, Jane? I was honestly surprised. The CoT baseline starts off better, no question. At zero added data, CoT beats Direct by about five BLEU points and seven xCOMET points. But then you start adding pseudo-labeled data, and Direct just keeps climbing.
Jane: While CoT peaks at twenty percent of the data and then starts falling. The paper shows that CoT’s transcription quality actually gets worse as you add more translation data. The WER, that’s the word error rate for transcription, goes up by as much as thirteen percent relative to baseline.
Tom: That’s a huge deal. It means the intermediate transcription step, which is supposed to be the strength of CoT, is actually becoming a bottleneck. The model is getting confused trying to do both tasks at once.
Jane: And the authors tested a variant where they didn’t train the transcription step, to see if that helped. It did help a bit, but CoT still degraded. Direct, on the other hand, kept its ASR performance stable while improving translation the whole way through.
Tom: So the summary is that Direct is more robust to scaling. It’s not the best at small data sizes, but it’s the one that benefits from more data. And that’s a really important insight for anyone building these systems.
Jane: It really is. And it makes you wonder, what’s the ceiling for Direct? We’ll talk about what this means for the field next.
Improvements: Tom: Alright, we’ve covered the summary. Now let’s talk about what this paper suggests we should do differently. And I want to bring in Lu and Meng for this one, because there are real engineering and research implications here.
Jane: Good idea. So the paper’s core suggestion is that we should reconsider Direct prompting as a primary strategy, especially as we scale up data. The authors point out that CoT forces the model to produce a transcription before translating, which constrains it to a fixed alignment between speech and text.
Lu: That’s the key insight for me, Jane. By forcing that intermediate transcription, you’re essentially making the model commit to a lossy representation. Speech carries prosody, emotion, emphasis, all sorts of paralinguistic information that doesn’t show up in a plain text transcript. CoT throws that away by design.
Tom: So you’re saying Direct could actually produce richer translations that capture things like sarcasm or urgency?
Lu: Exactly, Tom. The potential is there. But the paper is honest that this won’t happen automatically. Current training data is mostly just transcriptions and translations, so the model never learns to use those extra signals. We’d need datasets annotated specifically for speech-related cues.
Meng: And from an engineering standpoint, there’s a very practical benefit here too. Direct prompting requires roughly half the computational resources of CoT. You’re generating one sequence instead of two. That’s a massive saving in inference time and cost.
Jane: That’s a great point, Meng. And it’s not just about cost. It’s also about simplicity. Direct is easier to implement, easier to debug, and now we have evidence that it scales better. That’s a compelling combination.
Lu: I’d push back slightly on the simplicity point, Jane. The paper uses a two-stage training process with a frozen speech encoder and discrete tokens. That’s not trivial. But the prompting itself is simpler, yes.
Tom: Fair enough. So the improvements the paper suggests are really about shifting our focus. Instead of assuming CoT is always better, we should be building bigger and better direct S2TT datasets. And we should be thinking about what Direct can do that CoT fundamentally can’t.
Meng: And the numbers back that up. Direct at one hundred percent of the pseudo-labeled data nearly matches CoT’s peak for some language pairs, like Catalan to other languages. That’s with less compute and a simpler pipeline.
Jane: So the path forward is clear. More data, more direct training, and maybe a whole new class of translation quality that captures what speech actually sounds like.
Tom: I love that. Let’s wrap this up and look at the big picture.
Conclusion: Tom: Alright, we’ve reached the end of our discussion on “Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?” Jane, give us the final takeaway.
Jane: The takeaway is that Direct prompting, while starting behind CoT, scales much more consistently. The paper shows that as you add pseudo-labeled data, Direct keeps improving while CoT degrades, largely because the transcription step becomes a bottleneck.
Tom: And that’s a big deal because the field has been moving toward CoT as the default approach. This paper says, wait, let’s look at the data. Direct is cheaper, simpler, and has more headroom.
Lu: And I’ll add that the future is about exploiting what speech actually carries. Direct models aren’t forced through a text bottleneck, so they could produce translations that reflect tone, emphasis, and emotion. That’s a genuinely new capability.
Meng: From a practical side, the compute savings alone make this worth paying attention to. Half the resources for a model that scales better is a strong argument.
Lalam: If I may, there’s a cultural dimension here too. Translation that captures prosody could change how we localize media, preserve dialects, and convey the emotional weight of a speaker’s words across languages. That’s not just a technical improvement, it’s a richer form of human communication.
Jane: Beautifully put, Lalam. So we’re saying goodbye to this paper with a new question to carry forward. Can Direct prompting eventually surpass CoT entirely? The authors think it might, and the data suggests it’s worth trying.
Tom: And that’s where we leave it. Thanks to Lu, Meng, and Lalam for joining us. We’ll be back with the next paper soon. Until then, keep listening, keep learning, and keep questioning the default assumptions.
Jane: See you all next time.
Oriol Pareras, Gerard I. Gállego, Federico Costa, Cristina España-Bonet, Javier Hernando
Barcelona Supercomputing Center · Universitat Politècnica de Catalunya · DFKI GmbH · Saarland Informatics Campus
cs.CL, cs.SD
Submitted: 2026-01-23
Updated: 2026-08-18
Comments: To appear in Proc. ICASSP 2026, May 04-08, 2026, Barcelona, Spain
DOI: 10.1109/ICASSP55912.2026.11464161
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 81/100
The gist: The paper investigates the scaling behavior of two prompting strategies for End-to-End (E2E) Speech-to-Text Translation (S2TT) using LLM-based models: Chain-of-Thought (CoT) prompting, where the
Key concepts
- Speech-to-Text Translation
- This process involves taking spoken audio in one language (e.g., Spanish) and having a system output the translated text in another language (e.g., English).
- Direct Prompting
- A method where a model hears the audio and immediately translates it into text, performing the translation in a single step without an intermediate transcription.
- Chain-of-Thought (CoT)
- An approach where the model first transcribes speech into text, and then uses that resulting text to perform the actual translation. It is a two-step process.
- Pseudo-labeling
- A technique used by the hosts to create training data. They took an existing audio corpus (Common Voice) and translated its transcriptions into multiple languages using an LLM.
Terminology
Summary
The paper investigates the scaling behavior of two prompting strategies for End-to-End (E2E) Speech-to-Text Translation (S2TT) using LLM-based models: Chain-of-Thought (CoT) prompting, where the model first transcribes the speech and then translates it, versus Direct prompting, where the model translates directly without intermediate steps. The authors note that CoT typically outperforms direct prompting primarily because it can exploit abundant Automatic Speech Recognition (ASR) and Text-to-Text Translation (T2TT) datasets to explicitly model its steps.
However, they argue that the advantage of CoT is mainly tied to the large quantity of ASR and T2TT data currently used for training
and hypothesize that as S2TT data becomes increasingly available, direct translation may emerge as a stronger alternative.
To test this, the authors construct a pseudo-labeled S2TT dataset by translating the transcriptions of an ASR corpus (Common Voice 21.0) into six European languages (Catalan, German, English, Spanish, French, Italian) using a translation pipeline. This pipeline uses the same translation LLM that serves as the backbone for the S2TT model, followed by filtering with BLASER 2.0 for quality estimation (samples below 3.75 are removed) and GlotLID v3 for language identification (samples with expected language probability below 0.5 are discarded). The resulting S2TTpl dataset contains approximately 384M target tokens, with per-language hours ranging from 1,230 (Italian) to 8,875 (Catalan).
The S2TT model architecture uses mHuBERT from TWIST as the speech encoder, quantized into 500 discrete speech tokens via k-means clustering on 11th-layer representations at 25 Hz, and salamandraTA-7B-Instruct as the backbone LLM. Training occurs in two stages: stage 1 freezes the LLM and trains only the token embedding layer on ASR data; stage 2 trains the entire LLM on all data types. The authors train two baselines (DIRECT BASE and COT BASE) using all ASR, T2TT, and S2TT data, then incrementally add 20%, 40%, 60%, 80%, and 100% of the pseudo-labeled data to create AUG variants. For CoT, they also train a variant (COT† AUG) where the transcription step does not contribute to the loss, hypothesizing that each speech–transcription pair in the pseudo-labeled data is repeated six times, one for each target language
which may harm the ASR performance.
Results show that the COT BASE outperforms the DIRECT BASE across all languages, confirming the benefit of decomposing the task into transcription and translation in this scenario,
with an average gap of approximately 5 BLEU points and 7 xCOMET points. However, when scaling S2TTpl data, the two COT AUG variants peak at 20% of the S2TTpl data and then degrade as more data is added,
whereas DIRECT AUG consistently improves, achieving a better score with DIRECT AUG 100.
Although the peaks of CoT variants remain higher, DIRECT AUG demonstrates a steadier scaling trend, suggesting that this gap could be closed with additional data.
The authors also analyze sub-task performance. For ASR, WER remains stable with DIRECT AUG across all S2TTpl data scales, whereas both COT AUG and COT† AUG gradually increase it by up to +7.4% and +13% relative to their baseline,
correlating with the degradation in S2TT performance. For T2TT, DIRECT performs slightly worse... likely due to the lower proportion of this task in the training recipe,
but performance remains stable for all methods. The COT† AUG variant consistently underperforms COT AUG in ASR and S2TT,
supporting the hypothesis that training the transcription step may negatively impact the model.
Language-specific analysis for English, Catalan, and Italian confirms the overall trends. For English, performance on en→x remains nearly constant,
likely due to the large amount of English ASR data. For Catalan, DIRECT AUG 100 nearly matches the peak of COT AUG 20,
supporting the hypothesis that with sufficient data, Direct can reach CoT performance.
The authors conclude that COT struggles to improve with more data, regardless of whether the transcription step is trained or not,
while "DIRECT prompting demonstrates a consistent scaling trend: although its absolute performance has not yet surpassed CoT, it steadily improves as more pseudo-labeled data is added, indicating strong potential for larger scale trainings. They position Direct as a promising strategy due to its favorable scaling, requiring
roughly half the computational resources of CoT and being
simpler to implement. They also suggest that
by enforcing an intermediate transcription, CoT may constrain the model to a fixed alignment between speech and text, which may limit the potential of E2E models of leveraging paralinguistic information, and propose future work on
paralinguistic-aware translation where Direct prompting might better exploit
the full richness of the speech signal compared to CoT."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems, along with what the improved system can do:
Implementation:
Build a meta-controller that estimates the volume of available S2TT training data and automatically selects between Chain-of-Thought (CoT) and Direct prompting. The controller uses a threshold—when S2TT data exceeds 20% of the total training corpus (or when pseudo-labeled data reaches a scale where Direct shows consistent gains), it switches to Direct prompting.
What the improved system can do:
-
Automatically adapts its inference strategy to maximize translation quality without manual tuning.
-
Avoids the performance degradation seen in CoT when S2TT data scales (e.g., CoT dropped by up to 13% relative WER at 100% data, while Direct improved steadily).
-
Reduces computational cost by up to 50% (Direct uses half the tokens of CoT) when data is abundant.
The improved AI system can:
-
Automatically choose the optimal prompting strategy based on training data scale and audio quality, maximizing translation accuracy while minimizing computational cost.
-
Scale efficiently with pseudo-labeled data without the performance degradation seen in current CoT models, making it suitable for large-scale multilingual deployments.
-
Maintain robust ASR and T2TT performance even when trained primarily on S2TT data, ensuring versatility across tasks.
-
Handle low-resource languages effectively by combining CoT’s strengths at small data scales with Direct’s scalability at larger scales.
-
Reduce training and inference costs by up to 50% when Direct prompting is viable, without sacrificing quality.
Sources
- AudioPaLM: A Large Language Model That Can Speak and Listen
- Speech Translation with Large Language Models: An Industrial Practice
- Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning
- Deep Speech: Scaling up end-to-end speech recognition
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering