Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition".
Jane: The paper was written by Yi-Cheng Lin, Yu-Hsuan Li Liang, Hsuan Su, Tzu-Quan Lin, Shang-Tse Chen et al. from National Taiwan University and NTU Artificial Intelligence Center of Research Excellence.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. I'm Tom, and sitting across from me is the wonderful Jane. Today we're cracking open a paper that's been making waves in the speech recognition world: "Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition."
Jane: Tom, I have to say, the title alone got me excited. "Pseudo2Real" — that's such a clever name. It basically says, "we're turning fake labels into real ones," and that's exactly what the paper does. But let's break it down for our listeners who might not be deep in the ASR trenches.
Tom: Right. So, automatic speech recognition — that's the tech that turns your voice into text. Think voice assistants, captioning, transcription services. The big problem is that these systems work great on standard English, but they fall apart when they hear accented speech. And the reason is simple: you need labeled data to train them, and getting human transcriptions for every accent on Earth is basically impossible.
Jane: And that's where pseudo-labels come in. You take a pre-trained model, let it listen to unlabeled audio, and it generates its own transcriptions. Those are your "pseudo-labels." They're free, but they're noisy. The model makes systematic mistakes — it might consistently mishear certain sounds in a particular accent.
Tom: Exactly. And the authors — Yi-Cheng Lin, Yu-Hsuan Li Liang, Hsuan Su, Tzu-Quan Lin, Shang-Tse Chen, Yun-Nung Chen, and Hung-yi Lee from National Taiwan University — they had a really clever idea. Instead of trying to filter out the bad pseudo-labels, they decided to learn the *difference* between a model trained on real labels and one trained on pseudo-labels.
Jane: That difference is like a "correction vector" in the model's parameter space. Think of it this way: you have two models that started from the same point, but one learned from truth and one learned from lies. The direction between them tells you exactly where the lies pushed the model. If you apply that direction to a new model trained on pseudo-labels in a different accent, you can push it back toward the truth.
Tom: And the results are wild. On the AfriSpeech-two hundred dataset, which covers ten African accents, they got up to a thirty-five percent relative reduction in word error rate using the Whisper Tiny model. That's not a small improvement — that's a game-changer for low-resource accents.
Jane: What I love about this approach is that it doesn't need any target-domain ground truth. You learn the correction from a source domain where you do have real labels, and then you apply it to a completely new accent. It's like learning a general "anti-bias" direction and just adding it to whatever model you're adapting.
Tom: And that's the hook, folks. We've got the title and the big idea. Next up, we're going to dig into the actual summary of the paper and see how they set up their experiments. Stay with us.
Summary: Jane: Welcome back. We're still on "Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition." Tom, we talked about the core idea — the correction vector. But what does the paper actually do step by step?
Tom: Great question, Jane. So the setup is a cross-fold validation. They took ten accents from AfriSpeech-two hundred and split them into two folds. Fold one has Igbo, Swahili, Hausa, Zulu, and Twi. Fold two has Yoruba, Ijaw, Afrikaans, Idoma, and Setswana. They train on one fold as the source domain and test on the other as the target, then swap.
Jane: And the key insight is that these accents aren't just different pronunciations — they come from completely different language families. We're talking Niger-Congo, Afro-Asiatic, Indo-European. So the transfer isn't trivial. The model has to learn something general about pseudo-label bias, not just accent-specific quirks.
Tom: Right. And they used Whisper models — Tiny, Base, Small, Medium, and Large v2 — ranging from thirty-nine million parameters all the way up to one point five five billion. They fully fine-tuned them on the source data with real labels and pseudo-labels, computed the difference, and then applied that correction to target models fine-tuned on pseudo-labels.
Jane: The baseline comparisons are important too. They compared against confidence filtering, which just throws away low-confidence pseudo-labels, and error correction using a T5 model trained to rewrite bad transcriptions. Neither of those worked nearly as well. The confidence filtering barely moved the needle, and the T5 error correction actually made things worse in some cases.
Tom: Yeah, that's fascinating. The T5 approach — where you train a model to fix the teacher's mistakes — it just doesn't generalize across accents. It learns source-specific error patterns and then tries to apply them where they don't belong. But Pseudo2Real works in parameter space, which apparently transfers much better.
Jane: And the numbers back that up. For Whisper Tiny, the pseudo-label baseline had an average WER of eighty-nine point three. Pseudo2Real brought that down to fifty-seven point seven. That's a thirty-five percent relative improvement. For Whisper Small, it went from forty-seven point two down to forty-five point zero.
Tom: But here's the thing that really got me — sometimes Pseudo2Real actually beat the topline, which is a model trained on the *real* target labels. On Igbo, Twi, and Idoma with Whisper Tiny, the corrected model outperformed the oracle. That's not supposed to happen.
Jane: I know! And the paper digs into why. They looked at the error breakdown and found that Pseudo2Real makes far fewer insertion errors — those are hallucinations where the model adds words that aren't there. On Twi, the topline had six hundred seventy insertions, but Pseudo2Real only had two hundred sixty-two. The correction vector seems to carry over some anti-hallucination regularization from the source domain.
Tom: So it's not just fixing the pseudo-label noise — it's actually transferring beneficial training signals that help the model generalize better than even ground-truth training on a small target set. That's a pretty profound result.
Jane: It really is. And it sets up the next question perfectly: how does this scale when you have different teacher and student model sizes? That's what we're going to dig into next.
Improvements: Tom: Welcome back to our deep dive on "Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition." Jane, we've seen the main results. But the paper doesn't stop there — they also tested what happens when the teacher and student are different sizes.
Jane: Right, and that's a really practical question. In the real world, you might have a big, expensive model generating pseudo-labels and a small, efficient model that you actually deploy. Does the correction vector still work across that gap?
Tom: The answer is mostly yes, but with some interesting caveats. When the student is Whisper Tiny, pairing it with a Base teacher gave a twenty-one point seven percent relative improvement, and a Large teacher gave twenty-one point six percent. But pairing it with a Medium teacher only gave four point four percent, and it actually hurt on Igbo — an eleven point five percent degradation.
Jane: That's counterintuitive, right? You'd think a bigger teacher would always be better. But the paper suggests that the correction vector from a Medium teacher might encode biases that don't match the Tiny student's capacity to absorb them. It's like trying to give a small car a big engine's tuning — it just doesn't fit.
Tom: And for the Small student, the gains were more modest overall. The Medium teacher gave a five point two percent average improvement, but the Large teacher was basically flat — zero point one percent on average. And on Hausa, it actually degraded by eighteen point eight percent. So the relationship between teacher size and correction effectiveness is not monotonic.
Jane: But then they introduced Pseudo2Real-SC, which is the subgroup clustering variant. Instead of one correction vector, they split the source speakers into clusters using ECAPA-TDNN embeddings and k-means. Each cluster gets its own correction vector, and then they average them all together.
Tom: And that's where things get interesting. For the Large teacher paired with the Small student, the single vector was basically useless — zero point one percent improvement. But with eight clusters, Pseudo2Real-SC got a six point one percent average improvement. On Hausa, it jumped from sixty-five point three WER down to fifty-three point nine — a seventeen point five percent relative gain.
Jane: Why does clustering help so much? Because pseudo-label errors aren't uniform across speakers. Some speakers might have a consistent substitution pattern, while others have a different one. If you average them all together, you get a muddy correction that doesn't help anyone. But if you compute per-cluster corrections, you capture those fine-grained patterns.
Tom: And they showed that more clusters is generally better — they tested one two four and eight clusters, and WER kept dropping as the cluster count went up. But there's a cost: each cluster requires training two full ASR models, one on real labels and one on pseudo-labels. So it's a compute trade-off.
Jane: And they also looked at the scaling factor lambda, which controls how strongly you apply the correction. They found a sweet spot around zero point two to zero point three. Too small and you don't correct enough; too large and you over-correct and break the model. That's a critical practical detail.
Tom: It really is. And the case studies in the paper are beautiful — they show the teacher producing nonsense like "as a vif" for "survived," and the student inheriting that error. But Pseudo2Real restores the correct word. It's not just lowering a metric; it's fixing actual meaning.
Jane: And that's the kind of improvement that matters for real users. We're talking about making speech tech work for people who've been left out because of their accent. That's a big deal.
Tom: It is. And now we're at the conclusion, where we wrap this all up.
Conclusion: Jane: So here we are at the end of our discussion on "Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition." Tom, give us the final take — what did we learn?
Tom: We learned that you can correct systematic pseudo-label errors without ever seeing target-domain ground truth. You learn a correction vector in parameter space from a source domain where you have real labels, and you apply it to a new domain. It's simple, it's effective, and it works across model sizes.
Jane: And the numbers speak for themselves — up to thirty-five percent relative WER reduction on Whisper Tiny, and sometimes beating the topline that was trained on real target labels. That's not incremental progress; that's a paradigm shift for unsupervised domain adaptation.
Tom: The Pseudo2Real-SC extension with speaker clustering adds even more robustness, especially when you're dealing with a high-capacity teacher. And the ablation on the scaling factor gives practitioners a clear recipe for tuning.
Jane: I also appreciate that they were honest about limitations. The method needs a good source domain, the pseudo-label biases need to be somewhat consistent across domains, and the scaling factor has to be tuned. It's not a magic bullet, but it's a powerful tool.
Tom: And the implications go beyond just African accents. This could apply to any domain shift — different microphones, different recording conditions, different languages. The principle is general: if you can model the difference between real and synthetic supervision, you can correct for it.
Jane: We should also mention the broader impact. Making ASR work for more accents means making voice technology accessible to more people — healthcare, education, government services. That's not just a technical win; it's a fairness win.
Tom: Absolutely. And with that, we're going to say goodbye to "Pseudo2Real" and get ready for our next paper. Thanks for listening, everyone. We'll see you on the next episode.
Jane: Take care, and keep listening.
Yi-Cheng Lin, Yu-Hsuan Li Liang, Hsuan Su, Tzu-Quan Lin, Shang-Tse Chen, Yun-Nung Chen, Hung-yi Lee
National Taiwan University · NTU Artificial Intelligence Center of Research Excellence
eess.AS, cs.CL
Submitted: 2026-04-20
Updated: 2026-08-18
Comments: Accepted to ACL 2026 Findings
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 82/100
The gist: one on ground-truth labels (θ s real) and one on pseudo-labels (θ s pseudo).
Key concepts
- Pseudo-labels
- These are transcriptions generated by a pre-trained model listening to unlabeled audio. They are free but contain systematic mistakes because the model makes consistent errors when hearing specific accents or speech patterns.
- Correction Vector
- This represents the difference between a model trained on real labels and one trained on pseudo-labels, existing in the model's parameter space. Applying this vector pushes a pseudo-labeled model back toward the truth.
- Pseudo2Real-SC
- This extension uses speaker clustering to create separate correction vectors for different speaker groups. This helps capture fine-grained error patterns, leading to better performance than using a single correction vector.
Terminology
Summary
Summary
The paper introduces Pseudo2Real, a parameter-space correction method designed to mitigate systematic pseudo-label errors in Automatic Speech Recognition (ASR) domain adaptation without requiring target-domain ground truth. The authors address the problem that pseudo-labels inherit the teacher model’s systematic biases, such as under-recognizing rare words, accent-driven substitutions, or domain-specific mis-segmentations,
and that when used for adaptation, these errors can accumulate and degrade real-world performance.
The core idea is to construct a correction vector in parameter space. In a source domain containing both real and pseudo-labeled data, two ASR models are fine-tuned from the same pre-trained initialization: one on ground-truth labels (θ s real) and one on pseudo-labels (θ s pseudo). Their weight difference defines the correction vector: τ = θ s real − θ s pseudo. This vector captures systematic discrepancies introduced by pseudo-labeling.
When adapting to a new target domain, a student model is fine-tuned on pseudo-labeled target audio (θ t pseudo), and the correction vector is added with a scaling factor λ: θ t corrected = θ t pseudo + λτ.
The authors also propose an extension, Pseudo2Real-SC, which partitions the source domain into speaker subgroups using ECAPA-TDNN embeddings and k-means clustering. For each cluster c, a subgroup-specific correction vector is computed (τ c = θ s,c real − θ s,c pseudo), and the final correction is the average across all C clusters: θ t corrected = θ t pseudo + (λ/C) Σ c=1 C τ c.
Experiments are conducted on the AfriSpeech-200 dataset, using ten African accents (Igbo, Swahili, Hausa, Zulu, Twi, Yoruba, Ijaw, Afrikaans, Idoma, Setswana) in a cross-fold validation setting. The authors use Whisper models (TINY, BASE, SMALL, MEDIUM, LARGE v2) and fully fine-tune them. The scaling factor λ is selected via grid search on the source-domain development set.
Key results show that Pseudo2Real achieves up to a 35% relative Word Error Rate (WER) reduction on AfriSpeech-200 across ten African accents with the Whisper TINY model.
For Whisper TINY, the average WER drops from 89.3 (pseudo-label fine-tuning) to 57.7, a 35% relative improvement. For Whisper SMALL, the average WER decreases from 47.2 to 45.0. The method occasionally outperforms the topline trained on labeled target data,
which the authors attribute to the correction vector transferring additional anti-hallucination regularization from the source domain.
The paper also analyzes the scaling factor λ, finding a U-shaped trend where small values (λ = 0.2–0.3) yield the best performance, while excessive scaling causes over-correction. For Pseudo2Real-SC, results show that averaging subgroup-specific correction vectors yields additional improvements, particularly with high-capacity teachers (e.g., +6.1% relative improvement for LARGE→SMALL). An ablation on the number of clusters shows that increasing clusters from 1 to 8 reduces WER from 45.36 to 41.94 in the LARGE→SMALL setting.
The authors conclude that Pseudo2Real achieves substantial improvements across all accents, demonstrating the effectiveness of parameter-space correction
and that the correction vector effectively captures accent-specific pronunciation patterns and mitigates systematic pseudo-labeling errors that standard fine-tuning fails to address.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:
1. Add a Parameter-Space Correction Module (Pseudo2Real)
-
Implementation: After fine-tuning a student ASR model on pseudo-labels for a target domain, compute a correction vector
τ = θ source real − θ source pseudo(from a source domain where both real and pseudo-labels exist). Addλτto the target model’s weights, withλtuned on source dev data (grid search over 0.1–1.0). -
Benefit: Directly mitigates systematic pseudo-label biases (accent-specific substitutions, hallucinations) without needing target ground truth.
2. Add a Subgroup Correction Variant (Pseudo2Real-SC)
-
Implementation: Extract speaker embeddings (ECAPA-TDNN), cluster source utterances via k-means (k=4–8), train a real/pseudo model pair per cluster, average the resulting correction vectors, and apply the average to the target model.
-
Benefit: Captures heterogeneous pseudo-label errors across speaker subgroups, improving robustness when teacher capacity is high (e.g., LARGE→SMALL gains +6.1% relative WER).
3. Integrate a Scaling-Factor Scheduler
-
Implementation: Automatically select
λper target domain by evaluating WER on a held-out source dev set across a small grid (0.1–0.5). Optionally, use a U-shaped curve (optimal 0.2–0.3) to avoid over-correction. -
Benefit: Prevents performance degradation from excessive correction strength (e.g., λ>0.4 caused WER spikes in experiments).
4. Add an Error-Type-Aware Correction Filter
-
Implementation: During inference, compare the corrected model’s output against the pseudo-label model’s output. If the correction reduces insertion errors (a common hallucination pattern) while keeping substitutions/deletions stable, accept the correction; otherwise, revert to the uncorrected output.
-
Benefit: Reduces hallucinated tokens (e.g., “full stop”) and accent-induced confusions (e.g., “said” vs. “side”) more reliably, as shown in case studies.
5. Extend to Multi-Teacher Ensemble Correction
-
Implementation: When multiple teacher sizes are available (TINY, BASE, SMALL, MEDIUM, LARGE), compute correction vectors for each teacher–student pairing and average them (weighted by source-dev WER). Apply the averaged vector to the target student.
-
Benefit: Achieves consistent gains across model scales (e.g., +21.7% relative WER for TINY with BASE teacher, +5.2% for SMALL with MEDIUM teacher) while reducing variance from a single teacher’s biases.
-
Adapt to new accents with no labeled target data: Given unlabeled accented speech (e.g., Ijaw, Hausa), the system fine-tunes on pseudo-labels, then applies the correction vector to reduce WER by up to 35% relative (e.g., from 89.3% to 57.7% on Whisper TINY).
-
Correct systematic pseudo-label errors without target ground truth: It automatically fixes recurring teacher mistakes (e.g., mishearing “survives” as “as a vif”) by shifting parameters along a learned pseudo-to-real axis, restoring semantically correct transcriptions.
-
Handle heterogeneous speaker groups: By clustering speakers and computing subgroup-specific corrections, it improves performance on accents where teacher errors vary (e.g., +21.8% relative improvement on Hausa with MEDIUM teacher).
-
Avoid over-correction: The scaling-factor scheduler ensures the correction is applied at an optimal strength, preventing WER degradation from excessive parameter shifts.
-
Generalize across model sizes: The system works for Whisper TINY through LARGE, with gains observed for both low-capacity (TINY) and high-capacity (SMALL) students, and across different teacher–student pairings.
-
Reduce hallucinated tokens: The error-type-aware filter suppresses spurious insertions (e.g., “full stop”) that pseudo-label fine-tuning tends to propagate, improving transcription fidelity.
-
Operate in a fully unsupervised target-domain setting: All correction vectors are derived from source-domain supervision only; no target-domain labels are used, making it suitable for privacy-constrained or low-resource deployments.
These improvements are directly implementable in existing ASR pipelines (e.g., Whisper fine-tuning) with minimal architectural changes, requiring only an additional source-domain labeled dataset and a small grid-search for λ.
Abstract
Robust ASR under domain shift is crucial because real-world systems encounter unseen accents and domains with limited labeled data. Although pseudo-labeling offers a practical workaround, it often introduces systematic, accent-specific errors that filtering fails to fix. We ask: How can we correct these recurring biases without target ground truth? We propose a simple parameter-space correction: in a source domain containing both real and pseudo-labeled data, two ASR models are fine-tuned from the same initialization, one on ground-truth labels and the other on pseudo-labels, and their weight difference forms a correction vector that captures pseudo-label biases. When applied to a pseudo-labeled target model, this vector enhances recognition, achieving up to a 35% relative Word Error Rate (WER) reduction on AfriSpeech-200 across ten African accents with the Whisper tiny model.
Sources
- Privacy in Speech Technology
- Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
- Better Pseudo-labeling with Multi-ASR Fusion and Error Correction by SpeechLLM
- A correlation-permutation approach for speech-music encoders model merging
- Sotto Voce: Federated Speech Recognition with Differential Privacy Guarantees
- Building a Taiwanese Mandarin Spoken Language Model: A First Attempt
- Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping
- Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions