REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing
summary
The gist
This paper introduces REDDIT (REplay-based Distribution eDITing), a lightweight post-training framework designed to correct "non-speech-induced timestamp drift" in autoregressive Automatic Speech
In short
The episode discusses the REDDIT paper, which addresses timestamp drift in Automatic Speech Recognition models. Using a two-stage replay-based approach, researchers correct timing errors without degrading speech recognition accuracy. This method utilizes synthetic data and minimal parameter updates to significantly improve temporal precision while avoiding model "forgetting."
Key concepts
- Timestamp Drift
- The phenomenon where an ASR model correctly identifies spoken words but loses track of when they occurred. This often happens during silences, causing a disconnect where the text is accurate but the timing is out of sync with the audio.
- Forgetting
- A problem in fine-tuning where teaching a model a new skill causes it to lose its original capabilities, such as speech recognition accuracy. REDDIT avoids this by surgically editing temporal markers without disrupting the model's language understanding.
- Replay-Based Distribution Editing
- A two-stage method that uses the model's original output to maintain context and stability. It uses KL divergence to keep words unchanged while shifting timestamps, followed by a refinement stage using synthetic data—clean speech with known silent gaps—to ensure precise timing.
Terminology used across episodes
This episode discusses
- REDDIT: Forgetting-Resistant Correction of Timestamp Drift in ASR via Replay-Based Distribution Editing · Paper Radio
- Robust Speech Recognition via Large-Scale Weak Supervision
- Common Voice: A Massively-Multilingual Speech Corpus
- Word Level Timestamp Generation for Automatic Speech Recognition and Translation
- In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
- Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models
- Distilling the Knowledge in a Neural Network
- Tradition or Innovation: A Comparison of Modern ASR Methods for Forced Alignment
- CTC-Segmentation of Large Corpora for German End-to-end Speech Recognition
- LLM-ForcedAligner: A Non-Autoregressive and Accurate LLM-Based Forced Aligner for Multilingual and Long-Form Speech
- CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions
- Whisper Has an Internal Word Aligner
- rVAD: An Unsupervised Segment-Based Robust Voice Activity Detection Method
- Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down
- Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models
- From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection
- Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders
- A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
- Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models
- AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
- AudioBench: A Universal Benchmark for Audio Large Language Models
The paper
REDDIT: Forgetting-Resistant Correction of Timestamp Drift in ASR via Replay-Based Distribution Editing · Read on arXiv
National Taiwan University · Carnegie Mellon University · National Taiwan University Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing".
Jane: The paper was written by Cheng-Kang Chou, Ming-Douo Tchouang, Ke-Han Lu, Chan-Jan Hsu and Hung-yi Lee from National Taiwan University and Carnegie Mellon University and National Taiwan University Artificial Intelligence Center of Research Excellence (NTU AI-CoRE).
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome to the show, where we break down the latest breakthroughs in research, and today we're looking at a paper titled REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing.
Jane: That is quite a mouthful, Tom, but the problem they're solving is something we all encounter when watching videos.
Tom: You mean when the subtitles don't quite match the person speaking?
Jane: Exactly, it's that annoying delay or jump where the words are right but the timing is totally off.
Lu: It's actually a much deeper problem than just a glitchy subtitle, because it shows a fundamental disconnect in how these models perceive time.
Tom: So the authors, a team from NTU and CMU, are saying the model knows *what* was said but loses track of *when* it happened?
Jane: That's a great way to put it, especially since these models generate the time as part of their own text output.
Lu: I find the concept of "drift" so fascinating because it implies the model's internal clock is essentially drifting away from reality during silence.
Meng: From a practical standpoint, if I'm building a real-time transcription tool, that drift makes the whole system feel broken or unreliable.
Tom: And that's why this paper is so important, because they aren't just fixing the timing, they're doing it without breaking the actual speech recognition.
Jane: They call that "forgetting," which happens when you try to teach a model a new trick and it suddenly forgets how to do its old ones.
Meng: I've seen that happen in so many fine-tuning projects where the accuracy goes through the floor.
Lu: Imagine if we could surgically edit just the temporal part of the brain without touching the language part!
Lalam: That's the beauty of this research, as it moves us toward AI that can truly synchronize with the human experience of rhythm and pause.
Jane: It’s like giving a musician a metronome that only kicks in when they lose the beat.
Tom: Let's get into how they actually pull off this "surgical" edit without the model losing its mind.
Summary: Tom: We're moving into the meat of the paper now, looking at how REDDIT actually manages to fix that timing drift.
Jane: They use this two-stage approach that sounds almost like a rehearsal process for an actor.
Tom: An actor? How so?
Jane: Well, in the first stage, they let the model "replay" its own original, drifted output to keep the context stable.
Lu: It's a brilliant way to use the model's own existing knowledge as a foundation for the correction.
Meng: But how do they make sure the model doesn't just learn to repeat its own mistakes during that replay?
Lu: They use something called KL divergence, which basically acts like a tether to the original model's behavior for everything except the timestamps.
Jane: So, they're telling the model, "Keep your words exactly the same, but move these specific time markers to these new spots."
Tom: And then they have this second stage, right?
Jane: Yes, they take that corrected version and run a quick refinement stage using those new, correct timestamps as a starting point.
Meng: I'm curious about the data they used to train this, because you'd think you'd need humans to label all those timestamps.
Tom: That's the clever part—they didn't use any human annotations at all!
Jane: They just took clean speech, chopped it up, and inserted silent gaps where they already knew exactly how long the silence was.
Lu: It's a synthetic playground that allows for perfect, mathematical ground truth without any human error.
Meng: That sounds like a massive win for scaling, since you can generate infinite training examples this way.
Lalam: It allows the machine to learn the nuances of silence and speech transition in a way that feels natural to human culture.
Tom: It really is a clever way to build a teacher without needing a classroom of humans.
Improvements: Tom: The numbers in this paper are absolutely wild, especially when you look at the Whisper-tiny experiments.
Jane: I was just looking at that, and the jump in the long-gap mIoU from thirty-eight point seven percent to ninety-five point zero percent is just massive.
Tom: And the timing error, the AAS, dropped from two thousand seven hundred fifty-two milliseconds all the way down to two hundred twenty-three milliseconds!
Jane: That's the difference between a subtitle that's a huge distraction and one that's almost invisible.
Lu: What's even more impressive is that they did all of this by updating only one point six percent of the model's parameters.
Meng: Only one point six percent? That's incredibly efficient for an engineering team trying to optimize compute costs.
Tom: And they didn't fall into the "forgetting" trap that the other methods did.
Jane: Right, the standard fine-tuning actually caused the error rate to skyrocket to over five hundred percent in some cases!
Meng: That's a total disaster for any production environment.
Lu: It shows that the "distribution editing" approach is much more stable than just throwing more data at the problem.
Tom: They even showed it works on the larger Whisper-large-v3 model, though the scaling is a bit different there.
Jane: It seems like the two-stage process becomes even more critical as the models get bigger and more complex.
Lu: It's a scalable strategy that could eventually be applied to much larger audio-language models.
Lalam: This kind of precision will allow AI to participate in real-time translation and accessibility in ways that feel seamless and human.
Meng: If we can get this level of accuracy with such low overhead, it's a game changer for edge devices.
Tom: It really is, and it's hard to imagine going back to those drifting subtitles now.
Conclusion: Tom: We've covered a lot of ground today, from the frustration of drifting timestamps to this incredibly elegant solution.
Jane: It's one of those papers where the solution feels so much more "correct" than the brute-force methods we usually see.
Tom: They've really shown that you can edit a model's behavior without destroying its core intelligence.
Lu: This opens up so many doors for how we think about model adaptation and specialized task editing in the future.
Meng: I'm definitely looking at how we can implement this kind of parameter-efficient tuning in our own pipelines.
Lalam: Ultimately, this is about making technology more harmonious with the way we actually live and communicate.
Tom: That's a perfect place to wrap this up.
Jane: We'll be back with more research very soon, but thanks for joining us.
Tom: Goodbye, everyone!
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language