REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing

summary

Video file (mp4)

The gist

This paper introduces REDDIT (REplay-based Distribution eDITing), a lightweight post-training framework designed to correct "non-speech-induced timestamp drift" in autoregressive Automatic Speech

In short

The episode discusses the REDDIT paper, which addresses timestamp drift in Automatic Speech Recognition models. Using a two-stage replay-based approach, researchers correct timing errors without degrading speech recognition accuracy. This method utilizes synthetic data and minimal parameter updates to significantly improve temporal precision while avoiding model "forgetting."

Key concepts

Timestamp Drift
The phenomenon where an ASR model correctly identifies spoken words but loses track of when they occurred. This often happens during silences, causing a disconnect where the text is accurate but the timing is out of sync with the audio.
Forgetting
A problem in fine-tuning where teaching a model a new skill causes it to lose its original capabilities, such as speech recognition accuracy. REDDIT avoids this by surgically editing temporal markers without disrupting the model's language understanding.
Replay-Based Distribution Editing
A two-stage method that uses the model's original output to maintain context and stability. It uses KL divergence to keep words unchanged while shifting timestamps, followed by a refinement stage using synthetic data—clean speech with known silent gaps—to ensure precise timing.

Terminology used across episodes

This episode discusses

The paper

REDDIT: Forgetting-Resistant Correction of Timestamp Drift in ASR via Replay-Based Distribution Editing · Read on arXiv

National Taiwan University · Carnegie Mellon University · National Taiwan University Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing".

Jane: The paper was written by Cheng-Kang Chou, Ming-Douo Tchouang, Ke-Han Lu, Chan-Jan Hsu and Hung-yi Lee from National Taiwan University and Carnegie Mellon University and National Taiwan University Artificial Intelligence Center of Research Excellence (NTU AI-CoRE).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome to the show, where we break down the latest breakthroughs in research, and today we're looking at a paper titled REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing.

Jane: That is quite a mouthful, Tom, but the problem they're solving is something we all encounter when watching videos.

Tom: You mean when the subtitles don't quite match the person speaking?

Jane: Exactly, it's that annoying delay or jump where the words are right but the timing is totally off.

Lu: It's actually a much deeper problem than just a glitchy subtitle, because it shows a fundamental disconnect in how these models perceive time.

Tom: So the authors, a team from NTU and CMU, are saying the model knows *what* was said but loses track of *when* it happened?

Jane: That's a great way to put it, especially since these models generate the time as part of their own text output.

Lu: I find the concept of "drift" so fascinating because it implies the model's internal clock is essentially drifting away from reality during silence.

Meng: From a practical standpoint, if I'm building a real-time transcription tool, that drift makes the whole system feel broken or unreliable.

Tom: And that's why this paper is so important, because they aren't just fixing the timing, they're doing it without breaking the actual speech recognition.

Jane: They call that "forgetting," which happens when you try to teach a model a new trick and it suddenly forgets how to do its old ones.

Meng: I've seen that happen in so many fine-tuning projects where the accuracy goes through the floor.

Lu: Imagine if we could surgically edit just the temporal part of the brain without touching the language part!

Lalam: That's the beauty of this research, as it moves us toward AI that can truly synchronize with the human experience of rhythm and pause.

Jane: It’s like giving a musician a metronome that only kicks in when they lose the beat.

Tom: Let's get into how they actually pull off this "surgical" edit without the model losing its mind.

Summary: Tom: We're moving into the meat of the paper now, looking at how REDDIT actually manages to fix that timing drift.

Jane: They use this two-stage approach that sounds almost like a rehearsal process for an actor.

Tom: An actor? How so?

Jane: Well, in the first stage, they let the model "replay" its own original, drifted output to keep the context stable.

Lu: It's a brilliant way to use the model's own existing knowledge as a foundation for the correction.

Meng: But how do they make sure the model doesn't just learn to repeat its own mistakes during that replay?

Lu: They use something called KL divergence, which basically acts like a tether to the original model's behavior for everything except the timestamps.

Jane: So, they're telling the model, "Keep your words exactly the same, but move these specific time markers to these new spots."

Tom: And then they have this second stage, right?

Jane: Yes, they take that corrected version and run a quick refinement stage using those new, correct timestamps as a starting point.

Meng: I'm curious about the data they used to train this, because you'd think you'd need humans to label all those timestamps.

Tom: That's the clever part—they didn't use any human annotations at all!

Jane: They just took clean speech, chopped it up, and inserted silent gaps where they already knew exactly how long the silence was.

Lu: It's a synthetic playground that allows for perfect, mathematical ground truth without any human error.

Meng: That sounds like a massive win for scaling, since you can generate infinite training examples this way.

Lalam: It allows the machine to learn the nuances of silence and speech transition in a way that feels natural to human culture.

Tom: It really is a clever way to build a teacher without needing a classroom of humans.

Improvements: Tom: The numbers in this paper are absolutely wild, especially when you look at the Whisper-tiny experiments.

Jane: I was just looking at that, and the jump in the long-gap mIoU from thirty-eight point seven percent to ninety-five point zero percent is just massive.

Tom: And the timing error, the AAS, dropped from two thousand seven hundred fifty-two milliseconds all the way down to two hundred twenty-three milliseconds!

Jane: That's the difference between a subtitle that's a huge distraction and one that's almost invisible.

Lu: What's even more impressive is that they did all of this by updating only one point six percent of the model's parameters.

Meng: Only one point six percent? That's incredibly efficient for an engineering team trying to optimize compute costs.

Tom: And they didn't fall into the "forgetting" trap that the other methods did.

Jane: Right, the standard fine-tuning actually caused the error rate to skyrocket to over five hundred percent in some cases!

Meng: That's a total disaster for any production environment.

Lu: It shows that the "distribution editing" approach is much more stable than just throwing more data at the problem.

Tom: They even showed it works on the larger Whisper-large-v3 model, though the scaling is a bit different there.

Jane: It seems like the two-stage process becomes even more critical as the models get bigger and more complex.

Lu: It's a scalable strategy that could eventually be applied to much larger audio-language models.

Lalam: This kind of precision will allow AI to participate in real-time translation and accessibility in ways that feel seamless and human.

Meng: If we can get this level of accuracy with such low overhead, it's a game changer for edge devices.

Tom: It really is, and it's hard to imagine going back to those drifting subtitles now.

Conclusion: Tom: We've covered a lot of ground today, from the frustration of drifting timestamps to this incredibly elegant solution.

Jane: It's one of those papers where the solution feels so much more "correct" than the brute-force methods we usually see.

Tom: They've really shown that you can edit a model's behavior without destroying its core intelligence.

Lu: This opens up so many doors for how we think about model adaptation and specialized task editing in the future.

Meng: I'm definitely looking at how we can implement this kind of parameter-efficient tuning in our own pipelines.

Lalam: Ultimately, this is about making technology more harmonious with the way we actually live and communicate.

Tom: That's a perfect place to wrap this up.

Jane: We'll be back with more research very soon, but thanks for joining us.

Tom: Goodbye, everyone!

More episodes

← Home