REDDIT: Forgetting-Resistant Correction of Timestamp Drift in ASR via Replay-Based Distribution Editing

arXiv:2607.05364 · cs.CL, cs.AI, cs.SD · Submitted 2026-07-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing".

Jane: The paper was written by Cheng-Kang Chou, Ming-Douo Tchouang, Ke-Han Lu, Chan-Jan Hsu and Hung-yi Lee from National Taiwan University and Carnegie Mellon University and National Taiwan University Artificial Intelligence Center of Research Excellence (NTU AI-CoRE).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome to the show, where we break down the latest breakthroughs in research, and today we're looking at a paper titled REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing.

Jane: That is quite a mouthful, Tom, but the problem they're solving is something we all encounter when watching videos.

Tom: You mean when the subtitles don't quite match the person speaking?

Jane: Exactly, it's that annoying delay or jump where the words are right but the timing is totally off.

Lu: It's actually a much deeper problem than just a glitchy subtitle, because it shows a fundamental disconnect in how these models perceive time.

Tom: So the authors, a team from NTU and CMU, are saying the model knows *what* was said but loses track of *when* it happened?

Jane: That's a great way to put it, especially since these models generate the time as part of their own text output.

Lu: I find the concept of "drift" so fascinating because it implies the model's internal clock is essentially drifting away from reality during silence.

Meng: From a practical standpoint, if I'm building a real-time transcription tool, that drift makes the whole system feel broken or unreliable.

Tom: And that's why this paper is so important, because they aren't just fixing the timing, they're doing it without breaking the actual speech recognition.

Jane: They call that "forgetting," which happens when you try to teach a model a new trick and it suddenly forgets how to do its old ones.

Meng: I've seen that happen in so many fine-tuning projects where the accuracy goes through the floor.

Lu: Imagine if we could surgically edit just the temporal part of the brain without touching the language part!

Lalam: That's the beauty of this research, as it moves us toward AI that can truly synchronize with the human experience of rhythm and pause.

Jane: It’s like giving a musician a metronome that only kicks in when they lose the beat.

Tom: Let's get into how they actually pull off this "surgical" edit without the model losing its mind.

Summary: Tom: We're moving into the meat of the paper now, looking at how REDDIT actually manages to fix that timing drift.

Jane: They use this two-stage approach that sounds almost like a rehearsal process for an actor.

Tom: An actor? How so?

Jane: Well, in the first stage, they let the model "replay" its own original, drifted output to keep the context stable.

Lu: It's a brilliant way to use the model's own existing knowledge as a foundation for the correction.

Meng: But how do they make sure the model doesn't just learn to repeat its own mistakes during that replay?

Lu: They use something called KL divergence, which basically acts like a tether to the original model's behavior for everything except the timestamps.

Jane: So, they're telling the model, "Keep your words exactly the same, but move these specific time markers to these new spots."

Tom: And then they have this second stage, right?

Jane: Yes, they take that corrected version and run a quick refinement stage using those new, correct timestamps as a starting point.

Meng: I'm curious about the data they used to train this, because you'd think you'd need humans to label all those timestamps.

Tom: That's the clever part—they didn't use any human annotations at all!

Jane: They just took clean speech, chopped it up, and inserted silent gaps where they already knew exactly how long the silence was.

Lu: It's a synthetic playground that allows for perfect, mathematical ground truth without any human error.

Meng: That sounds like a massive win for scaling, since you can generate infinite training examples this way.

Lalam: It allows the machine to learn the nuances of silence and speech transition in a way that feels natural to human culture.

Tom: It really is a clever way to build a teacher without needing a classroom of humans.

Improvements: Tom: The numbers in this paper are absolutely wild, especially when you look at the Whisper-tiny experiments.

Jane: I was just looking at that, and the jump in the long-gap mIoU from thirty-eight point seven percent to ninety-five point zero percent is just massive.

Tom: And the timing error, the AAS, dropped from two thousand seven hundred fifty-two milliseconds all the way down to two hundred twenty-three milliseconds!

Jane: That's the difference between a subtitle that's a huge distraction and one that's almost invisible.

Lu: What's even more impressive is that they did all of this by updating only one point six percent of the model's parameters.

Meng: Only one point six percent? That's incredibly efficient for an engineering team trying to optimize compute costs.

Tom: And they didn't fall into the "forgetting" trap that the other methods did.

Jane: Right, the standard fine-tuning actually caused the error rate to skyrocket to over five hundred percent in some cases!

Meng: That's a total disaster for any production environment.

Lu: It shows that the "distribution editing" approach is much more stable than just throwing more data at the problem.

Tom: They even showed it works on the larger Whisper-large-v3 model, though the scaling is a bit different there.

Jane: It seems like the two-stage process becomes even more critical as the models get bigger and more complex.

Lu: It's a scalable strategy that could eventually be applied to much larger audio-language models.

Lalam: This kind of precision will allow AI to participate in real-time translation and accessibility in ways that feel seamless and human.

Meng: If we can get this level of accuracy with such low overhead, it's a game changer for edge devices.

Tom: It really is, and it's hard to imagine going back to those drifting subtitles now.

Conclusion: Tom: We've covered a lot of ground today, from the frustration of drifting timestamps to this incredibly elegant solution.

Jane: It's one of those papers where the solution feels so much more "correct" than the brute-force methods we usually see.

Tom: They've really shown that you can edit a model's behavior without destroying its core intelligence.

Lu: This opens up so many doors for how we think about model adaptation and specialized task editing in the future.

Meng: I'm definitely looking at how we can implement this kind of parameter-efficient tuning in our own pipelines.

Lalam: Ultimately, this is about making technology more harmonious with the way we actually live and communicate.

Tom: That's a perfect place to wrap this up.

Jane: We'll be back with more research very soon, but thanks for joining us.

Tom: Goodbye, everyone!

National Taiwan University · Carnegie Mellon University · National Taiwan University Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)

cs.CL, cs.AI, cs.SD

Submitted: 2026-07-06

Updated: 2026-09-15

Comments: Accepted to IEEE Spoken Language Technology Workshop (SLT 2026)

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 84/100

The gist: This paper introduces REDDIT (REplay-based Distribution eDITing), a lightweight post-training framework designed to correct "non-speech-induced timestamp drift" in autoregressive Automatic Speech

Key concepts

Timestamp Drift
The phenomenon where an ASR model correctly identifies spoken words but loses track of when they occurred. This often happens during silences, causing a disconnect where the text is accurate but the timing is out of sync with the audio.
Forgetting
A problem in fine-tuning where teaching a model a new skill causes it to lose its original capabilities, such as speech recognition accuracy. REDDIT avoids this by surgically editing temporal markers without disrupting the model's language understanding.
Replay-Based Distribution Editing
A two-stage method that uses the model's original output to maintain context and stability. It uses KL divergence to keep words unchanged while shifting timestamps, followed by a refinement stage using synthetic data—clean speech with known silent gaps—to ensure precise timing.

Terminology

Summary

This paper introduces REDDIT (REplay-based Distribution eDITing), a lightweight post-training framework designed to correct non-speech-induced timestamp drift in autoregressive Automatic Speech Recognition (ASR) systems. As modern models increasingly generate timestamps as decoded tokens, they risk assigning plausible content while assigning it to the wrong region of the audio following long silent spans. REDDIT provides a method for forgetting-resistant temporal editing, allowing models to correct their temporal grounding without degrading their primary speech recognition capabilities.

The challenge of timestamp drift

The authors identify a failure mode where the transcript remains linguistically acceptable, but the decoded time axis drifts away from the audio. This is distinct from lexical hallucination; while hallucination emits unsupported words, timestamp drift can place supported words at the wrong time. This problem is often exacerbated by training data that lacks exposure to long non-speech prefixes, which strengthens a prior for early timestamp tokens near zero seconds.

Furthermore, standard approaches to fixing this issue are problematic. The researchers found that naive timestamp-corrected fine-tuning can improve alignment on targeted data but severely degrade non-target ASR behavior, exposing a catastrophic forgetting problem. Consequently, the task must be treated as a model-editing problem that moves timestamp predictions to correct positions while preserving the decoder’s original non-timestamp distribution.

The REDDIT framework

REDDIT is a two-stage post-training pipeline that constructs correction supervision without requiring human transcripts or manual timestamp annotations. Instead, it creates synthetic training examples by combining:

  • VAD-trimmed speech spans.

  • Inserted non-speech gaps.

  • Known concatenation offsets.

The framework operates through two specific stages to ensure stability and accuracy. In Stage 1, the model edits timestamp targets under a cached replay context while anchoring non-timestamp behavior to the frozen base distribution via KL divergence. This ensures that timestamp positions are optimized... while matching the frozen base distribution on non-timestamp tokens. Stage 2 then applies a short edited-prefix refinement from the Stage-1 checkpoint to consolidate these corrected transitions, allowing the model to learn from its own edited prefixes.

Experimental performance and scalability

Evaluations on Whisper-tiny demonstrate that REDDIT achieves high temporal accuracy with minimal parameter updates. By updating only 1.6% of parameters, the framework achieved significant improvements:

  • Raised long-gap mIoU from 38.7% to 95.0%.

  • Reduced mixed-gap out-of-domain AAS from 2752 ms to 223 ms.

  • Preserved CV-en MER at 41.3%, whereas ordinary SFT decoder tuning resulted in a catastrophic 524.2% MER.

The study also indicates that REDDIT is a scalable strategy. When applied to Whisper-large-v3, the two-stage schedule proved particularly effective, with the Stage-2 refinement helping larger decoders use the corrected timestamp trajectory while retaining the Stage-1 replay anchor. This confirms that REDDIT can successfully separate temporal correction from recognizer preservation, maintaining high recognition quality even when temporal behavior is heavily edited.

Improvements for AI systems

Improvements to implement:

  • Implement a Two-Stage Post-Training Framework (REDDIT):

  • Stage 1 (Replay-based Distribution Editing): Train the model using cached replay contexts from the frozen base model. Use cross-entropy loss to optimize corrected timestamp targets while simultaneously applying KL divergence to anchor non-timestamp tokens to the original frozen teacher distribution. This prevents the model from drifting away from its original lexical knowledge.

  • Stage 2 (Edited-Prefix Refinement): Perform a secondary refinement stage where the corrected timestamps from Stage 1 are used as teacher-forcing prefixes to consolidate temporal transitions and stabilize inference-time behavior.

  • Deploy an Automated Timestamp Supervision Pipeline: Construct training datasets by concatenating VAD-trimmed speech spans with sampled non-speech segments at known, precise temporal offsets. This allows for the generation of ground-truth timestamp targets without requiring expensive human transcription or manual timecode annotations.

  • Apply Targeted Parameter-Efficient Fine-Tuning (PEFT): Limit weight updates strictly to the last cross-attention layers and layer norms (targeting <2% of total parameters). This minimizes the risk of catastrophic forgetting in the decoder's language modeling capabilities.

Capabilities of the improved AI system:

  • Native Temporal Grounding: The system can generate highly accurate, self-contained timestamps as part of its decoded output, eliminating the need for external inference-time modules like Voice Activity Detection (VAD), forced aligners, or post-processing alignment scripts.

  • Drift-Resistant Long-Form Transcription: The system can maintain precise temporal alignment even when encountering long periods of silence, background noise, or non-speech spans that typically cause timestamp drift in standard autoregressive models.

  • High-Fidelity Subtitling: It provides millisecond-level accuracy for start and end timecodes (low AAS/MAE) while preserving high lexical transcription quality (low MER), making it suitable for professional-grade automated subtitling.

  • Robust Out-of-Domain Performance: The system maintains its ability to transcribe diverse languages and acoustic environments accurately, even after being specifically optimized for temporal precision in a target domain.

Sources

Related papers