TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models".
Jane: The paper was written by Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Utathya Aich, Ramani Duraiswami et al. from University of Maryland, College Park, USA and CNH Industrial India.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: In our second segment, we're exploring the core mechanics of TEMPO: Temporally-ground Multi-task Post-training for Large Audio-Language Models. How does a model actually learn these intricate temporal skills from scratch?
Jane: The researchers use a supervised fine-tuning process that introduces three key innovations to teach the model time, which is where things get really clever.
Lu: They’ve introduced atomic timestamp tokens, essentially forcing the AI to predict timestamps as discrete categories at zero point one-second intervals rather than relying on its standard word fragmentation.
Meng: That atomic token approach is crucial for ensuring that the temporal supervision isn't broken down by the BPE vocabulary, which is a major headache in large language models.
Lalam: It's about enabling AI to perceive the rhythm of an audio stream, not just its content, allowing us to build better tools for appreciating pacing in music or conversation flow.
Tom: And Jane mentioned that this unified framework allows all five tasks to be related together as one integrated system.
Jane: Exactly; it’s not just running five different pipelines but making one model that recognizes how these tasks—like diarization and ASR—are fundamentally related in time, which is a huge step up from previous silo-based approaches.
Lu: The paper says they also use a time-aware multi-modal projector to inject sinusoidal wall-clock encodings into the audio frames, providing the model with temporal context that its standard features lack.
Meng: This is an elegant solution for grounding, ensuring the practical impact of having strong initial capabilities across all tasks without needing excessive pre-processing.
Lalam: It allows us to create a more cohesive experience where the AI can understand speech, sound, and music in a single stream without losing track of time or context.
Tom: This unified structure is what gives TEMPO: Temporally-ground Multi-task Post-training for Large Audio-Language Models its powerful framework.
Jane: We're going to look at how this training regimen translates into real performance gains in the next segment.
Improvements: Tom: Moving into our third segment, we’re looking at the concrete performance improvements that make TEMPO: Temporally-ground Multi-task Post-training for Large Audio-Language Models stand out against the current state of the art.
Jane: The data shows significant gains across all five tasks compared to models like Qwen3-Omni and Audio Flamingo Next, which is a huge win for users who are using these tools.
Lu: The most impressive wins are in speech tasks, specifically looking at multi-speaker ASR where they’ are cutting the Word Error Rate from zero point seven zero down to an incredible zero point four four—a massive improvement in accuracy.
Meng: That level of reduction in error means much more reliable transcription for any professional use case, allowing us to trust the AI output when it matters most, like in legal or scientific review.
Lalam: It’s about the ability to trust the AI's output and knowing that it has accurately placed that speech in time, enabling us to build better digital archives of human interaction for future generations.
Tom: These gains are driven by those three innovations we discussed—atomic timestamp tokens, a time-aware projector, and a distance-aware Gaussian loss.
Jane: The atomic timestamp tokens are key because they prevent the BPE vocabulary from fragmenting our temporal supervision, forcing the model to predict accurate time intervals directly.
Lu: That prevents the fragmentation of our temporal data, which is vital for accurately tracking multiple events over time rather than just seeing them as a single blob of audio.
Meng: And when you combine that with the distance-aware Gaussian loss, it’s essentially teaching the AI to reward near-miss predictions, which is much more aligned with how humans perceive timing than just penalizing every single incorrect timestamp.
Lalam: This allows us to move beyond "close enough" and start aiming for true precision in time, enabling a culture of highly accurate digital documentation.
Tom: We also need to talk about their reinforcement learning approach, which is another major addition to the process.
Jane: They are using Group Relative Policy Optimization or GRPO as a refinement stage on top of the SFT model, which is fascinating because it’s designed to optimize for real-world evaluation metrics.
Lu: This is interesting because they are applying RL with verifiable temporal rewards, directly optimizing against the actual metrics we use in testing, not just some general reasoning score.
Meng: This shows a very targeted approach; instead of trying to make the model smarter generally, it’s making its timestamps accurate and useful for real-world tasks.
Lalam: It helps us refine our understanding of time, making the whole system feel more reliable and trustworthy when we can verify that a prediction is temporally sound.
Tom: We'll see how this refinement translates into the final performance metrics in the last segment.
Conclusion: Tom: In our fourth segment, let's look at what the final results of TEMPO: Temporally-ground Multi-task Post-training for Large Audio-Language Models say about its success and what it all means for the future of AI.
Jane: The data clearly shows that TEMPO performs significantly better than other state-of-the-art models across all five tasks when we test against our held-out benchmarks.
Lu: The fact that they are achieving these gains primarily through their careful SFT design, rather than just brute force or massive training budgets, is a huge indicator of skill in model architecture.
Meng: It’s clear that the combination of atomic tokens and the time-aware projector is proving to be a very efficient way to solve these complex temporal problems practically without needing massive compute resources.
Lalam: This means we are moving toward an era where AI can truly grasp the timeline of audio, which is vital for how we process our collective history and artistic works across cultures.
Tom: We also noticed that even though they added this extra RL refinement stage, the performance gains were modest compared to SFT, right?
Jane: That suggests that while reinforcement learning helps with fine-tuning boundaries, the primary lever for adding this crucial timestamping capability is indeed the initial supervised training.
Lu: It’s a fascinating insight into how these different stages of training interact and what each contributes to the overall system's success in a truly multi-task way.
Meng: This confirms that we are finding more robust ways to teach models timing, allowing for reliable systems that have been specifically designed for temporal accuracy.
Lalam: It’s about empowering AI to truly understand the story told by sound, even when multiple stories are happening at once in a single stream.
Tom: Before we wrap up, Lu, any final thoughts on this work?
Lu: I think we're seeing a new way to treat time as an explicit part of the core model training.
Meng: Just ensuring that every practical implementation can now handle these high levels of temporal accuracy is a major relief for me.
Lalam: It’s about empowering AI to truly understand the story told by sound, even when multiple stories are happening at once in a single stream.
Tom: And Jane, you?
Jane: I think it' really makes us excited to see what other applications this will unlock for our listeners.
Wrap-up: Tom: Well, that’s a lot of ground covered today on the paper titled TEMPO: Temporally-ground Multi-task Post-training for Large Audio-Language Models.
Jane: It’s clear that this is more than just a small improvement; it's a foundational shift in how we approach audio understanding in AI, giving us tools that are much more reliable.
Lu: The paper is showing us that we can teach AI not just the *meaning* of sound, but the precise *structure* by mastering these complex temporal tasks.
Meng: I believe this allows for incredibly robust and reliable systems across all practical applications where timing matters.
Lalam: It’s about ensuring that the next generation of audio-based AI is respectful of the flow of life and sound, creating a better digital world for our listeners.
Tom: We want to thank all the authors for this work and hope they' are having a successful rollout in their future research endeavors.
Jane: We're looking forward to seeing how other researchers build on this foundation, using TEMPO: Temporally-ground Multi-task Post-training for Large Audio-Language Models as is.
Lu: It’s a powerful example of the impact that detailed engineering can have on the big picture AI capabilities in this field.
Meng: I hope we’re able to integrate these advancements into real-world products quickly now thanks to this approach.
Lalam: I believe that this will lead to a more nuanced and temporally aware digital culture for all our listeners as we move forward.
Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Utathya Aich, Ramani Duraiswami, Dinesh Manocha
University of Maryland, College Park, USA · CNH Industrial India
cs.SD, cs.AI, cs.LG
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: Accepted at EMNLP 2026 Main Conference. Project page - https://kaousheik-26.github.io/tempo
Project page: https://kaousheik-26.github.io/tempo
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: TEMPO (Temporally-grounded Multitask Post-training) is the first unified Large Audio-Language Model (LALM) designed to handle timestamping across audio, speech, and music.
Key concepts
- Atomic Timestamp Tokens
- These tokens force the AI to predict timestamps as discrete categories at zero point one-second intervals. This method is crucial because it prevents the standard BPE vocabulary from breaking down the temporal supervision, allowing for accurate time perception.
- Time-aware Multi-modal Projector
- This component injects sinusoidal wall-clock encodings into audio frames. It provides the model with essential temporal context that standard audio features typically lack, helping ground the model's understanding in time.
- Temporally-ground Multi-task Post-training (TEMPO)
- TEMPO is a unified framework that trains a single model to recognize how multiple tasks—like diarization and ASR—are fundamentally related in time. This is an improvement over previous silo-based approaches.
- Group Relative Policy Optimization (GRPO)
- This reinforcement learning technique acts as a refinement stage on the main model. It is designed to optimize the system's performance directly against real-world evaluation metrics, ensuring practical accuracy.
Terminology
Summary
TEMPO (Temporally-grounded Multitask Post-training) is the first unified Large Audio-Language Model (LALM) designed to handle timestamping across audio, speech, and music. While existing LALMs typically treat audio understanding as a clip-level task
and discard temporal structure, TEMPO addresses this foundational capability
gap by enabling precise temporal grounding. This is essential for downstream applications such as speech recognition and dense audio captioning.
The Five Timestamping Tasks
TEMPO is designed to perform a wide range of temporal tasks within a single decoder, distinguishing them via task-specific tags in the prompt. The model is trained to handle:
-
Multi-speaker automatic speech recognition (ASR)
-
Speaker diarization
-
Audio temporal grounding
-
Dense audio captioning
-
Timestamped music captioning
By unifying these tasks, the model can predict who, what, and when within a general-purpose framework
spanning speech, sound, and music.
Supervised Fine-Tuning Innovations
The authors introduce a supervised fine-tuning (SFT) stage built on three core innovations to teach a model with no prior temporal supervision to predict accurate timestamps.
These innovations include:
-
Atomic timestamp tokens: To prevent BPE from
fragmenting timestamps across subword pieces,
the tokenizer is extended with dedicated tokens at 0.1 s resolution. -
Time-aware multi-modal projector: This component injects
sinusoidal wall-clock encodings into each frame embedding,
providing the wall-clock information that frame-indexed features lack. -
Distance-aware Gaussian loss: To address the fact that token-level cross-entropy
penalizes all incorrect timestamps equally,
this auxiliary loss assigns partial credit to near-miss predictions, encouraging the model to learn theordinal structure of the timestamp vocabulary.
The training follows a two-stage synthetic-to-real curriculum
to establish temporal calibration before exposure to real-world data.
Reinforcement Learning via GRPO
Beyond SFT, the researchers present the first application of reinforcement learning to unified audio timestamping
using Group Relative Policy Optimization (GRPO). This stage acts as a refinement stage on top of the SFT checkpoint,
utilizing verifiable temporal rewards
that directly optimize the evaluation objectives. The reward structures are task-specific and designed to align the model with temporal metrics:
-
ASR and diarization rewards average metrics such as (1 − WER), symmetric mIoU, and boundary scores.
-
Dense audio captioning rewards combine soft event F1, symmetric mIoU, and METEOR.
-
Music captioning rewards utilize a weighted sum of chord accuracy, instrument accuracy, tempo accuracy, and statistics accuracy.
Performance and Results
On a benchmark containing 10K samples, TEMPO substantially outperforms Audio Flamingo Next and Qwen3-Omni,
which are state-of-the-art LALMs explicitly trained on timestamped data. The most significant gains are observed in speech tasks, where TEMPO cuts WER from 0.70 to 0.44 on multi-speaker ASR
and improves diarization mIoU from 0.44 to 0.71. The experiments confirm that while SFT delivers most of these gains,
the GRPO refinement provides consistent but moderate refinements
to temporal boundaries.
Improvements for AI systems
1. Implementation of Atomic Timestamp Tokenization
-
Improvement: Replace standard BPE-tokenized numeric timestamps with a dedicated vocabulary of atomic timestamp tokens (e.g., t) at a fixed 0.1s resolution.
-
Capability: The system can perform precise, single-token temporal predictions, eliminating the semantic fragmentation and inconsistent tokenization errors that occur when Large Language Models attempt to predict numeric strings.
2. Integration of a Time-Aware Multi-Modal Projector
-
Improvement: Replace standard frame-indexed projectors with a time-aware architecture that injects sinusoidal wall-clock encodings (logarithmically spaced to cover sub-second to long-range durations) into audio frame embeddings.
-
Capability: The system can bridge the gap between frame-level audio features and absolute wall-clock time, enabling accurate temporal grounding and long-range musical or speech structure recognition.
3. Adoption of Distance-Aware Gaussian Loss
-
Improvement: Replace standard token-level cross-entropy for timestamp prediction with an auxiliary distance-aware Gaussian loss that applies a soft-label distribution centered at the ground-truth time.
-
Capability: The system learns the ordinal structure of time, ensuring that
near-miss
temporal predictions are treated as partially correct rather than being penalized equally todistant-miss
errors, directly aligning training with temporal Intersection-over-Union (tIoU) metrics.
4. Deployment of a Synthetic-to-Real Training Curriculum
-
Improvement: Implement a two-stage training pipeline that begins with a large-scale synthetic corpus (to establish temporal calibration and token usage) followed by fine-tuning on high-quality, human-annotated real-world data.
-
Capability: The system achieves high temporal accuracy and linguistic fluency by first mastering the mechanics of timestamping in a controlled environment before adapting to the complexities and noise of real-world audio.
5. Application of GRPO with Verifiable Temporal Rewards
-
Improvement: Apply Group Relative Policy Optimization (GRPO) during the post-training stage, utilizing reward functions derived directly from task-specific temporal metrics (e.g., symmetric mIoU, Word Error Rate, Diarization Error Rate, and chord accuracy).
-
Capability: The system undergoes direct optimization of its evaluation objectives, allowing it to refine boundary precision and linguistic accuracy across a unified multi-task framework including multi-speaker ASR, speaker diarization, audio temporal grounding, dense audio captioning, and structured music analysis.
Abstract
Large audio-language models (LALMs) describe audio at the clip level but cannot assign timestamps to the events, speakers, or sounds they identify. Despite being essential for downstream tasks like speech recognition and dense audio captioning, timestamping remains a key limitation of most LALMs. We present TEMPO (Temporally-grounded Multi-task Post-training), the first unified model to handle audio, speech, and music timestamping tasks. Our core contribution is a supervised fine-tuning (SFT) stage built on three innovations: atomic timestamp tokens, a time-aware projector that injects sinusoidal wall-clock encodings into audio frame embeddings, and a distance-aware Gaussian loss. Our training is based on a synthetic-to-real curriculum. We further introduce, to our knowledge, the first application of reinforcement learning to unified audio timestamping, using GRPO with verifiable temporal rewards that directly optimize the evaluation objectives. Rather than serving as the primary source of performance gains, GRPO acts as a refinement stage on top of the SFT checkpoint, providing modest additional improvements. To support this work, we build a training dataset containing 119K samples and an evaluation benchmark containing 10K samples, drawn from established corpora across five tasks. On this benchmark, TEMPO outperforms Audio Flamingo Next and Qwen3-Omni, two state-of-the-art LALMs explicitly trained on timestamped data. Experiments confirm that SFT delivers most of these gains, with GRPO providing consistent but moderate refinements.
Sources
- Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
- madmom: a new Python Audio and Music Signal Processing Library
- BEATs: Audio Pre-Training with Acoustic Tokenizers
- Qwen2-Audio Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Clotho: An Audio Captioning Dataset
- Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
- Music Flamingo: Scaling Music Understanding in Audio Language Models
- Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
- GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
- Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
- Kimi-Audio Technical Report
- Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
- A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models
- TAC: Timestamped Audio Captioning
- Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
- Cutting Music Source Separation Some Slakh: A Dataset to Study the Impact of Training Data Quality and Quantity
- GPT-4 Technical Report
- The Benefit Of Temporally-Strong Labels In Audio Event Classification
- TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment