Relative Time Intervals Representation for Word-level Timestamping with Masked Training

arXiv:2608.24041 · cs.AI · Submitted 2026-08-25 · Read on arXiv

cs.AI

Submitted: 2026-08-25

Updated: 2026-08-25

Comments: ICASSP2026 Accpeted

Code: https://github.com/tangquanwei/Timestamp-Aware-Speech-LLM

License: http://creativecommons.org/licenses/by/4.0/

The gist: Although Speech Large Language Models (SpeechLLMs) excel at speech understanding and generation, their capacity for fine-grained, temporally aligned outputs remains underexplored.

Terminology

Abstract

Although Speech Large Language Models (SpeechLLMs) excel at speech understanding and generation, their capacity for fine-grained, temporally aligned outputs remains underexplored. Our work addresses this gap by enabling SpeechLLMs to jointly model speech content and temporal structure, effectively transforming them from ``content understanding machines" into ``temporal-aware content understanding machines". Specifically, we replace traditional absolute timestamps with relative timestamps, achieving a more compact vocabulary and stronger generalization capabilities. To efficiently infuse timestamp prediction ability into pre-trained large language models, we introduce a hybrid fine-tuning strategy: full-parameter fine-tuning of the timestamp-augmented embedding layer and language model head, combined with LoRA fine-tuning of the decoder layers. Moreover, we design a masked timestamp training objective, preventing the model from over-relying on ground-truth timestamps, and thereby enhancing robustness against noisy real-world annotations. Extensive experiments demonstrate that our approach achieves significant improvements in timestamp prediction accuracy while maintaining strong speech transcription performance.

Sources

Related papers