Augmenting Large Audio-Language Models with Frame-Level Grounding for Fine-Grained Temporal Perception

arXiv:2609.15215 · cs.SD, cs.AI · Submitted 2026-09-14 · Read on arXiv

cs.SD, cs.AI

Submitted: 2026-09-14

Updated: 2026-09-14

Comments: Submitted to ICASSP 2027

Project page: https://saber5203.github.io/frame-level-grounding

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Sources

Related papers