MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues
cs.CV, cs.AI
Submitted: 2026-05-21
Updated: 2026-09-27
Project page: https://ddz16.github.io/mllmsknowwhen.github.io
Terminology
Sources
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- An Empirical Study on How Video-LLMs Answer Video Questions
- GPT-4o System Card
- Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
- See What You Are Told: Visual Attention Sink in Large Multimodal Models
- Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
- VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
- TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
- Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
- OpenAI GPT-5 System Card
- Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
- Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
- MiMo-VL Technical Report
- Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration
- Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models