RESUME: Recurrent State Updates from Motion and Residual Signals for Efficient Video Language Modeling
cs.CV
Submitted: 2026-09-30
Updated: 2026-09-30
Project page: https://microsoft.github.io/CoPE
Terminology
Sources
- LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
- Qwen2.5-VL Technical Report
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- AdaCodec: A Predictive Visual Code for Video MLLMs
- VideoChat: Chat-Centric Video Understanding
- TempCompass: Do Video LLMs Really Understand Videos?
- CoPE-VideoLM: Leveraging Codec Primitives For Efficient Video Language Modeling
- LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
- OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence
- TEAM-Net: Multi-modal Learning for Video Action Recognition with Partial Decoding
- VisionZip: Longer is Better but Not Necessary in Vision Language Models
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models