When Should the Count Change? Learning State Maintenance for Causal Video Counting
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Project page: https://placeholder.github.io/StaMina
Terminology
Sources
- Layer Normalization
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
- Gaussian Error Linear Units (GELUs)
- OpenVLA: An Open-Source Vision-Language-Action Model
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Kimi K2.5: Visual Agentic Intelligence
- Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning
- OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
- Decoupled Weight Decay Regularization
- ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better
- End-to-End Test-Time Training for Long Context
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Cambrian-S: Towards Spatial Supersensing in Video
- Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models