From Priors to Perception: Grounding Video-LLMs in Physical Reality
cs.CV
Submitted: 2026-05-06
Updated: 2026-09-19
Code: https://github.com/LiamZhao326/From-Priors-to-Perception
Terminology
Sources
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
- Video-R1: Reinforcing Video Reasoning in MLLMs
- NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
- WorldModelBench: Judging Video Generation Models As World Models
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
- TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
- SAM 2: Segment Anything in Images and Videos
- IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Kling-Omni Technical Report
- InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
- VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
- SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models