SYNCR: Diagnosing and Learning Cross-Video Reasoning from Simulation
cs.CV, cs.LG
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
- ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models
- LLaVA-OneVision: Easy Visual Task Transfer
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
- SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents
- CLEVRER: CoLlision Events for Video REpresentation and Reasoning
- MLVU: Benchmarking Multi-task Long Video Understanding
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- CVBench: Benchmarking Cross-Video Synergies for Complex Multimodal Reasoning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models