What 30,000 Hours of Ego-centric Video Does Not Teach
cs.CV
Submitted: 2026-10-08
Updated: 2026-10-08
Project page: https://fidelity-gap.github.io
Terminology
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Qwen3-VL Technical Report
- PointZero: 3D Point Track Completion for Learning Transferable 3D Dynamics
- Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control
- Mastering Diverse Domains through World Models
- Training Compute-Optimal Large Language Models
- Scaling Laws for Neural Language Models
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- Cosmos 3: Omnimodal World Models for Physical AI
- EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
- Towards Accurate Generative Models of Video: A New Metric & Challenges
- OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics
- EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
- PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
- EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models