HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration
cs.CV, cs.AI, cs.GR
Submitted: 2026-06-26
Updated: 2026-09-05
Comments: Accepted to ECCV 2026. 15 pages of main text and 39 pages of appendices. Project page: https://lijiaxin0111.github.io/HAT4D/
Project page: https://lijiaxin0111.github.io/HAT4D
License: http://creativecommons.org/licenses/by/4.0/
The gist: Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs.
Terminology
Abstract
Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for scaling Embodied AI and training VLAs. However, existing monocular 4D reconstruction methods primarily focus on isolated objects, often failing under the severe occlusions and complex dynamics inherent in multi-object interactions. To bridge this gap, we propose HAT-4D, the first agentic framework designed to reconstruct the 3D geometry, temporal dynamics, and physical interactions of multiple objects from a single video. By integrating VLMs with a multi-level human-in-the-loop feedback mechanism, HAT-4D efficiently resolves depth ambiguities and interaction-induced occlusions during 3D generation and 4D propagation, yielding physically plausible assets without relying on expensive multicamera rigs. As a scalable data engine, HAT-4D facilitates the creation of MVOIK-4D, an open-world benchmark for monocular 4D interaction reconstruction, accompanied by a novel multi-dimensional evaluation protocol focused on physical plausibility and temporal consistency. Extensive experiments demonstrate that HAT-4D achieves SOTA performance on most evaluation metrics, while maintaining competitive semantic alignment. Ablation studies show that introducing a small amount of human feedback improves interaction reconstruction. Moreover, the data produced by HAT-4D effectively improves baseline performance when used for fine-tuning. Our data and code are available at https://lijiaxin0111.github.io/HAT4D/
Sources
- Recent Progress on a Manifold Damped and Detuned Structure for CLIC
- Qwen3-VL Technical Report
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- Segment Anything in 3D with Radiance Fields
- Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis
- V2M4: 4D Mesh Animation Reconstruction from a Single Monocular Video
- Affordance Grounding from Demonstration Video to Target Image
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- LRM: Large Reconstruction Model for Single Image to 3D
- 4D Visual Pre-training for Robot Learning
- VBench: Comprehensive Benchmark Suite for Video Generative Models
- HY3D-Bench: Generation of 3D Assets
- Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis
- Consistent4D: Consistent 360{\deg} Dynamic Object Generation from Monocular Video
- Segment Anything
- Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
- FB-4D: Spatial-Temporal Coherent Dynamic 3D Content Generation with Feature Banks
- SGTR: End-to-end Scene Graph Generation with Transformer
- 4D-LRM: Large Space-Time Reconstruction Model From and To Any View at Any Time
- DINOv3
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models