IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training
cs.AI, cs.RO
Submitted: 2026-08-31
Updated: 2026-08-31
License: http://creativecommons.org/licenses/by/4.0/
The gist: World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions.
Terminology
Abstract
World models have made remarkable progress in action-conditioned future prediction for embodied agents, yet still struggle to model physically plausible interactions. Existing approaches address this limitation by constraining the generation process with external representations encoding motion, geometry, or semantics. Obtaining these spatiotemporally dense representations typically requires auxiliary estimators or manual annotations, limiting training scalability. We instead revisit the training objective and identify a supervision-allocation mismatch under the globally averaged mean squared error (MSE) denoising objective: prevalent static content dominates the optimization signal, leaving sparse dynamic-object regions critical to interaction generation disproportionately under-supervised. Motivated by this observation, we introduce IMPACT, a scalable Interaction-aware Model training framework with Prior-guided Attention Calibration and Targeting. IMPACT uses cross-attention associated with manipulated-object tokens as an internal spatiotemporal prior for action-conditioned changes. It samples candidate regions from this prior, calibrates them with detached local prediction errors to construct an interaction map, and uses the map to reweight denoising supervision, requiring neither external representations nor inference-time modifications. Extensive experiments on robot-arm and human-hand manipulation, spanning diverse control modalities and DiT backbones, show that IMPACT consistently outperforms the corresponding MSE-trained baselines, improving interaction fidelity, physical plausibility, and visual quality.
Sources
- AlayaWorld: Long-Horizon and Playable Video World Generation
- Genie: Generative Interactive Environments
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework
- Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control
- Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
- Multiplayer Interactive World Models with Representation Autoencoders
- VACE: All-in-One Video Creation and Editing
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Geometry-aware 4D Video Generation for Robot Manipulation
- CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
- Cosmos World Foundation Model Platform for Physical AI
- Qwen2.5 Technical Report
- WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
- WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory and In-Context Learning
- STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation
- Towards Accurate Generative Models of Video: A New Metric & Challenges
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection