LPA-CWM: A Learned Physical Adjudicator for Motion Reasoning with Counterfactual World Models
cs.CV, cs.AI, cs.RO
Submitted: 2026-09-12
Updated: 2026-10-02
Comments: A quick overview is available at https://LPA-CWM.github.io
Project page: https://lpa-cwm.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Counterfactual world models (CWM) extract motion from pretrained video predictors by comparing factual and intervened predictions.
Terminology
Abstract
Counterfactual world models (CWM) extract motion from pretrained video predictors by comparing factual and intervened predictions. However, responses generated under different target-frame masks vary in reliability, while uniform aggregation weights them equally. We formulate response aggregation as candidate reliability learning and propose LPA-CWM with a lightweight Learned Physical Adjudicator (LPA). Trained on dense MOVi-F trajectories, the 3.0M-parameter LPA compares visual context and response structure across an unordered candidate set to predict relative weights, while the CWM predictor and intervention generator remain frozen. The weighted responses undergo windowed localization and one paired re-evaluation to recover motion. We also introduce Completeness-aware Motion Correspondence (CMC), a ground-truth-anchored evaluation protocol that jointly measures localization, trajectory completeness, visibility, and continuity, counting missing predictions as failures on visible dynamic points. On the evaluated DAVIS and Kinetics subsets, LPA-CWM improves DCA avg over Uniform CWM by 60.0% and 29.0%, respectively, and also improves tracking accuracy under TAP-Vid First. A quick overview is available at https://LPA-CWM.github.io.
Sources
- Revisiting Feature Prediction for Learning Visual Representations from Video
- CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models
- World Models
- The Kinetics Human Action Video Dataset
- Taming generative video models for zero-shot optical flow extraction
- Fr'echet Video Motion Distance: A Metric for Evaluating Motion Consistency in Videos
- The 2017 DAVIS Challenge on Video Object Segmentation
- WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
- RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation
- VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models