PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
cs.CV, cs.AI
Submitted: 2026-08-27
Updated: 2026-09-03
Code: https://github.com/Wan-Video/Wan2.2
Project page: https://pawbench.github.io
Terminology
Sources
- Cosmos 3: Omnimodal World Models for Physical AI
- VideoPhy: Evaluating Physical Commonsense for Video Generation
- Genie: Generative Interactive Environments
- On Calibration of Modern Neural Networks
- T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation
- World Models
- LTX-2: Efficient Joint Audio-Visual Foundation Model
- VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
- Couple to Control: Joint Initial Noise Design in Diffusion Models
- How Far is Video Generation from World Model: A Physical Law Perspective
- Improved Precision and Recall Metric for Assessing Generative Models
- Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
- How Confident are Video Models? Empowering Video Models to Express their Uncertainty
- World Models That Know When They Don't Know - Controllable Video Generation with Calibrated Uncertainty
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
- String Seed of Thought: Prompting LLMs for Distribution-Faithful and Diverse Generation
- Reliable Fidelity and Diversity Metrics for Generative Models
- Do Deep Generative Models Know What They Don't Know?
- Cosmos World Foundation Model Platform for Physical AI
- PICABench: How Far Are We from Physically Realistic Image Editing?
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models