PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
cs.CV
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/Helen1p/PhysVista
Terminology
Sources
- Phi-4 Technical Report
- ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
- Qwen3-VL Technical Report
- VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
- CoPhy: Counterfactual Learning of Physical Dynamics
- IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
- PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
- CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
- "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
- PAC Bench: Do Foundation Models Understand Prerequisites for Executing Manipulation Policies?
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- GRASP: A novel benchmark for evaluating language GRounding And Situated Physics understanding in multimodal language models
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
- 4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models
- PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
- IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning
- WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
- PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
- Kimi K2.5: Visual Agentic Intelligence
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models