RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
cs.CV, cs.CL, cs.RO
Submitted: 2026-06-01
Updated: 2026-08-31
Project page: https://huiqiongli.github.io/RoboTrustBench
Terminology
Sources
- Rethinking Video Generation Model for the Embodied World
- World Simulation with Video Foundation Models for Physical AI
- Qwen3-VL Technical Report
- Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Test
- GigaWorld-0: World Models as Data Engine to Empower Embodied AI
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- Large Video Planner Enables Generalizable Robot Control
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models
- WorldEval: World Model as Real-World Robot Policies Evaluator
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- Robotic Video World Models: A Survey of Applications, Research Challenges, Future Directions
- Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
- Causal World Modeling for Robot Control
- Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
- Advancing Open-source World Models
- WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- HunyuanVideo 1.5 Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models