A Comprehensive Review of Generative Physical Artificial Intelligence
cs.RO, cs.AI, cs.CL, cs.CV, cs.LG
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 25 pages, 8 figures
Journal ref: IEEE Internet of Things Journal, vol. 13, no. 10, pp. 20452-20476, 15 May 2026
DOI: 10.1109/JIOT.2026.3671268
License: http://creativecommons.org/licenses/by/4.0/
The gist: The integration of large-scale foundation models with physical embodiments has led to significant advancements in robotics known as Generative Physical Artificial Intelligence (GPAI).
Terminology
Abstract
The integration of large-scale foundation models with physical embodiments has led to significant advancements in robotics known as Generative Physical Artificial Intelligence (GPAI). These agentic AI systems autonomously perceive, reason, and act in complex real-world situations. This survey comprehensively analyzes GPAI systems, focusing on their architectural foundations, current applications, and key limitations. We introduce a taxonomy of five distinct approaches: Robot Foundation Models (RFMs) for cross-platform skill transfer; Vision-Language Action (VLA) models for end-to-end multi-modal perception and control; Large Behavior Models (LBMs) for human-like movement generation; Diffusion Policy Models (DPMs) for diffusion model-based temporally coherent action generation; and World Foundation Models (WFMs) for physics-compliant simulation and data generation. We examine how these approaches complement each other: WFMs generate training data for VLAs and DPMs, RFMs enable cross-platform deployment of learned policies, while LBMs provide motion priors for natural behavior. Through examples across autonomous vehicles, industrial automation, healthcare robotics, and humanoid systems, we identify significant performance improvements and summarize promising research directions in data-efficient learning, sim-to-real transfer, edge-compatible architectures, and safety frameworks. These insights advance embodied AI for IoT-connected environments where intelligent agents interact with networked sensors, actuators, and edge devices.
Sources
- Generative Artificial Intelligence in Robotic Manipulation: A Survey
- Generative Physical AI in Vision: A Survey
- A Survey on Robotics with Foundation Models: toward Embodied AI
- A Comprehensive Survey on World Models for Embodied AI
- Industry 6.0: New Generation of Industry driven by Generative AI and Swarm of Heterogeneous Robots
- Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges
- REASSEMBLE: A Multimodal Dataset for Contact-rich Robotic Assembly and Disassembly
- Cosmos World Foundation Model Platform for Physical AI
- Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets
- Proximal Policy Optimization Algorithms
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
- REALM: A Real-to-Sim Validated Benchmark for Generalization in Robotic Manipulation
- Defining and Evaluating Physical Safety for Large Language Models
- RT-1: Robotics Transformer for Real-World Control at Scale
- Zero-Shot Whole-Body Humanoid Control via Behavioral Foundation Models
- GraspClutter6D: A Large-scale Real-world Dataset for Robust Perception and Grasping in Cluttered Scenes
- Unified Vision-Language-Action Model
- A Generalist Agent
- Exploring GPT-4 for Robotic Agent Strategy with Real-Time State Feedback and a Reactive Behaviour Framework
- Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving