PhysAI-Bench: A Benchmark for LLM-Based Agentic Decision-Making in Autonomous UAV-Centric Physical AI
cs.AI
Submitted: 2026-09-20
Updated: 2026-09-20
Code: https://github.com/maferrag/physai-bench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent advances in Physical AI have accelerated the use of foundation models in autonomous systems such as unmanned aerial vehicles (UAVs), which must perceive, reason, plan, and act in dynamic
Terminology
Abstract
Recent advances in Physical AI have accelerated the use of foundation models in autonomous systems such as unmanned aerial vehicles (UAVs), which must perceive, reason, plan, and act in dynamic environments. Existing benchmarks assess physical perception, intuitive physics, embodied navigation, and collaborative reasoning, but rarely evaluate the agentic decision-making required for reliable autonomy. We introduce PhysAI-Bench, a benchmark for evaluating this capability. It contains 10,178 standardized decision instances automatically extracted from conversational traces of autonomous UAV missions. Each instance preserves mission context, temporal dependencies, physical constraints, Model Context Protocol (MCP) tool calls, Agent-to-Agent (A2A) interactions, sensor observations, and AI-native 6G network conditions, including latency, packet loss, throughput, edge load, and network slicing. We expose only information preceding each decision, preventing future-event leakage and approximating online decision-making. We evaluate 29 foundation models using a two-stage protocol. We select model-specific configurations from 12 combinations of zero-, three-, and five-shot prompting and four temperatures, tested in three runs on a 35-instance, human-verified development set. We then freeze each selected configuration and evaluate it in three runs on a fixed, episode-disjoint set of 500 instances. GPT-5.3 achieves the highest accuracy (52.00%), followed by GPT-5.2 (49.40%) and Grok 4.5 (49.07%). Few-shot prompting generally improves performance, while temperature has limited influence. The results demonstrate that reliable agentic decision-making in Physical AI remains an open challenge. The dataset is available at https://github.com/maferrag/physai-bench
Sources
- Gemini Robotics: Bringing AI into the Physical World
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
- ACWM-Phys: Investigating Generalized Physical Interaction in Action-Conditioned Video World Models
- Cosmos 3: Omnimodal World Models for Physical AI
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
- T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation
- IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
- Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration
- RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation
- ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
- SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection