PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud".
Jane: The paper was written by Chenghua Wang, Daliang Xu, Dongqi Cai, Duojin Sun, Hao Zhang et al. from Beijing University of Posts and Telecommunications and Nanjing University and Peking University and Tsinghua University and MingTi Technology and ModelBest.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the channel, everyone. Today we’re digging into a paper that’s been making waves in the robotics world, and it’s called “PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud.” Jane, I have to say, the title alone tells you they’re trying to solve two very different problems at once.
Jane: Absolutely, Tom. And that’s exactly what caught my eye. When we talk about physical AI, we mean robots that actually move and act in the real world, not just chatbots. The authors are saying, look, we need one system that works both when the robot is right there on the factory floor and when you’re training it in the cloud with thousands of GPUs.
Tom: Right, and that’s a big deal because usually those are two completely separate worlds. You have the onboard computer that has to be fast and cheap, and then you have the cloud training cluster that cares about throughput. The paper’s authors, from places like Beijing University of Posts and Telecommunications and Tsinghua, they’re basically saying, why do we need two different software stacks?
Jane: Exactly. They built something called PhyAI, which is a unified inference engine. Think of it like having one engine that can power both a go-kart and a race car. The go-kart needs to be nimble and responsive, the race car needs raw power, but the core mechanics are the same. That’s what they’re trying to do for robot brains.
Tom: And the numbers they’re seeing are pretty wild. They claim speedups anywhere from one point four times to over four point six times compared to the official implementations of models like π0, π0 point 5, and GR00T. That’s not a small improvement, that’s a game changer for real-time control.
Jane: But here’s the thing I love about this paper, Tom. They’re not just bragging about being the fastest. They’re honest about it. They say, look, in some configurations, other specialized runtimes are still faster than us. Their goal is to have one runtime that’s competitive everywhere, not the absolute fastest in every single case.
Tom: That’s a really mature take. And it speaks to the bigger picture, which is that the robotics field is fragmenting. Every new model has its own inference code, and that’s slowing everything down. PhyAI is trying to be the common ground, the one piece of software that can run all these different robot brains without needing to rewrite everything from scratch.
Jane: And that’s what we’re going to dig into for the rest of the show. How they actually pulled this off, what it means for real robots, and whether this could be the foundation that finally lets physical AI scale beyond the lab. Stick around, because this gets really interesting.
Summary: Tom: So Jane, we’ve set the stage with the title. Now let’s get into what this paper actually does. And I want to bring in Lu, our senior researcher, because I think the architecture here is what makes it special.
Lu: Thanks, Tom. So the core idea is pretty elegant. They split the system into two parts. On one side, you have model adapters, which know all the specific details of a particular robot model, like how it processes camera images or how it runs its action solver. On the other side, you have the shared runtime, which handles all the boring but critical stuff like memory management, kernel optimization, and parallel execution.
Jane: So it’s like having a universal power outlet. Each robot model brings its own plug adapter, but the wall socket and the wiring behind it are the same for everyone. That means when a new model comes out, like MiniCPM-Robot, they can add support for it in a single day, which is exactly what they did.
Meng: And that’s where I get excited as an engineer. Because in my world, the nightmare is always that you optimize something for one model, and then the next model comes along and breaks everything. PhyAI is saying, no, we’ve got a stable foundation, and we just swap in the model-specific logic.
Tom: But it’s not just about being easy to maintain. The paper shows real performance gains. They tested on everything from a Jetson Thor, which is a small edge device, up to eight H20 GPUs in the cloud. And they saw speedups across the board.
Lu: Right, and the most striking example to me is Cosmos3, which is a world-action model. That means it doesn’t just predict actions, it also predicts future video frames. On eight H20 GPUs, they cut the latency from two point four six seconds down to one point one eight seconds. That’s a two point zero eight times speedup, and that’s huge for a model that’s doing that much work.
Jane: And the paper is really careful about explaining why these models behave so differently. They did these detailed phase profiles, breaking down where the time actually goes. For π0 point 5, the action expert is only eight point eight percent of the FLOPs, but at batch size one, it takes up fifty-seven point two percent of the time. That’s because it’s launching tons of tiny kernels, and that overhead dominates.
Meng: Yeah, that’s the classic small-batch problem. You’re spending more time telling the GPU what to do than actually doing it. But when they scale up to batch size thirty-two that same action expert drops to thirteen point five percent of the time, and throughput hits about one hundred samples per second.
Lu: And that’s the key insight. Different models have different bottlenecks. Cosmos3 is compute-bound from the start, so batching doesn’t help it much. π0 point 5 is launch-bound at small batches, so batching helps a lot. A one-size-fits-all approach would miss these nuances, but PhyAI lets each model pick the right execution policy.
Tom: So it’s not just a faster engine, it’s a smarter engine that understands what each model actually needs. And that brings us to the really interesting part, which is how they think about control time and whether inference speed even matters in the real world. That’s coming up next.
Improvements: Jane: So Tom, we’ve covered what PhyAI does and why it’s fast. But the part I find most thought-provoking is their “control-time Roofline.” That’s the idea that just making inference faster doesn’t always make the robot faster. And I want to bring in Meng for this, because it’s a very practical engineering concern.
Meng: Yeah, and this is something I think about constantly. You have a robot that executes an action, and that takes a certain amount of time. Then it runs inference to figure out the next action. If the action execution takes longer than the inference, then inference speed doesn’t matter for the control rate. You’re environment-bound, not inference-bound.
Tom: So it’s like if you’re baking cookies and the oven takes twenty minutes, but your recipe takes five minutes to prepare. Speeding up the recipe to two minutes doesn’t make the cookies come out any faster. The oven is the bottleneck.
Jane: Exactly. And the paper actually measured this with π0 point 5 on four LIBERO task suites. They found that the model inference was already faster than the environment execution. So those robots are environment-bound. That means you could actually use a slower, cheaper GPU and still hit the same control rate.
Lu: But that’s not the whole story. For Cosmos3, which does that heavy video prediction, inference is still the bottleneck. So there, every millisecond you save on inference directly translates to faster control. The Roofline helps you figure out which regime you’re in, so you know where to spend your optimization effort.
Meng: And that’s the improvement I appreciate. Instead of just blindly optimizing inference speed, PhyAI gives you a framework to decide if that’s even worth it. If you’re environment-bound, you can spend your money on a better robot arm instead of a better GPU.
Tom: And they take it further with real-time chunking, which they call RTC. The idea is that you generate the next action chunk while the robot is still executing the current one. So you overlap the inference with the action execution, hiding the latency entirely.
Jane: But there’s a catch, right? Because the next chunk is based on an older observation. The robot is acting on information that’s slightly stale. So the paper says you have to evaluate the trade-off between latency and task success. You can’t just assume overlap is always better.
Lu: And that’s the kind of nuance that makes this paper stand out. They’re not just throwing out a speedup number and calling it a day. They’re saying, here’s the tool to understand whether that speedup matters for your specific robot, your specific task, and your specific hardware budget.
Meng: And for me, that’s the real improvement. It turns inference optimization from a black art into a systematic engineering decision. You can measure, you can predict, and you can make the right call for your deployment.
Tom: And that’s exactly the kind of thinking that’s going to push physical AI from the lab into real factories and homes. But we’ve got one more thing to talk about, which is what this means for the future of robot training and deployment. Let’s wrap this up.
Conclusion: Tom: Alright, let’s bring it home. We’ve been talking about “PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud,” and I think we’ve only scratched the surface. Jane, what’s the big takeaway for you?
Jane: For me, it’s the unification. This paper is saying that the same codebase can run a robot on a Jetson Thor in your living room and a massive training rollout on eight H20 GPUs in a data center. That’s a huge step toward making physical AI practical, because you’re not maintaining two separate systems that can drift apart.
Lu: And I’d add that the control-time Roofline is a genuinely new way of thinking. It forces you to ask whether your bottleneck is the model or the environment. That’s going to change how people allocate resources, both in terms of hardware budget and engineering time.
Meng: From my side, the fact that they’re honest about the limitations is refreshing. They admit that specialized runtimes are still faster in some cases. But they’re building for the long term, for a world where you have many different models and you need one stable foundation.
Tom: And the numbers speak for themselves. Speedups from one point four to four point six five times across eleven different model-hardware pairs. And they even showed that in a simulated RL rollout, using PhyAI as the inference backend could cut the step time by over twenty-six percent, which means training faster.
Jane: But it’s not just about speed. It’s about the ecosystem. They’re releasing the code and benchmark harnesses, so other researchers can build on this. And they’ve already shown they can add a brand-new model in a single day. That’s the kind of momentum that could really accelerate the field.
Lu: And I think the future work section is exciting too. They’re talking about mega-kernels to fuse those tiny operations, kernel agents to automatically optimize for new hardware, and even a standard protocol for robot model serving. That’s a roadmap for the next few years.
Tom: So as we say goodbye to this paper, I think the message is clear. Physical AI is moving from one-off demos to a real engineering discipline. And PhyAI is building the infrastructure to make that happen. We’ll be keeping an eye on this one.
Jane: Absolutely. And thanks to everyone for listening. Next up, we’ve got a paper on world models that I think is going to blow your mind. See you then.
Chenghua Wang, Daliang Xu, Dongqi Cai, Duojin Sun, Hao Zhang, Haoze Qian, Huaiyuan Zhang, Jinshuo Cui, Junbo Cui, Kezhao Zhao, Longxi Gao, Mengwei Xu, Rongjie Yi, Tam Sikyuen, Tianyue Zhang, Weikai Xie, Xuanzhe Liu, Yingying Qin, Yiwen Lu, Yuan Yao, Yuezhi Zu, Yunhan Guo, Yuxin Zheng, Ziqi Guo
Beijing University of Posts and Telecommunications · Nanjing University · Peking University · Tsinghua University · MingTi Technology · ModelBest
cs.AI, cs.RO
Submitted: 2026-08-14
Updated: 2026-08-17
Comments: 25 pages, 9 figures
Code: https://github.com/mingti-org/phyai
Project page: https://opengalaxea.github.io/G05/Galaxea_
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 65/100
The gist: "Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning (RL) rollout, edge GPU serving, and onboard deployment.
Key concepts
- Physical AI
- Refers to robots that move and act in the real world, distinguishing them from chatbots. The goal is to create systems that can operate both on local edge hardware (like factory floors) and in powerful cloud training environments.
- Unified Inference Engine
- A single software system, like PhyAI, designed to power various robot models. It standardizes the core mechanics (memory management, kernel optimization) so that new models can be added without rewriting the entire software stack.
- Control-Time Roofline
- A concept that suggests simply making inference faster doesn't guarantee a faster robot. It helps determine if a system is limited by computation (inference) or by the physical environment's speed (action execution).
- Edge vs. Cloud Deployment
- The challenge of running AI models in two different environments: fast, cheap onboard computers (the Edge) versus high-throughput, powerful data centers (the Cloud). PhyAI aims to solve this discrepancy.
Terminology
Summary
Summary
The paper introduces PhyAI, a unified inference runtime for Physical AI policies, designed to provide a single inference path across the entire model lifecycle, including offline evaluation, cloud reinforcement learning (RL) rollout, shared edge serving, and onboard deployment. The authors state: "Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning (RL) rollout, edge GPU serving, and onboard deployment. Although these settings use the same checkpoint and action semantics, they often rely on separate inference programs. To provide one inference path across these settings, we build PhyAI, an Physical AI inference engine with a unified runtime."
The core architectural principle is a separation between model-specific logic and shared execution services. The paper explains: PhyAI keeps architecture-specific conditioning, solver, cache, and output logic in model adapters while sharing graph execution, kernels, memory management, and parallel services.
This design allows the same codebase to run vision-language-action (VLA) models and world-action models (WAMs) on single or multiple GPUs across onboard, edge, and cloud deployments. The authors note that The same codebase currently runs pi0, pi0.5, GR00T N1.7, Cosmos3-Nano-Policy, Cosmos3-Edge, and MiniCPM-Robot on devices ranging from Jetson Thor to RTX 5090, A40, H100, and multi-GPU H20 and A100 servers.
They also highlight the extensibility of the adapter interface: The adapter interface also allowed us to add MiniCPM-Robot on the day of its release without introducing a separate runtime.
The paper reports significant performance improvements over official implementations. The authors state: PhyAI achieves 1.40× to 4.65× speedups over the official implementations of pi0, pi0.5, GR00T N1.7, and MiniCPM-Robot.
Specific results include: On Cosmos3-NanoPolicy-DROID, it reduces latency from 2.46 to 1.18 s on eight H20 GPUs with CFG=2 and TP=4, a 2.08× speedup.
The paper provides a detailed table of speedups: "On RTX 5090, PhyAI is 4.05× faster for pi0, 1.82× for pi0.5, 2.28× for GR00T, and 2.02× for MiniCPM-Robot. On Thor, the speedups are 2.64×, 1.67×, 1.40×, and 2.78×, respectively. We also measured a 4.31× speedup for GR00T on A40, 4.65× for MiniCPM-Robot on H100, and 2.08× for Cosmos3-Nano-Policy-DROID on eight H20 GPUs with CFG=2 and TP=4. The authors are careful to note that
Specialized runtimes remain faster in several configurations, so our goal is one runtime with competitive latency rather than the fastest result in every case."
The paper provides detailed phase and batch profiles to explain why different models require different execution policies. For instance, "On a Hopper-series GPU at batch size one, the pi0.5 action expert accounts for 8.8% of estimated FLOPs but 57.2% of profiled latency. At batch size 32, its share falls to 13.5%, throughput reaches about 100 samples/s, and the full batch takes about 320ms. In contrast,
Cosmos3 remains generation dominated and gains only 14.3% throughput as batch size increases from 1 to 16. For GR00T, the analysis shows that
At batch size one, Action Head accounts for 59.2% of profiled time and Backbone for 40.7%. Backbone becomes the larger phase at batch size four and reaches 69.7% at batch size 32."
A key conceptual contribution is the control-time Roofline,
which the authors introduce to distinguish inference-bound from environment-bound control.
The formal definition is: "Let Linference denote the inference time and Lenv the time required to execute the current action and advance the environment. Under ideal overlap, the slower stage determines the control period: Loverlap = max(Linference, Lenv). The paper explains the implications:
When inference is slower, the system is inference bound and lower model latency directly improves control. When environment execution is slower, the system is environment bound and further acceleration creates timing margin rather than a higher ideal control rate. Applying this analysis, the authors find that
The measured pi0.5 points on four LIBERO suites are environment bound, while Cosmos3 remains inference bound."
The paper also presents a simulated RL rollout workload analysis. The authors describe: "We further simulated an integrated RL rollout workload with PhyAI as its inference backend. The setup used eight A100 GPUs, a batch size of 40, and 41 policy-inference calls per RL step. Inference accounted for 53.1% of the baseline step time and 36.2% with PhyAI, a decrease of 16.9 percentage points, or 31.8% relative. With the non-inference time held fixed in the simulation, this change corresponds to a 26.5% reduction in rollout-step latency and about 1.36× higher training throughput."
The paper positions PhyAI relative to existing systems. It notes that "LeRobot covers data collection, training, evaluation, and a generalized asynchronous policy-inference stack that can decouple remote action prediction from robot control. It does not, however, provide the same kernel-level, multi-architecture execution engine targeted by PhyAI. Similarly,
vLLM and SGLang offer mature token-oriented LLM serving, but their execution model does not directly represent reusable multimodal conditions, iterative action solvers, evolving world latents, or classifier-free-guidance branches."
The paper outlines several areas for future work, including Mega Kernels
to address small-operator overhead, Climbing the Performance Mountain with Kernel Agents
for architecture-specific kernel tuning, making PhyAI as an RLinf Rollout Backend,
expanding Broader Model and Accelerator Support,
and developing A Production-Grade Inference Serving Protocol.
The authors conclude: "We therefore want one runtime that can follow a model wherever it needs to run. The same model path should work for robot MaaS and cloud RL rollouts, as well as edge-cloud and fully onboard execution. Model adapters keep its behavior intact, while the runtime adjusts batching, parallelism, kernels, and placement for each setting."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to an AI system, along with what the improved system can do:
Improvement: Replace the current fragmented inference stack (separate code paths for evaluation, RL rollout, edge serving, and onboard deployment) with a single runtime that separates model-specific logic (adapters) from shared execution services (graph replay, kernels, memory management, parallelism).
What the improved AI system can do:
-
Run the same policy checkpoint across cloud, edge, and onboard hardware without reimplementation or behavior drift.
-
Add new models (e.g., MiniCPM-Robot) within a day by implementing only the adapter, not the entire runtime.
-
Achieve 1.40×–4.65× speedups over official implementations on matched hardware (e.g., π0.5: 1.82× on RTX 5090; Cosmos3: 2.08× on 8×H20 with CFG=2, TP=4).
Improvement: Introduce a control-time Roofline that distinguishes inference-bound from environment-bound control, using the formula: L overlap = max(L inference, L env). This provides a clear stopping point for optimization.
Improvement: Use profiled operator selection and phase-level analysis to dynamically choose kernels and batch sizes based on the bottleneck phase (e.g., action expert vs. vision-language backbone vs. generation).
Improvement: Implement a flexible parallel execution model that supports data parallelism (DP), tensor parallelism (TP), and classifier-free-guidance (CFG) parallelism in a unified way, including a 2×N TP mesh for CFG branches.
Improvement: Implement explicit validity scopes for reusable state (e.g., vision-language prefix KV cache, timestep tables) and capture the full denoising loop in CUDA Graphs.
Improvement: Support W4A16, W8A16, W8A8, and W4A8 quantization via RTN, using high-performance operators from FlashInfer and Humming.
Improvement: Integrate PhyAI as a drop-in inference backend for RLinf, replacing the default Hugging Face path.
Improvement: Prototype a timed, stateful serving protocol with keyed observations, delta updates, action horizons, and deadline-aware backpressure.
The improved AI system can:
-
Run any VLA or WAM model on any device (Jetson Thor to multi-GPU H20/A100) with one codebase.
-
Achieve 1.4×–4.65× latency improvements over official baselines.
-
Automatically adapt batching, parallelism, and kernels to the bottleneck phase.
-
Distinguish inference-bound vs. environment-bound control and optimize accordingly.
-
Integrate seamlessly into RL training pipelines, reducing rollout time and improving training throughput.
-
Deploy in edge, cloud, or onboard settings with consistent behavior and measurable control margins.
These improvements are directly actionable and grounded in the paper’s measured results, making them safe to implement in production systems.
Sources
- Cosmos 3: Omnimodal World Models for Physical AI
- ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RT-1: Robotics Transformer for Real-World Control at Scale
- LeRobot: An Open-Source Library for End-to-End Robot Learning
- Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
- AVO: Agentic Variation Operators for Autonomous Evolutionary Search
- StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
- RynnBrain: Open Embodied Foundation Models
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel
- OpenVLA: An Open-Source Vision-Language-Action Model
- RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
- CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
- OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection