PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration
summary
The gist
PhysCaP introduces a Physics-Informed Code-as-Policy agent designed for active perception in robotic manipulation, addressing the limitations of vision-language policies by integrating
In short
PhysCaP enhances Code-as-Policy agents for robotics by adding physics-informed exploration. It uses training-free modules to estimate hidden physical properties like an object's mass and stiffness directly from robot movements. This allows the agent to intelligently decide when and where to interact with the environment, leading to better task performance with fewer interactions.
Key concepts
- Physical Property Extraction Modules (PhysX)
- These are modules that estimate hidden physical traits, such as an object's mass or stiffness, using only data from the robot's joints. For instance, they calculate mass by analyzing the difference in joint torques during a grasp. This works without needing extra sensors.
- Planner Agent
- This agent decides the strategy for exploration. It looks at what physical information is missing and sets rules for when to start exploring and, crucially, when to stop interacting with the environment. It acts as a stopping criterion based on gathered evidence.
- Prioritizer Agent
- This agent refines the exploration plan proposed by the Planner. It uses visual clues and heuristics to filter out poor or redundant interaction candidates. This ensures that the agent focuses its limited interactions only on those most likely to yield important physical information.
Terminology used across episodes
This episode discusses
- PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration · Paper Radio
- Octo: An Open-Source Generalist Robot Policy
- pi 0.5: a Vision-Language-Action Model with Open-World Generalization
- MolmoAct2: Action Reasoning Models for Real-world Deployment
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Gemini Robotics: Bringing AI into the Physical World
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
- MEM: Multi-Scale Embodied Memory for Vision Language Action Models
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Phys2Real: Fusing VLM Priors with Interactive Online Adaptation for Uncertainty-Aware Sim-to-Real Manipulation
- SAVOR: Skill Affordance Learning from Visuo-Haptic Perception for Robot-Assisted Bite Acquisition
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- Qwen3-VL Technical Report
The paper
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration · Read on arXiv
National Taiwan University · NVIDIA Research
We present PhysCaP, a Physics-Informed Code-as-Policy agent system for active perception in robotic manipulation. While vision-language-action policies excel at imitating demonstrations, they rely on passive observation and fail to infer latent physical properties critical for manipulation. PhysCaP augments code-as-policy frameworks with a physics-informed exploration layer that enables explicit information-seeking through interaction. Our method introduces training-free physical property extraction modules that estimate object mass and stiffness from robot proprioception without additional sensors. To balance exploration costs and the efficiency of information obtained, PhysCaP employs a multi-agent design: a Planner that decides when to explore and when to stop, and a Prioritizer that filters implausible interactions and ranks the remainder using a heuristic priority score, enabling efficient, targeted exploration. We evaluate PhysCaP on three real-world tabletop manipulation tasks and a simulated task in LIBERO. The results show that existing passive and naive interactive baselines either fail when physical properties are hidden or over-explore, whereas PhysCaP achieves comparable performance with fewer interactions and reduced execution time. Ablation studies further validate the effectiveness of the proposed physical property extraction modules. Project page: https://physcap.github.io
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration".
Dev: PhysCaP introduces a Physics-Informed Code-as-Policy agent designed for active perception in robotic manipulation, addressing the limitations of vision-language policies by integrating physics-informed exploration to infer latent physical properties.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So we’ve just walked through the paper "PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration," which is really about giving these code as policy agents a way to actually learn what things are made of without needing extra sensors.
Dev: I agree, Rosa, it seems like the core idea is that vision and language models are great at looking at pictures, but they struggle when you need to know something physical about the object itself, like its weight or how stiff it is.
Taro: Exactly. The paper addresses that gap by adding this physics-informed exploration layer so the agent can actively seek out those missing properties through interaction with the environment.
Rosa: That sounds like a big step because it moves us from just guessing based on what we see to actually measuring the physical reality of what we're handling.
Dev: The methodology they propose involves training-free modules, specifically Physical Property Extraction Modules, that can estimate things like object mass and stiffness directly from the robot's joint torques and proprioceptive feedback.
Taro: That’s compelling because it means we don't need to build complex tactile sensors just to figure out basic properties like density or rigidity; the system infers them from how the robot moves.
Rosa: And then they pair that extraction with a dual-agent design where a Planner decides when to explore and stop, and a Prioritizer filters out bad interactions using heuristics.
Dev: That planning agent acting as a dynamic stopping criterion is interesting because it means the exploration isn't just running until it runs out of plan; it stops once the gathered information is deemed sufficient for the task.
Taro: And the Prioritizer refining that plan by filtering implausible interactions sounds like a smart way to manage interaction costs, ensuring we aren't wasting time on things that don't lead us closer to the answer.
Rosa: So, instead of just blindly exploring everything in a room, the agent uses these physics constraints to decide exactly which physical interactions are worth making next.
Dev: It’s about balancing the cost of interaction against the information gained, which is crucial for any real-world deployment where time and energy matter.
Taro: If you think about what happens when the world misbehaves, like an object behaving unexpectedly or resisting a certain movement, this framework should allow it to query those physical properties to adapt its strategy.
Rosa: Speaking of adaptation, the paper shows how this approach handles tasks where visual cues are ambiguous, like distinguishing between similar containers.
Title and authors: Dev: The results show that PhysCaP can achieve comparable performance with fewer interactions and reduced execution time compared to the baselines they tested, which is a solid metric for efficiency.
Taro: That reduction in interaction count is significant because every physical interaction takes time and energy, so reducing that overhead directly translates to better real-world feasibility.
Rosa: I'm curious about how this works outside of a perfectly controlled lab setting; can these physical property estimates hold up when the environment gets messy?
Dev: The paper tests it on three challenging tabletop tasks—finding hidden cubes, identifying empty cans, and selecting ripe avocados—and also includes a simulated empty-can task in LIBERO.
Taro: The fact that it’s tested on real-world scenarios like finding a hidden cube shows that the framework isn't just theoretical; it’s grounded in actual manipulation challenges.
Rosa: The results are pretty impressive when they show that existing passive baselines often fail when physical properties are hidden or if they just keep exploring too much.
Dev: I noticed they detail how mass is estimated using a formula like mˆ = (Jz · ∆τ)/(gJz2) by isolating the gravitational contribution of the object from joint torque differentials, which sounds very concrete.
Taro: That specific measurement technique for mass based purely on torque differentials provides a solid physical prior that feeds into the downstream reasoning tasks, which is exactly what we needed.
Rosa: Then there’s stiffness estimation, where they use a two-phase procedure: first checking for true contact via a backoff test to rule out friction, and then measuring the displacement needed to reach a target effort threshold.
Dev: That backoff test is important because it helps ensure that the stiffness measurement isn't just an artifact of surface friction, which would skew the result.
Taro: If we consider future autonomy research, this capability means an agent could potentially assess material properties in real-time during a complex manipulation sequence without needing pre-programmed knowledge of those materials.
Rosa: What about the limitations mentioned in the paper? The authors flag a few things they don't cover perfectly or where it falls short.
Dev: They point out that one limitation is the reliance on commercial Vision-Language Model APIs, which introduces latency into the loop, and another issue is dependence on 2D predictions for object localization <ref:2608.21031#pg0>.
Taro: I think those are real hurdles; if you're relying on a 2D model for pose estimation, you introduce uncertainty that could cause trajectory divergence in physical tasks <ref:2608.21031#pg0>.
Rosa: And they also mention hardware communication latency affecting the trajectories, so we can't just assume perfect real-time performance without careful tuning.
Title and authors: Dev: So while the framework is powerful, you still have those practical constraints of current infrastructure and model outputs to deal with before it’s fully deployed in a high-stakes setting.
Taro: The paper clearly states that future work will aim to address these issues by transitioning to locally hosted models and incorporating multi-view or three dee-native models, which shows they have a clear path forward <ref:2608.21031#pg0>.
Rosa: So, wrapping up the main points of "PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration," it’s an agent that actively acquires missing physical information through interaction using physics extraction modules and a dual-agent planning system.
Dev: It successfully balances exploration cost against information efficiency by stopping when sufficient evidence is gathered, leading to fewer interactions and less time spent on the task.
Taro: The implication for autonomy is that we can build agents that don't just rely on visual semantics but ground their actions in measurable physical truths like mass and stiffness, making them more robust in unpredictable settings.
Rosa: It certainly seems like a very promising direction for achieving more reliable manipulation capabilities outside of perfectly sterile laboratory environments.
Dev: We need to keep watching how they tackle those latency and localization issues, because those are the practical bottlenecks for getting this kind of active perception into widespread use.
Taro: I think the long-term impact is enabling a level of environmental understanding that was previously only accessible through direct physical measurement, which opens up a lot more complex manipulation possibilities.
Rosa: That's what we were talking about, Taro—moving from observation to active sensing guided by physics. We’ve covered the title and authors, then walked through how the framework works and what it can do.
Dev: We also touched on the specific improvements they propose regarding efficient exploration via the Prioritizer and how it cuts down on redundant interactions.
Taro: And we looked at their conclusion, which summarizes that PhysCaP provides a way to achieve grounding in the physical world by integrating code-as-policy with physics-informed exploration.
Rosa: It really is an important piece of work because it shows how to make agents actively seek out the data they need rather than just passively observing what's there.
Dev: I think we’ve got a good handle on the mechanics and the trade-offs between performance and computational cost discussed in this paper.
Taro: We should keep an eye on their future work regarding those local models, because that will be key to moving this from a strong research concept to something genuinely deployable for complex autonomous systems.
The paper's summary: Rosa: So, to recap what we've heard so far, PhysCaP is an agent that uses physics knowledge to help a code-as-policy system figure out what physical objects are actually made of through active interaction.
Dev: Right, and the core mechanism involves these training-free modules that let the AI infer properties like mass and stiffness just by looking at how the robot moves and reacts.
Taro: That’s interesting because it shifts the focus from just interpreting visual data to actually sensing physical reality during manipulation.
Rosa: Exactly, and what I find compelling is how they use a dual-agent system—a Planner that decides when to explore and a Prioritizer that filters out bad ideas—to keep those physical interactions efficient.
Dev: That efficiency is key for me; if the loop rate drops because the agent is over-exploring, the entire control system falls apart, so minimizing those interactions sounds like a big win for real-time execution.
Taro: I agree with Dev there; and from an autonomy standpoint, this active information seeking means the AI can handle unexpected situations in ways that purely passive vision models simply can't.
Rosa: It really moves us toward systems that aren't just reacting to what they see, but are actively probing the environment to build a true physical model of the task at hand.
Dev: And those physics-informed priors, like knowing an object’s mass beforehand, should make the downstream reasoning tasks much more stable and less prone to visual ambiguities.
Taro: If we can reliably tell if something is empty or ripe based on measured stiffness rather than just a visual guess, that opens up entirely new capabilities for complex environments.
Rosa: It makes me wonder how long this kind of active exploration strategy will be viable outside of the highly controlled tabletop experiments the authors describe.
Dev: That’s a fair question, Rosa; we have to keep watching those latency and localization issues mentioned in the paper to see if they can push it into more demanding real-world scenarios without too much tuning.
Taro: My focus stays on how robust this active probing is when things get messy, like when an object behaves unexpectedly or is partially obscured.
Rosa: We'll definitely keep that in mind as we look at the next part of the paper, which dives into those specific experiments they ran on different manipulation tasks.
The paper's improvements: Tom: We've talked about how PhysCaP uses physics to help an agent understand objects, and now we're looking at what they suggest to make it even better than it already is.
Rosa: So, essentially, the authors are proposing specific enhancements to the dual-agent framework and the property extraction modules that could push this technology further.
Dev: Right, and I'm keen to hear how these improvements tackle those practical issues we talked about earlier regarding loop rate and stability in control loops.
Taro: From an autonomy viewpoint, I’m really interested in how they suggest making the agent's decision-making process more robust when things get physically weird or unpredictable.
Rosa: The paper suggests refining the Prioritizer Agent by incorporating more sophisticated visual heuristics to make the filtering of bad exploration plans even smarter and faster.
Dev: That sounds like it should directly translate into reduced computational load because it means we're discarding implausible paths earlier in the decision cycle, which is exactly what we need for a tight control loop.
Taro: If those heuristics are tied to physical intuition—like knowing that a certain shape shouldn't behave in a certain way—that gives the AI more of a sense of "common sense" in its exploration strategy.
Rosa: Furthermore, they propose making the physical property estimation modules more adaptable so they can generalize better across different types of objects without needing completely new training for every single item.
Dev: Generalization is crucial; if we have to retrain the property estimator from scratch every time we introduce a new material or object type, that defeats the purpose of having a reusable physics prior.
Taro: That generalization capability means this framework could theoretically be applied to entirely new domains, not just the specific manipulation tasks they tested on in their paper.
Rosa: I’m also intrigued by their suggestions for making the system more model-agnostic, which means the core exploration logic can stay consistent even if we swap out the underlying language model.
Dev: That's a big deal for deployment; if the core physics reasoning is decoupled from a specific VLM, we have much more flexibility in choosing how to deploy it across different hardware platforms.
Taro: That decoupling means the autonomy layer can evolve independently of the perception layer, which should allow us to build more flexible systems that can adapt their interaction strategy on the fly.
Rosa: It sounds like these improvements are focused on making the system less brittle and more broadly applicable to a wider variety of physical problems.
Dev: I'm still thinking about how much reduction in latency those heuristic refinements actually yield when running at high frequencies, which is what I need to get this into a production-level environment.
Taro: We need concrete data on the performance gains under stress, though; it’s great to see the theoretical potential for better robustness in handling novel physical interactions.
Rosa: So we’ve seen that their next steps focus on making the system more flexible and smarter in its filtering mechanisms, which is a solid direction.
Conclusion: Rosa: So we're wrapping up our discussion on PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration, which shows how adding physical awareness to code policies can solve real problems in manipulation.
Dev: Indeed, we’ve seen how this architecture balances the need for physical understanding against the practical constraints of execution speed and latency in a control loop.
Taro: I'm still thinking about the autonomy aspect; if we can reliably use measurable physical properties to guide exploration when the environment throws curveballs, that means agents could become much more resilient in complex, real-world settings.
Rosa: It really does suggest that future agents won't just be visual interpreters; they’ll be systems grounded in verifiable physical facts about what they are touching.
Dev: And from an engineering standpoint, the efficiency gains in terms of interaction count and execution time are significant for any practical robot deployment where battery life or cycle speed matters.
Taro: When we think about the broader impact, this could allow us to tackle a much wider variety of physical tasks that currently require deep prior knowledge that we can’t just teach them through observation alone.
Rosa: It seems like the authors have laid a very solid foundation for how code-as-policy agents can start actively sensing and reasoning about their physical interactions in a way we haven't seen before.
Dev: I hope the future work addresses those latency issues head-on because if the loop rate degrades too much, all this careful planning just becomes irrelevant.
Taro: We should definitely keep an eye on how they address those hardware communication delays; that’s where a lot of real-world performance will be won or lost.
Rosa: That sounds like a great focus for the next paper, so I'm excited to see how they tackle those practical deployment challenges.
More episodes
- 2610.12154-Stochastic Distribution Network Reconfiguration under Load Uncertainty
- 2607.00148-3D Point World Models: Point Completion Enables More Accurate Dynamics Learning
- 2607.02403-ACID: Action Consistency via Inverse Dynamics for Planning with World Models
- 2510.26623-A Sliding-Window Filter for Online Continuous-Time Continuum Robot State Estimation
- 2406.13267-The Kinetics Observer: A Tightly Coupled Estimator for Legged Robots
- 2511.02147-Census-Based Population Autonomy For Distributed Robotic Teaming
- 2603.08260-Seed2Scale: A Self-Evolving Data Engine with Parallel Worlds Expansion for Scalable Robot Learning
- 2602.14032-RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation
- 2602.15397-ActionCodec: What Makes for Good Action Tokenizers
- 2607.01819-Koopman operator theory: fundamentals, control, and applications