FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment
summary
The gist
The gist FOCUS, a single-stage PPO framework, trains an actor to use privileged state information during training while automatically regulating whether it collects rollouts from RGB-D or privileged
In short
FOCUS is a single-stage PPO framework that trains an actor using privileged state information while intelligently regulating whether to collect rollouts from RGB-D or privileged state latents. It achieves this by aligning their representations and using a KL-guided switching mechanism to control the modality during training, leading to better sample efficiency and superior test performance without needing privileged states at test time.
Key concepts
- KL-Guided Modality Switching
- This mechanism uses cross-modal policy disagreement to decide the probability of using RGB-D rollouts. It measures how much the actor's actions change when switching from privileged state input to RGB-D input, converting this disagreement into a probability ($\lambda_r$) for collecting an RGB-D rollout. The switch is smoothed over time to ensure regulated exposure.
- Representation Alignment Objective
- This objective ensures that the representations of the privileged state ($z_{t,s}$) and RGB-D observations ($z_{t,o}$) are compatible during training. It minimizes policy-output KL divergence between these two modalities, forcing them to induce similar action distributions when fused into the actor's input. This alignment is crucial for making training inputs work together.
- Single-Stage PPO Framework
- FOCUS is structured as a single-stage Proximal Policy Optimization (PPO) framework. It trains an actor that utilizes privileged state information during training but employs a controlled switching mechanism to balance the use of this information against RGB-D data for rollout collection. This structure allows it to manage the trade-off between training efficiency and test-time requirements.
Terminology used across episodes
This episode discusses
- FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment · Paper Radio
- Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning
- End-to-End Training of Deep Visuomotor Policies
- Provable Partially Observable Reinforcement Learning with Privileged Information
- Asymmetric Actor Critic for Image-Based Robot Learning
- Robust Asymmetric Learning in POMDPs
- Distilling the Knowledge in a Neural Network
- Actor-Mimic: Deep Multitask and Transfer Reinforcement Learning
- TWIST: Teacher-Student World Model Distillation for Efficient Sim-to-Real Transfer
- DextrAH-G: Pixels-to-Action Dexterous Arm-Hand Grasping with Geometric Fabrics
- DextrAH-RGB: Visuomotor Policies to Grasp Anything with Dexterous Hands
- Privileged Sensing Scaffolds Reinforcement Learning
- CURL: Contrastive Unsupervised Representations for Reinforcement Learning
- Reinforcement Learning with Augmented Data
- Proximal Policy Optimization Algorithms
- Learn to Teach: Sample-Efficient Privileged Learning for Humanoid Locomotion over Diverse Terrains
- AACC: Asymmetric Actor-Critic in Contextual Reinforcement Learning
- A Theoretical Justification for Asymmetric Actor-Critic Algorithms
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable Model
- Gated Multimodal Units for Information Fusion
- Attention Is All You Need
The paper
FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment · Read on arXiv
Filip Grigorov, Kourosh Darvish, Nandita Vijaykumar
University of Toronto
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment".
Dev: The gist FOCUS, a single-stage PPO framework,
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So we're talking about this paper called "FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment". It tackles the problem that training a robot with high-resolution RGB-D cameras is super slow because those inputs are just too noisy and have too much data to handle.
Dev: Exactly. And you get this privileged state information, which is basically what you see in simulation, which makes training way faster, but then when you take that trained robot outside to test it, it can't use that privileged state anymore because there isn't any privileged state there.
Taro: So the core problem they’re looking at is this gap between the training environment where we have this extra helpful data and the real world testing environment where we only have RGB-D. It’s how to bridge that gap so the robot learns well in training but works on its own when it only sees what's out there.
Rosa: And they propose FOCUS, which is a single-stage PPO framework designed to train the critic using that privileged state while automatically deciding whether the actor should collect rollouts from either that privileged state or the actual RGB-D data.
Dev: The big claim here is that they regulate this switching based on how much the actor's behavior changes when you swap out one input for the other, which is measured through KL divergence between those two inputs.
Taro: So it’s not just about switching randomly; they use this measure of disagreement to decide if it’s time to collect data from RGB-D or stick with the privileged state latents. That sounds like a smart way to control exposure during training.
Rosa: And they also have this representation alignment objective, which is designed to make sure that even when the actor is using different inputs, those inputs result in compatible action distributions.
Dev: They use a policy-output alignment loss, specifically KL divergence between the policy induced by the privileged state and the one induced by RGB-D observations. This loss helps keep things consistent across both modalities during training.
Taro: So it’s not just controlling *when* to use RGB-D, but also making sure that when it does use it, the resulting actions look like they belong with what the robot is already trained to do from the privileged state.
Rosa: Right. And they find that these two things—the switching mechanism and the alignment objective—work together to limit RGB-D rollouts when those distributions don't match up well, which means it only increases RGB-D exposure when it makes sense for learning.
Dev: The experimental results show this works well across five different manipulation tasks, raising average test success from zero point seven one to zero point nine three compared to a baseline that just used RGB-D at test time on those tasks.
Paper summary: Taro: And looking at the budget normalized training success AUC, it jumps from zero point four seven up to zero point six five across all the pipelines they tested, which shows a real improvement in how much learning happens for the same amount of data used in training.
Rosa: On a specific task called Pick-and-Place, they see test success go up from zero point four seven to zero point eight six and the AUC increase five times, from zero point one two to zero point six one, which is pretty significant for that kind of manipulation work.
Dev: Now looking at the ablation studies, they showed that a fixed schedule ramp without the KL feedback from the switching mechanism still helps, but it just lets the RGB-D probability increase based on a predefined warmup schedule w(r).
Taro: And when they looked at FOCUS-simple, which only kept that KL-guided switch and the image-encoder policy output KL term, it suggested that having controlled visual exposure and aligning the image encoder through policy output KL are the most reliable parts of this system.
Rosa: So, to put it plainly, FOCUS uses a single PPO setup to manage whether you collect data from privileged states or RGB-D latents during training, using disagreement as a signal for switching and alignment to keep things consistent.
Dev: It’s about solving that training versus testing gap by making sure the actor is adapted from the privileged state rollouts toward actual RGB-D control, without letting it see too much of the real world too early.
Taro: And what this means for us is that we can get better sample efficiency in simulation while still having a robot policy that actually works when it’s interacting with real visual data at test time.
Rosa: It also suggests that the way we structure the training signal—how we blend simulated knowledge with real-world sensor data—is a key factor in making these kinds of skills more robust.
Dev: The paper also points out a limitation, which is that they don't fully address how this system would perform when you move it to physical robots or more complex manipulation suites outside of simulation.
Taro: That makes sense. They focused on the latent space dynamics in simulation, and the real world adds a whole new layer of uncertainty that’s hard to model perfectly with just these switching rules.
Rosa: So, while this is a strong result for achieving high performance metrics in simulated settings with limited real-world data, it leaves open questions about its direct applicability to physical robots where sensor noise and dynamics are much messier.
Dev: The name of the paper, "FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment," really sums up the core idea of this paper.
Taro: It’s a controlled system that manages the flow of data based on how well the learned actions from different sources align.
Rosa: So, if you're interested in how we can build systems that learn from simulation but perform reliably in reality without needing perfect real-world data for every step, this paper is worth checking out.
Conclusion: Rosa: So we're wrapping up this look at FOCUS, which is this PPO framework that manages how the robot gets its training data from privileged states versus actual RGB-D images during learning and testing.
Dev: Yeah, Rosa, so the main point is this single-stage setup that uses a switching mechanism based on policy disagreement to decide when to use the RGB-D data instead of those simulation latents.
Taro: It means they're not just blindly collecting data from one source; they are actively regulating that exposure based on how much the robot’s behavior changes between those two inputs.
Rosa: Right, so for someone just listening, it boils down to this system making smart choices about which input to use during training so the robot can actually perform well when it only sees real-world camera feeds at test time.
Dev: And they're saying this improves sample efficiency because you get the benefit of fast simulation training while still hitting better performance metrics on the real hardware.
Taro: The authors focused on getting these two things—the switching and the representation alignment—working together to ensure that whatever input is being used, it leads to compatible actions.
Rosa: It seems like their conclusion is that this controlled approach successfully bridges that gap between where training happens and where the robot actually operates.
Dev: And they show strong results across five different manipulation tasks, with test success jumping up significantly compared to just using RGB-D at test time on its own.
Taro: But they also pointed out something important, which is that while this works well in simulation for these specific tasks, it’s still not fully tested on physical robots or much more complex scenarios outside of what they set up.
Rosa: So the big implication is that this shows a really clever way to structure training signals when you have access to high-fidelity simulation data but need it to translate into good real-world robot skills.
Dev: It suggests that how you blend simulated knowledge with real sensor data is a critical design choice for making those kinds of skills robust.
Taro: And the next thing we need to look at is whether this kind of controlled switching can handle the messy, unpredictable nature of a completely physical robot interacting with a real room.
More episodes
- 2610.10846-Cross-Embodiment Robot Foundation World Models with Latent Actions
- 2610.10601-Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection
- 2610.10637-TacHair: Tactile Contact-Distribution Guided Online Correction for Robotic Hair Stroking and Perception
- 2610.10646-Masked Generative Motion Planning with Geometry-Guided Token Search
- 2610.10812-Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation
- 2610.10801-Same Action, Different Outcome: Variability in Dynamic Cloth Manipulation
- 2610.10810-Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams
- 2610.10748-TAPNAV: Humanoid Navigation through Tactile Active Perception
- 2610.10855-OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video
- 2610.11003-ActiveReg: Information-Driven Active Regional Probing for Partial-to-Full Bone Registration