FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment

summary

Video file (mp4)

The gist

The gist FOCUS, a single-stage PPO framework, trains an actor to use privileged state information during training while automatically regulating whether it collects rollouts from RGB-D or privileged

In short

FOCUS is a single-stage PPO framework that trains an actor using privileged state information while intelligently regulating whether to collect rollouts from RGB-D or privileged state latents. It achieves this by aligning their representations and using a KL-guided switching mechanism to control the modality during training, leading to better sample efficiency and superior test performance without needing privileged states at test time.

Key concepts

KL-Guided Modality Switching
This mechanism uses cross-modal policy disagreement to decide the probability of using RGB-D rollouts. It measures how much the actor's actions change when switching from privileged state input to RGB-D input, converting this disagreement into a probability ($\lambda_r$) for collecting an RGB-D rollout. The switch is smoothed over time to ensure regulated exposure.
Representation Alignment Objective
This objective ensures that the representations of the privileged state ($z_{t,s}$) and RGB-D observations ($z_{t,o}$) are compatible during training. It minimizes policy-output KL divergence between these two modalities, forcing them to induce similar action distributions when fused into the actor's input. This alignment is crucial for making training inputs work together.
Single-Stage PPO Framework
FOCUS is structured as a single-stage Proximal Policy Optimization (PPO) framework. It trains an actor that utilizes privileged state information during training but employs a controlled switching mechanism to balance the use of this information against RGB-D data for rollout collection. This structure allows it to manage the trade-off between training efficiency and test-time requirements.

Terminology used across episodes

This episode discusses

The paper

FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment · Read on arXiv

Filip Grigorov, Kourosh Darvish, Nandita Vijaykumar

University of Toronto

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment".

Dev: The gist FOCUS, a single-stage PPO framework,

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: So we're talking about this paper called "FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment". It tackles the problem that training a robot with high-resolution RGB-D cameras is super slow because those inputs are just too noisy and have too much data to handle.

Dev: Exactly. And you get this privileged state information, which is basically what you see in simulation, which makes training way faster, but then when you take that trained robot outside to test it, it can't use that privileged state anymore because there isn't any privileged state there.

Taro: So the core problem they’re looking at is this gap between the training environment where we have this extra helpful data and the real world testing environment where we only have RGB-D. It’s how to bridge that gap so the robot learns well in training but works on its own when it only sees what's out there.

Rosa: And they propose FOCUS, which is a single-stage PPO framework designed to train the critic using that privileged state while automatically deciding whether the actor should collect rollouts from either that privileged state or the actual RGB-D data.

Dev: The big claim here is that they regulate this switching based on how much the actor's behavior changes when you swap out one input for the other, which is measured through KL divergence between those two inputs.

Taro: So it’s not just about switching randomly; they use this measure of disagreement to decide if it’s time to collect data from RGB-D or stick with the privileged state latents. That sounds like a smart way to control exposure during training.

Rosa: And they also have this representation alignment objective, which is designed to make sure that even when the actor is using different inputs, those inputs result in compatible action distributions.

Dev: They use a policy-output alignment loss, specifically KL divergence between the policy induced by the privileged state and the one induced by RGB-D observations. This loss helps keep things consistent across both modalities during training.

Taro: So it’s not just controlling *when* to use RGB-D, but also making sure that when it does use it, the resulting actions look like they belong with what the robot is already trained to do from the privileged state.

Rosa: Right. And they find that these two things—the switching mechanism and the alignment objective—work together to limit RGB-D rollouts when those distributions don't match up well, which means it only increases RGB-D exposure when it makes sense for learning.

Dev: The experimental results show this works well across five different manipulation tasks, raising average test success from zero point seven one to zero point nine three compared to a baseline that just used RGB-D at test time on those tasks.

Paper summary: Taro: And looking at the budget normalized training success AUC, it jumps from zero point four seven up to zero point six five across all the pipelines they tested, which shows a real improvement in how much learning happens for the same amount of data used in training.

Rosa: On a specific task called Pick-and-Place, they see test success go up from zero point four seven to zero point eight six and the AUC increase five times, from zero point one two to zero point six one, which is pretty significant for that kind of manipulation work.

Dev: Now looking at the ablation studies, they showed that a fixed schedule ramp without the KL feedback from the switching mechanism still helps, but it just lets the RGB-D probability increase based on a predefined warmup schedule w(r).

Taro: And when they looked at FOCUS-simple, which only kept that KL-guided switch and the image-encoder policy output KL term, it suggested that having controlled visual exposure and aligning the image encoder through policy output KL are the most reliable parts of this system.

Rosa: So, to put it plainly, FOCUS uses a single PPO setup to manage whether you collect data from privileged states or RGB-D latents during training, using disagreement as a signal for switching and alignment to keep things consistent.

Dev: It’s about solving that training versus testing gap by making sure the actor is adapted from the privileged state rollouts toward actual RGB-D control, without letting it see too much of the real world too early.

Taro: And what this means for us is that we can get better sample efficiency in simulation while still having a robot policy that actually works when it’s interacting with real visual data at test time.

Rosa: It also suggests that the way we structure the training signal—how we blend simulated knowledge with real-world sensor data—is a key factor in making these kinds of skills more robust.

Dev: The paper also points out a limitation, which is that they don't fully address how this system would perform when you move it to physical robots or more complex manipulation suites outside of simulation.

Taro: That makes sense. They focused on the latent space dynamics in simulation, and the real world adds a whole new layer of uncertainty that’s hard to model perfectly with just these switching rules.

Rosa: So, while this is a strong result for achieving high performance metrics in simulated settings with limited real-world data, it leaves open questions about its direct applicability to physical robots where sensor noise and dynamics are much messier.

Dev: The name of the paper, "FOCUS: From Privileged States to RGB-D with Controlled Modality Switching and Representation Alignment," really sums up the core idea of this paper.

Taro: It’s a controlled system that manages the flow of data based on how well the learned actions from different sources align.

Rosa: So, if you're interested in how we can build systems that learn from simulation but perform reliably in reality without needing perfect real-world data for every step, this paper is worth checking out.

Conclusion: Rosa: So we're wrapping up this look at FOCUS, which is this PPO framework that manages how the robot gets its training data from privileged states versus actual RGB-D images during learning and testing.

Dev: Yeah, Rosa, so the main point is this single-stage setup that uses a switching mechanism based on policy disagreement to decide when to use the RGB-D data instead of those simulation latents.

Taro: It means they're not just blindly collecting data from one source; they are actively regulating that exposure based on how much the robot’s behavior changes between those two inputs.

Rosa: Right, so for someone just listening, it boils down to this system making smart choices about which input to use during training so the robot can actually perform well when it only sees real-world camera feeds at test time.

Dev: And they're saying this improves sample efficiency because you get the benefit of fast simulation training while still hitting better performance metrics on the real hardware.

Taro: The authors focused on getting these two things—the switching and the representation alignment—working together to ensure that whatever input is being used, it leads to compatible actions.

Rosa: It seems like their conclusion is that this controlled approach successfully bridges that gap between where training happens and where the robot actually operates.

Dev: And they show strong results across five different manipulation tasks, with test success jumping up significantly compared to just using RGB-D at test time on its own.

Taro: But they also pointed out something important, which is that while this works well in simulation for these specific tasks, it’s still not fully tested on physical robots or much more complex scenarios outside of what they set up.

Rosa: So the big implication is that this shows a really clever way to structure training signals when you have access to high-fidelity simulation data but need it to translate into good real-world robot skills.

Dev: It suggests that how you blend simulated knowledge with real sensor data is a critical design choice for making those kinds of skills robust.

Taro: And the next thing we need to look at is whether this kind of controlled switching can handle the messy, unpredictable nature of a completely physical robot interacting with a real room.

More episodes

← Home