CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization".
Dev: Controlling an agent with vision requires separating useful task information from irrelevant background noise, and this work introduces Controllability Factorized JEPA (CF-JEPA),
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So, we're diving into the paper "CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization," which tackles a real problem in vision-based world models where background noise messes things up. I want to know if this approach is practical for real-world deployment and how long it can maintain that robustness before we have to worry about drift or something similar.
Dev: That’s a good starting point, Rosa, because the core engineering challenge with any latent space model is usually the loop rate and latency; I’m curious if this factorization introduces significant overhead in terms of computational load or predictive lag compared to standard JEPA architectures.
Taro: From an autonomy standpoint, what I'm most interested in is how this system handles unexpected world misbehavior; specifically, when the agent encounters visual distractions that aren't just random noise but actual dynamic elements it has to navigate around.
Rosa: Well, the paper introduces Controllability Factorized JEPA as a way to separate what’s important for the task from what’s just distracting background information using two distinct latent subspaces. It seems they found that by isolating the noise into one region and keeping the relevant control information in another, you get a much more stable model overall.
Dev: That separation sounds promising, but I need to understand exactly how this mathematical split translates into concrete dynamics for the control loop; we need to see how that T(s's, a) = T c(s c' s c, a) T u(s u' s u) structure actually impacts the real-time performance of the predictor.
Taro: If it successfully isolates the uncontrollable part, does that mean when something unexpected happens in the scene, like a sudden moving object, the agent doesn't lose its grasp on its immediate control objectives? I’m pushing on what this system does when things go sideways.
Rosa: The paper shows that this approach allows them to capture all that distracting information within the uncontrollable region while focusing only on the control-relevant latent information for executing the task, which is what gives it comparable performance across 2D and three dee control tasks under nominal conditions <ref:2610.00727#pg0>.
Dev: I noticed they use a specific set of loss functions to enforce this structure, like the Forward Dynamics Loss and the Adversarial Dynamics Loss; I want to know how those specific losses keep that separation intact during training without causing instability in the prediction process.
Taro: Focusing on the learning side, if we look at how they train it, are there specific scenarios where this controllability factorization helps it maintain its state when the environment starts behaving unpredictably or presents visual clutter?
Title and authors: Rosa: They introduce several specialized losses to enforce this idea, including a Forward Dynamics Loss that ensures the controllable region only influences itself, and an Inverse Dynamics Loss to focus on control ability using only the current and next controllable latent.
Dev: That inverse dynamics loss is key for setting up the structure, but what about predicting actions in that uncontrollable space? I'm concerned about action predictability there because if it predicts things there well, it might just be learning irrelevant motions.
Taro: The paper also includes an Adversarial Dynamics Loss specifically to penalize the action predictability of the uncontrollable space, which suggests they are actively trying to make sure that part of the model doesn't learn actions based on noise.
Rosa: Exactly, and they use a predictor phi to predict actions between those uncontrollable latents; it’s a way of keeping that region from becoming too active or unpredictable in ways that don't serve the task.
Dev: So, looking at the overall objective function L = L fwd + alpha L inv + beta L adv + gamma L SIGReg, I’m wondering about those specific weights; how do those parameters alpha, beta, and gamma affect the trade-off between task accuracy and robustness against visual noise?
Taro: The paper states that these weights are tuned per task, which is important because a general setting might not be optimal for every single control scenario an agent faces.
Rosa: And the results show that by tuning them appropriately—for instance, setting beta = one and gamma = zero point zero nine in some cases—they manage to maintain high success rates even when distractors are introduced, and crucially, CF-JEPA is the only model tested that does not experience latent collapse under any of the tested distractor configurations.
Dev: That lack of latent collapse is a significant finding for stability; it means we aren't dealing with those catastrophic failures where the entire representation just breaks down when presented with visual interference.
Taro: I'm really interested in their analysis of the latent space itself; does that diagnostic decoder give us any clear insight into *why* this factorization works, or is it just a mathematical trick to get away from collapse?
Rosa: The latent space analysis confirms the effectiveness of the factorization by showing that the controllable latent of CF-JEPA actually "preserves the agent," whereas baseline latents lose that information under distractors.
Dev: That preservation of agent state seems critical; if you lose track of where your arm is, no amount of visual noise filtering will help if the underlying representation is corrupted.
Title and authors: Taro: Furthermore, a qualitative analysis using a decoder shows that the controllable latent subspace reconstructs the arm well, while its uncontrollable latent subspace only blurs it, which paints a very clear picture of what's happening in those two regions.
Rosa: This separation is quantified by the participation ratio metric for tasks like Reacher: LeWM and SMWM collapse significantly under distractors with a PR of forty-eight point two, whereas CF-JEPA maintains a much higher PR at twelve point nine in nominal settings, which strongly indicates this successful separation of control-relevant information from distractors in the latent space.
Dev: That performance difference in the participation ratio is telling; it’s not just that CF-JEPA works better on one task; it’s fundamentally showing a better structure for how the model organizes its knowledge under stress.
Taro: So, when we look ahead, what does this suggest for future work in autonomous systems? Are there specific areas where this controllability concept could be applied beyond visual control?
Rosa: The implication is that we can build world models that are inherently more robust to visual disturbances by explicitly architecting the latent space to handle noise and focus on task relevance, which opens up possibilities for deploying agents in much noisier, real-world environments.
Dev: From an engineering standpoint, the immediate impact is a more reliable predictive pipeline; if the controllable part is stable, we can trust the resulting actions more in dynamic situations where latency might be tight.
Taro: I think this moves us closer to systems that can operate effectively even when their sensory input is degraded or corrupted by irrelevant background features during complex manipulation.
Rosa: To wrap up our discussion on "CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization," the authors have demonstrated a method where explicitly factorizing the latent space into controllable and uncontrollable subspaces significantly improves robustness against visual disturbances, proving that we can capture all distractor information while keeping the task-relevant information clean.
Dev: It’s a solid step for making these models more reliable for deployment in cluttered or noisy physical settings, provided we can manage the training complexity effectively.
Taro: It suggests a path toward world models that aren't just good at reconstruction but are structurally sound regarding what they choose to prioritize when faced with sensory overload.
Rosa: That’s all the time we have for this paper; it really shows how carefully structuring the latent space can lead to a model that handles real-world visual noise much better than previous methods.
The paper's summary: Rosa: So, to recap, Controllability Factorized JEPA splits the agent's latent space into two distinct zones—one for task control and one for irrelevant noise—to make world models much tougher against visual distractions.
Dev: Right, and what I find interesting from that summary is how they formalized this split by partitioning the underlying state into a controllable part s c and an uncontrollable part s u, meaning the agent's action only directly influences that controllable section.
Taro: From an autonomy angle, that structure really suggests a clear separation of concerns; if we can isolate the noise in the uncontrollable region, then our control policy doesn't have to constantly fight against visual clutter trying to figure out what’s relevant.
Rosa: Exactly, and they show this works by using a specific set of losses during training that force that separation—like making sure the controllable part only influences itself through the Forward Dynamics Loss.
Dev: That mathematical enforcement is where my engineering curiosity kicks in; I’m looking at how those loss functions interact with the dynamics to ensure we don't just get a neat theoretical split on paper, but a stable predictor in real-time.
Taro: And when we look at the results, it’s not just about nominal conditions; they test it under actual visual perturbations with random rectangles acting as distractors, and the fact that CF-JEPA is the only one they tested that doesn't suffer from latent collapse is really telling for robustness.
Rosa: That lack of collapse is significant because it means we’re dealing with a model that maintains its identity even when overwhelmed by irrelevant background features, which speaks directly to real-world deployability.
Dev: It pushes us toward thinking about hardware; if the latent representation remains stable under visual noise, it might mean we can push these models into environments with lower-quality cameras or more chaotic lighting without needing extensive retraining cycles.
Taro: I wonder how this translates to complex maneuvers; for tasks like reaching or pushing an object, does this factorization allow the agent to focus its processing power entirely on the geometry of the task rather than trying to interpret every pixel in its view?
Rosa: That's precisely what they demonstrated qualitatively; their diagnostic decoder showed that the controllable latent subspace actually reconstructs critical parts of the agent, like its arm, while ignoring or blurring out background details.
Dev: So, if we can trust that s c represents the essential task state, we can build more reliable control loops where latency doesn't cause catastrophic failure when visual input is ambiguous.
Taro: It implies that for future autonomous systems, instead of just learning a monolithic representation of the world, we should be designing architectures that inherently prioritize control-relevant information over everything else.
Rosa: That’s the big picture they are aiming for; they are showing us how to build a world model that doesn't just see the scene but understands what part of the scene matters for executing our specific mission.
The paper's improvements: Rosa: So, to wrap up our discussion on CF-JEPA's structure, it’s important to look at the specific improvements they suggest for making these world models more robust than what we had before.
Dev: Right, and beyond just the latent space split itself, what are these explicit loss functions—like Lfwd and Ladv—actually doing to keep the system from drifting or collapsing during long training runs?
Taro: I think the real improvement lies in how they’ve framed it as an exogenous block MDP; that structure means we are explicitly modeling where action inputs have their effects, which should give us better interpretability when things go wrong in complex scenarios.
Rosa: It really does, and this leads to a major implication: we’re moving toward models that can handle visual noise because they aren't just guessing; they are structurally designed to ignore irrelevant input while focusing on the control signals.
Dev: From a deployment standpoint, that means less need for constant fine-tuning when an agent encounters a new type of clutter in the field because its core mechanism for handling distraction is baked into the latent space design itself.
Taro: And this separation also helps with understanding failure modes; if something goes wrong, we can look at whether the error originated in the controllable region or if it was just noise bleeding into that part of the model.
Rosa: Exactly, and that diagnostic power is huge because it moves us past just observing performance numbers to actually understanding *why* a model succeeds or fails under stress.
Dev: I’m looking at how they tune those loss function weights—the alpha, beta, and gamma parameters—to balance task accuracy against this robustness, which is critical for setting the right trade-off for latency-sensitive control loops.
Taro: That tuning aspect is key because it shows that the architecture itself is flexible enough to be optimized for different types of visual environments, not just one specific setup.
Rosa: This opens up a lot of future work in applying this concept to other sensory modalities, not just vision, but maybe integrating tactile data into that same controllable subspace for better manipulation.
Dev: If we can successfully implement this factorization on hardware with low latency, it could drastically improve the reliability of complex robotic systems operating in unstructured environments where visual input is inherently messy.
Taro: I think the implication for autonomy is that world models become less brittle; instead of failing entirely when presented with novel visual stimuli, they adapt by isolating what they need to learn and what they can safely ignore.
Conclusion: Rosa: So, to wrap up our discussion on "CF-JEPA: Improving Robustness of JEPA World Models via Controllability Factorization," it really boils down to how this architecture fundamentally changes how we build world models by separating what's necessary from what's just visual clutter.
Dev: That’s right, and I think the most concrete implication for us engineers is that if we can stabilize that controllable subspace, we get a much more reliable predictive pipeline, which directly helps manage latency in real-time control loops.
Taro: I agree with Dev; the ability to explicitly isolate noise means our autonomy systems won't get derailed by unexpected visual elements in the field, which is huge for complex navigation tasks.
Rosa: It really is; this work suggests that we can build models that are intrinsically more robust to real-world visual noise because they prioritize task relevance over irrelevant background features.
Dev: I’m still thinking about the training side, and how those tailored loss functions help keep things stable during long training periods without introducing unwanted instability into the action prediction process.
Taro: That level of stability is what we need for any serious autonomy; if the model collapses under distraction, it's useless in a real-world setting where sensory input is never perfectly clean.
Rosa: Absolutely, and this success with visual control tasks means we can start thinking about applying this structural separation to other domains, maybe even integrating tactile information into that controllable part of the latent space later on.
Dev: That’s a big direction for future work; if we can keep the loop rate high while maintaining this factorization structure, it could make our robotic control systems much more resilient to environmental changes.
Taro: So, moving forward, it shows that designing the latent space with a clear hierarchy—controllable versus uncontrollable—is a powerful way to guide how an AI learns to understand and act in a complex world.
Rosa: And that’s the core message of this paper; by focusing on controllability factorization, we get models that are not just accurate in perfect simulations but actually hold up when deployed in messy, real-world conditions.
Dev: It’s a solid piece of research because it gives us a clear mechanism for improving robustness rather than just relying on brute-force data collection to fix problems.
Taro: I think the structural insight into the latent space itself is what makes this interesting; it tells us *how* the model is organizing its knowledge, which is a much deeper level of understanding than just seeing better performance metrics.
Morgan Byrd, Robert Wright, Sehoon Ha
cs.RO, cs.LG
Submitted: 2026-09-30
Updated: 2026-09-30
Comments: Website: https://morganbyrd03.github.io/cf-jepa/
Project page: https://worldmodels.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: Controlling an agent with vision requires separating useful task information from irrelevant background noise, and this work introduces Controllability Factorized JEPA (CF-JEPA), a world model that
Key concepts
- Controllability Factorization
- This is the core innovation where the latent space is explicitly divided into two parts: a controllable section ($s_c$) and an uncontrollable section ($s_u$). The goal is to ensure that only the controllable part affects the agent's action, while the uncontrollable part captures all irrelevant visual noise.
- Exogenous Block MDP
- The dynamics of CF-JEPA are modeled as an exogenous block Markov Decision Process. This structure formalizes how the agent's action input specifically influences only the controllable state components ($s_c$), allowing the uncontrollable components ($s_u$) to evolve autonomously, capturing background information.
- Forward Dynamics Loss (Lfwd)
- This loss function is used during training to enforce the factorization. It ensures that the dynamics of the controllable latent space only depend on itself, meaning $z_{t+1}$ is predicted based only on $z_t$ and the action $a_t$ within the controllable subspace.
- Latent Collapse
- Latent collapse occurs when a model's latent space loses its meaningful structure or representation, often due to irrelevant information overwhelming the signal. CF-JEPA is specifically designed to prevent this by isolating task information from distractors, maintaining a stable and useful latent representation.
Terminology
Summary
Controlling an agent with vision requires separating useful task information from irrelevant background noise, and this work introduces Controllability Factorized JEPA (CF-JEPA), a world model that splits its latent space into controllable and uncontrollable subspaces to improve robustness against visual distractors.
The gist
CF-JEPA is a JEPA-style world model that splits the latent space into controllable and uncontrollable subspaces, allowing it to capture all distractor information in the uncontrollable region while using control-relevant latent information for the task, resulting in comparable performance across 2D and 3D control tasks under nominal conditions and improved performance under distracted conditions where it is the only tested model that does not experience latent collapse.
How it works
The core innovation of CF-JEPA is the explicit factorization of the latent space into a control-relevant, action-conditioned region and a control-irrelevant, action-free region. This structure is formalized by partitioning the underlying state into a controllable part, denoted as s c, and an uncontrollable part, denoted as s u. The dynamics are modeled such that the agent's action input influences only the controllable section:
T(s′s, a) = Tc(s c' s c, a) Tu(s u' s u),
where q is the emission function mapping state s into an image x. This structure is described as an exogenous block MDP.
Loss Functions and Dynamics
The model training incorporates several specialized loss functions to enforce this factorization and improve robustness:
-
Forward Dynamics Loss (Lfwd): This loss is computed on a persubspace basis, ensuring the controllable region only influences the controllable region, while the uncontrollable region progresses autonomously: Lfwd = (z t+1 − zˆ t+1)2.
-
Inverse Dynamics Loss (Linv): This loss is used to set up the latent space controllability structure and focus on control ability by predicting action given only the current and next controllable latent: Linv = (a t − aˆ c t)2, where ψ(z c t, zc t+1) = ˆa c t.
-
Adversarial Dynamics Loss (Ladv): This loss penalizes the action predictability of the uncontrollable space by using a predictor ϕ to predict actions between the uncontrollable latents: Ladv = (a t − aˆ u t)2, where ϕ(z u t, z u't+1) = ˆa u t.
-
SIGReg Loss: The same SIGReg loss used in LeWM is utilized to prevent the collapse of the latent space.
The full objective function is defined as L = Lfwd + αLinv + βLadv + γLSIGReg (6), with specific weights tuned per task, such as β = 1 and γ = 0.09.
Performance and Results
The paper compares CF-JEPA against LeWM and SMWM across four visual control tasks (Reacher, PushT, OGBench Cube, TwoRoom) under nominal conditions and under visual perturbations with two randomly placed rectangles acting as distractors. Under nominal conditions (Table II), CF-JEPA ranks second on OGBench Cube and TwoRoom. Under distracted conditions (Table III), LeWM collapses on Reacher and PushT, while SMWM also collapses on Reacher. CF-JEPA is the only model tested that does not experience latent collapse under any of the tested distractor configurations, maintaining high success rates even when distractors are introduced, and in some cases performing better than under nominal conditions (e.g., PushT: 0.87 vs 0.78).
Latent Space Analysis
Analysis of the latent space confirms the effectiveness of the factorization. A diagnostic decoder shows that the controllable latent of CF-JEPA preserves the agent,
whereas baseline latents lose it under distractors. Furthermore, a qualitative analysis using a decoder reveals that CF-JEPA's controllable latent subspace reconstructs the arm
well, while its uncontrollable latent subspace only blurs it.
The participation ratio (PR) metric for the Reacher task demonstrates this separation: LeWM and SMWM collapse significantly under distractors, whereas CF-JEPA maintains a higher PR (12.9 vs 48.2 in nominal settings). This indicates that CF-JEPA successfully separates control-relevant information from distractors in the latent space.
Conclusion
CF-JEPA improves robustness to visual disturbances by explicitly factorizing the latent space into controllable and uncontrollable subspaces, allowing for focused learning on task-relevant information while ignoring irrelevant background features.
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements that can be made to existing AI systems by adopting or extending CF-JEPA:
-
Enhanced Robustness to Visual Distractors in Latent World Models (LWMs): CF-JEPA explicitly separates the latent space into a controllable subspace and an uncontrollable subspace.
-
Improved Performance Under Visual Perturbations: The system will maintain high control performance even when the environment is corrupted by irrelevant visual noise or moving background objects, unlike standard LWMs which suffer from latent collapse under such conditions.
-
Task-Specific Robustness Enhancement: The model's ability to ignore distractors is optimized for the specific control task at hand, leading to superior success rates on tasks like PushT and OGBench Cube when visual noise is present.
-
Preservation of Agent State Information: CF-JEPA ensures that the controllable latent subspace explicitly retains information critical for controlling the agent (e.g., arm configuration in Reacher), whereas baseline models lose this state under distraction.
-
Improved Latent Space Structure Analysis: The factorization provides a clearer diagnostic tool to understand which parts of the latent space encode control-relevant information versus irrelevant background features, allowing researchers to identify and isolate noise sources within the learned representations.
This improved AI system (CF-JEPA) can perform the following specific actions:
-
Control complex robotic tasks (like grasping or reaching) in real-world or simulated environments while being unaffected by random visual disturbances, such as moving objects in the background or extraneous textures on the scene.
-
Execute precise manipulation when an agent's view is partially occluded or cluttered with non-task-relevant information, maintaining goal-directed behavior where other models would fail due to representation collapse.
-
Be deployed in environments where visual input is noisy (e.g., low-quality cameras, high background clutter) without requiring extensive retraining or manual scene filtering, as the model intrinsically learns to prioritize task-relevant signals over distractors.
-
Provide a more reliable and interpretable latent representation for downstream tasks by clearly delineating the
action-conditioned
state from theaction-free
environmental context.
Sources
- Dream to Control: Learning Behaviors by Latent Imagination
- Learning Invariant Representations for Reinforcement Learning without Reconstruction
- Sensorimotor World Models: Perception for Action via Inverse Dynamics
- Mastering Atari with Discrete World Models
- Mastering Diverse Domains through World Models
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
- Learning from Reward-Free Offline Data: A Case for Planning with Latent Dynamics Models
- LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
- Behavioral Cloning from Observation
- Zero-Shot Visual Imitation
- ImDy: Human Inverse Dynamics from Imitated Observations
- Inverse Dynamics Pretraining Learns Good Representations for Multitask Imitation
- Metrics for Finite Markov Decision Processes
- Denoised MDPs: Learning World Models Better Than the World Itself
- Unsupervised Domain Adaptation by Backpropagation
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
- stable-worldmodel: A Platform for Reproducible World Modeling Research and Evaluation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving