DeepJEPA: Scaling World Models from Within
summary
The gist
World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition.
In short
DeepJEPA addresses how to scale world models by focusing computation internally rather than just rolling farther. It introduces a method where the model decides whether to perform deeper latent updates for each potential future based on its expected value to the planner. This means refinement is concentrated only on critical events, like physical contact, leading to better performance without uniformly increasing prediction depth.
Key concepts
- Budgeted Recurrent Transitions
- Each predicted transition is treated as a computation budget. The model uses a recurrent cell and a 'continue head' to decide if computing another latent update for that specific candidate is worth the cost, based on whether it improves the predicted outcome beyond a certain threshold.
- Value-of-Computation (VOC)
- This concept frames computation allocation as an economic decision. The system calculates the expected value of one more latent update by comparing it to the current state's value. Computation is only performed if this net value is positive, meaning the update is expected to benefit the planner.
- Elite-Margin Stability
- This result shows when internal updates are actually useful. If a candidate's correction isn't large enough to cross a specific 'elite margin,' the set of best candidates remains stable. This proves that refinement only matters when it changes which action is chosen by the planner.
- Interaction-Aligned Refinement
- Useful computation is not random; it aligns with important scene dynamics, such as physical contact or sustained interaction. The system learns to focus its deeper computations precisely at these decision-critical moments rather than applying uniform depth across all predictions.
Terminology used across episodes
This episode discusses
- DeepJEPA: Scaling World Models from Within · Paper Radio
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
- PonderNet: Learning to Ponder
- Revisiting Feature Prediction for Learning Visual Representations from Video
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- RT-1: Robotics Transformer for Real-World Control at Scale
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- WorldVLA: Towards Autoregressive Action World Model
- Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models
- Universal Transformers
- Deep Visual Foresight for Planning Robot Motion
- World Models for Learning Dexterous Hand-Object Interactions from Human Videos
- Adaptive Computation Time for Recurrent Neural Networks
- World Models
- Mastering Diverse Domains through World Models
- Selecting Computations: Theory and Applications
- When to Trust Your Model: Model-Based Policy Optimization
- Adapting World Models with Latent-State Dynamics Residuals
- Light-WAM: Efficient World Action Models with State-Fusion Action Decoding
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels
The paper
DeepJEPA: Scaling World Models from Within · Read on arXiv
Zijian Jin, Yunbei Zhang, Yuanzhe Liu, Ming Liu, Baian Chen, Weirui Ye
New York University · Tulane University · UIUC Department of Computer Science and Engineering, Princeton University, MIT, Columbia University, Stanford University
World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition. We show that making every transition uniformly deeper wastes computation and can degrade planning because useful refinement is concentrated at a small set of decision-critical events. We introduce DeepJEPA, a weight-tied joint-embedding predictive world model that treats transition depth as an inner test-time scaling axis and learns when another recurrent update is worth computing for each candidate and rollout step. Across five visual-control settings, DeepJEPA improves or matches the strongest fixed-depth planner while averaging only 1.00-1.26 updates per transition. Its additional computation concentrates at contact onset and sustained object interaction, where latent corrections can change which candidates enter the planner's elite set and which action is selected. Representation probes further show that improved planning does not require uniformly better object-state decodability. DeepJEPA therefore reframes world-model scaling as a problem of allocating internal computation where it can change the planner's decision: think deeper at decision-critical transitions instead of making every rollout uniformly deeper or longer.
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "DeepJEPA: Scaling World Models from Within".
Dev: World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: To recap, DeepJEPA tackles the problem of scaling world models by suggesting that uniform depth is inefficient because useful refinement is highly concentrated at decision-critical events (<ref:2610.00368#pg0>). The paper introduces a weight-tied joint-embedding predictive world model that treats transition depth as an inner testtime scaling axis, learning when another recurrent update is worth computing based on its expected value to the planner (<ref:2610.00368#pg2>). It shows that this method can improve or match the strongest fixed-depth planner across five visual-control settings by averaging only one point zero zero to one point two six updates per transition (<ref:2610.00368#pg2>).
Dev: I see what you mean regarding the efficiency gain from that averaging, Rosa, but I still need more clarity on how this predictive world model actually works compared to other latent planning methods like V-JEPA or LeWM mentioned in the introduction (<ref:2610.00368#pg1>). What's the core mechanism of this joint-embedding self-supervision that allows it to turn prediction in representation space into a route toward world models?
Taro: The paper focuses on action-conditioned JEPAs, which takes the principle of learning representations without reconstructing pixels and turns it into a world model by having a planner roll candidate actions through latent space and select those whose predicted terminal state matches the visual goal (<ref:2610.00368#pg1>). That turns the JEPA principle into something that can drive model-predictive control with CEM (<ref:2610.00368#pg1>).
Rosa: So, it’s taking those latent representations and using them to drive planning through a predictive loop, but DeepJEPA adds this layer of dynamic computation management on top of the transition model itself (<ref:2610.00368#pg2>). It's about managing the internal computation rather than just scaling up the search space externally.
Dev: That internal management is key for loop rate stability, but I still wonder how this ties into real-time constraints on hardware. If we are running a high-frequency control loop, the decision to compute another recurrent update based on a marginal gain threshold η has to be extremely fast and low latency, otherwise we introduce unacceptable lag into our feedback cycle.
Taro: And that brings up the misbehavior aspect again. When things go wrong dynamically, does this budgeted approach offer a mechanism for adaptation beyond just refining contact events? We need assurance that if the world violates expectations in an unforeseen way, the system has a way to allocate compute effectively instead of stalling or making poor decisions.
Rosa: The paper suggests that useful computation concentrates around physical interaction and acts through planner rankings rather than uniform gains in state decodability (<ref:2610.00368#pg2>). This means the system is prioritizing what directly impacts the planner's ranking of actions, which sounds like a targeted way to handle dynamic situations.
Dev: Targeted refinement is good for efficiency, but my concern remains about deployment fidelity. If this system is trained on certain contact dynamics, will it generalize its decision-critical focus reliably when deployed in a completely different physical environment where the contact physics are different?
Taro: That generalization depends on whether the underlying joint-embedding representations capture enough fundamental dynamics to recognize *any* type of interaction as potentially critical, even if it’s not exactly what was seen during training.
Rosa: It seems DeepJEPA is proposing a way to make world models more computationally frugal while maintaining or improving planning quality by being selective about where they spend their compute budget (<ref:2610.00368#pg2>). This selective approach is what makes it interesting for real-world application, even if the generalization still needs rigorous testing.
Conclusion: Dev: Looking at the conclusion of DeepJEPA: Scaling World Models from Within, the authors are essentially saying that we should stop thinking about scaling by just increasing trajectory counts or extending horizons because that wastes computation (<ref:2610.00368#pg1>). Instead, the focus should be on how to allocate computation intelligently across internal transition depths based on its expected value to the planner (<ref:2610.00368#pg2>).
Rosa: I agree with that sentiment; it shifts the design principle from brute-force scaling outward to a more nuanced internal allocation strategy (<ref:2610.00368#pg2>). The title itself, DeepJEPA: Scaling World Models from Within, perfectly captures this idea of controlling the internal process rather than just pushing the boundaries of what we can imagine externally.
Taro: From an autonomy perspective, if we accept that useful refinement is concentrated at decision-critical events like contact onset (<ref:2610.00368#pg0>), it implies that a truly autonomous system doesn't need a perfect, uniformly deep understanding of every single pixel state to make good decisions (<ref:2610.00368#pg2>). It suggests focusing on the moments that directly influence the action selection process.
Dev: That makes sense for latency management because we aren't trying to compute everything at maximum detail simultaneously; we are computing what is valuable right now (<ref:2610.00368#pg2>). But I still have my concerns about deployment fidelity—how reliably that mechanism works when the environment throws a completely novel dynamic at it.
Rosa: The paper’s main contribution seems to be identifying this internal depth axis as a distinct scaling variable and framing its allocation as a value-of-computation problem (<ref:2610.00368#pg2>). It’s less about achieving the most complex model possible and more about achieving the most efficient planning capability.
Taro: So, the implication for future work might be developing better criteria for that continue head mechanism—finding a way to make it robust enough to detect dynamic shifts beyond just physical contact, ensuring it handles misbehavior gracefully (<ref:2610.00368#pg2>).
Dev: I’m hoping future work will also address the computational cost of running this allocation logic itself, because if the metareasoning process becomes too heavy for real-time hardware, the entire efficiency gain vanishes (<ref:2610.00368#pg2>).
Rosa: Exactly. The whole point of DeepJEPA is to show that we can get better planning performance by being selective about where we spend our compute budget, which points toward a more efficient architecture for future world models (<ref:2610.00368#pg2>).
More episodes
- 2610.10846-Cross-Embodiment Robot Foundation World Models with Latent Actions
- 2610.10601-Teaching a Robot Dog New Tricks: Diverse Quadruped Skills via Combined Reinforcement and Imitation Learning with Adversarial Task Selection
- 2610.10637-TacHair: Tactile Contact-Distribution Guided Online Correction for Robotic Hair Stroking and Perception
- 2610.10646-Masked Generative Motion Planning with Geometry-Guided Token Search
- 2610.10812-Skill-SLM: Agent Skill-driven Small Language Models for Reliable Robot Operation
- 2610.10801-Same Action, Different Outcome: Variability in Dynamic Cloth Manipulation
- 2610.10810-Diagnosing and Recovering from Observation-Space Shift at Long-Horizon Skill Seams
- 2610.10748-TAPNAV: Humanoid Navigation through Tactile Active Perception
- 2610.10855-OmniHOI: Dexterous Hand-Object Interaction from Monocular Human Video
- 2610.11003-ActiveReg: Information-Driven Active Regional Probing for Partial-to-Full Bone Registration