Video Replanning via Latent Embedding Refinement and Rejection-Based Sampling
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Video Replanning via Latent Embedding Refinement and Rejection-Based Sampling".
Dev: Video-based representations have gained prominence in planning and decision-making due to their ability to encode rich spatiotemporal dynamics and geometric relationships,
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So we’re diving into "Video Replanning via Latent Embedding Refinement and Rejection-Based Sampling" today. This paper tackles the problem of video planning systems that struggle when they encounter unexpected failures during actual interaction in a real environment. It introduces a way to integrate that interaction data directly into the planning process by updating model parameters online and filtering out plans that have already failed before generating new ones.
Dev: That sounds really interesting from an engineering standpoint, Rosa; dealing with online parameter updates and plan rejection means we have to be mindful of the loop rate and any latency introduced by those refinement steps. How does this approach handle the uncertainty about those unknown system parameters?
Taro: It tackles that uncertainty head-on by aiming for implicit state estimation without needing to explicitly model every unknown variable, which is a big shift from traditional methods, Dev.
Rosa: Exactly; they are trying to use the interaction data itself to refine an internal latent embedding that captures those hidden parameters like mass or friction. This allows the system to adapt dynamically as it learns more about the environment through trial and error.
Dev: If the system is updating parameters online, what’s the mechanism for deciding which parameter subset to optimize at any given moment? We need a clear process for how it decides what's unknown versus what needs explicit modeling.
Taro: The paper suggests that instead of assuming you have access to the parameterization of theta, they assume that these underlying parameters can be inferred directly from the interaction data itself, which is a key part of their problem formulation.
Rosa: That inference happens through two main components: a retrieval module and a refining state embedding step. The retrieval module looks at past videos to get state embeddings for objects and then samples the most relevant one based on how close the current interaction features are to those stored embeddings.
Dev: Sampling from past data sounds computationally intensive; what’s the trade-off there between getting a very accurate state embedding and keeping the inference time low enough for real-time planning?
Title and authors: Taro: They use this retrieval mechanism to pull in prior experience, but then they follow up with an optimization step where an identification module generates all possible outcomes, including unsuccessful ones.
Rosa: And during replanning, they freeze those identification module parameters and optimize the state embedding to minimize the discrepancy between its generated rollouts and what actually happened in the current interaction using a denoising diffusion loss.
Dev: Optimizing against observed data through that loss function sounds like it could introduce some instability if the loss landscape is too rough, Rosa; we need to make sure that doesn't lead to catastrophic failure modes during execution.
Taro: The rejection mechanism really helps here because instead of just picking one plan from the generated candidates, they select the plan that is most different from past failures in a data buffer. This actively steers the system away from repeating strategies it already knows don't work.
Rosa: That rejection strategy is quite clever; by selecting plans with the largest distance to past failures, they are encouraging novel actions rather than just re-running old attempts, which should lead to more exploration during planning.
Dev: So you’re saying the rejection module isn’t just a filter but an active driver for finding new strategies? That implies we’re not just relying on the generator to be good enough on its own, right?
Taro: Precisely; it forces the system to consider possibilities that have proven unsuccessful before, which is crucial when you're operating in a partially observable setting where things don't always go according to plan.
Rosa: Moving toward the conclusion of this discussion on "Video Replanning via Latent Embedding Refinement and Rejection-Based Sampling," we’ve seen how they use interaction data to implicitly estimate parameters, adapt online, and use rejection to guide exploration. This whole framework aims for a more robust way for video planning systems to handle real-world unpredictability.
Dev: From an engineering standpoint, the success of this hinges on how quickly that refinement step converges; if the optimization takes too long, we’re back to high latency issues that we saw with things like ProbeFlow when dealing with iterative ODE solving.
Title and authors: Taro: I think it shows a solid direction because it moves away from just learning policies from state-action pairs and instead focuses on extracting representations of interaction dynamics directly from videos, including the failed trials.
Rosa: And the implications are that we might see systems that don't need a massive amount of pre-collected simulation data to get good results; they can learn a lot just by observing interactions in the real world.
Dev: If this works outside the lab for extended periods, I’d want to know how stable those learned embeddings are when the visual context changes drastically, because that’s where my concerns about loop rate and failure modes really kick in.
Taro: The paper also shows how these state embeddings can generalize across different object appearances and even long-horizon, multi-mode environments without needing retraining for every new scenario.
Rosa: That generalization is compelling because it suggests we could have much more versatile robotic systems that aren't brittle when they encounter objects or situations slightly outside their initial training set.
Dev: So, while the performance metrics on the simulated task set are strong—they improved replanning performance compared to baselines like AVDC—the authors did admit a limitation regarding execution noise in real-world scenarios.
Taro: They specifically stated that they attribute failures solely to planning errors, meaning if there’s actual messy physical interaction noise not captured by the model, the system might still struggle.
Rosa: That’s a fair point; we have to remember that this framework focuses heavily on the planning side and doesn't account for every single bit of real-world execution noise.
Dev: It sounds like it’s a very strong step forward in handling uncertainty during the planning phase, but it’s not yet ready to handle the full complexity of physical execution noise in a purely autonomous setting without further refinement.
Taro: Still, I think the concept of using past failed interactions to inform *future* successful plans through rejection is a powerful idea for autonomy research.
Rosa: Absolutely; the way they integrate retrieval and rejection into this video planning pipeline offers a new way to handle uncertainty in these complex spatiotemporal tasks. That wraps up our discussion on "Video Replanning via Latent Embedding Refinement and Rejection-Based Sampling."
The paper's summary: Rosa: So, to recap, this paper is all about taking interaction videos—the messy data from real experiments—and using them to dynamically update an internal representation of the environment, which then lets the AI plan better by rejecting bad ideas in real-time.
Dev: That’s a pretty concise way to put it; essentially, they are building a closed loop where observation directly informs planning and parameter estimation without needing a separate, explicit model for every physical variable. I'm still wondering about the practical implications of that online parameter updating you mentioned earlier; how stable is that system when the physical dynamics shift suddenly?
Taro: The core idea is implicit state estimation, meaning the AI infers things like friction or mass just by watching videos succeed and fail, which is huge because we don't always have perfect physics models for our robots. If the system can learn those dynamics online, it means a robot could adapt to a slightly slippery floor or a heavier object without needing an engineer to manually tweak its parameters beforehand.
Rosa: That’s exactly what excites me; imagine a field robot trying to navigate uneven terrain where the ground friction changes constantly, and this system just adjusts its plan on the fly because it remembers how those interactions felt. It moves beyond fixed models entirely.
Dev: From an engineering standpoint, that online refinement process sounds like it could introduce significant jitter into our control loop if the optimization step isn't fast enough; we’re dealing with one hundred twenty-eight times one hundred twenty-eight video data and trying to make real-time decisions. What’s the actual latency profile they report for those crucial refinement steps?
Taro: They are focusing on making sure that when things go wrong, the rejection module kicks in fast enough to stop repeating those failures before they derail a whole mission. By rejecting plans that look like past mistakes, it keeps the exploration focused on genuinely novel strategies.
Rosa: And I’m really looking forward to seeing this outside of a controlled lab setting. Can we expect this framework to be robust enough for long-horizon tasks, like navigating a complex warehouse or performing a sequence of intricate manipulation steps?
Dev: That’s the real test, Rosa; I worry about generalization. If the system learns the dynamics for one specific object interaction—say, grasping a glass—will it seamlessly transition to grasping a completely different shape without needing retraining on that new object?
Taro: The paper suggests that by using learned state embeddings from diverse interactions, the system gains this kind of generalization across different objects and even long sequences of movements. It’s not just memorizing one path; it’s understanding the underlying physics enough to adapt.
Rosa: That would be incredible for widespread robotic deployment; systems that aren't brittle when faced with slightly new conditions are what we need in the field. We could see this applied to everything from complex search and rescue scenarios to general domestic tasks, provided we can nail the real-world execution noise issue they mentioned.
Dev: I agree that adaptability is key, but I still want to know how resilient it is when things get truly unexpected—like a sudden change in lighting or unexpected external forces that aren't part of the modeled interaction dynamics. That’s where my concern about failure modes really ramps up.
Taro: The authors acknowledge that they are primarily attributing failures to planning errors, which hints at where the current boundary is; if the physical execution itself introduces noise that isn't captured by their model parameters, the planning might still fail even with perfect state estimation.
Rosa: So, while it’s a huge step toward smarter planning driven by interaction data, we still have to tackle that messy reality of physical execution noise and how long this system can reliably operate outside of a perfectly controlled environment. We've got some serious potential here for making robots truly autonomous.
The paper's improvements: Rosa: So, to wrap up on the methodology, these authors propose a few key enhancements that really beef up this video planning framework by making the estimation process more robust and the plan generation smarter.
Dev: I’m listening; what exactly are those improvements you’re referring to? Are we talking about fixing the latency issues we saw with ProbeFlow, or something deeper into how they handle uncertainty?
Taro: They're focused on making that implicit state estimation more reliable by refining the embedding against observed data and then using a rejection strategy that actively seeks out plans different from past failures. It’s like training the AI to be skeptical of its own history when generating new routes.
Rosa: That active rejection mechanism is really clever; it ensures the system doesn't just fall back on old, known bad habits when things get tricky in a novel situation. It keeps the planning process genuinely exploratory rather than repetitive.
Dev: If that rejection module is working well, it means the AI isn't stuck in local minima where it just keeps trying the same sequence of actions that didn't work before; that’s a big win for stability during replanning, even if the refinement step itself takes a moment longer.
Taro: Exactly; that prevents getting trapped in those suboptimal loops, which is critical when we think about autonomy in unpredictable real-world environments where the rules aren't perfectly defined yet.
Rosa: From a broader perspective, these improvements suggest that video-based planning can become much more resilient to the inevitable messiness of physical interaction compared to methods that rely on a single, static plan.
Dev: I’m still curious about the trade-off between refinement accuracy and computational cost; how do they balance getting that highly optimized state embedding against maintaining a fast enough loop rate for responsive control?
Taro: The authors argue that the benefits of having a physically plausible plan outweigh the computational overhead, especially since they are using prior interaction videos for retrieval to get started, which prunes the search space significantly.
Rosa: That makes sense; if you can skip generating millions of plans by sampling relevant past experiences first, then spending time optimizing only the most promising candidates with that refinement loss function is a very smart way to manage resources.
Dev: It sounds like they’re building a sophisticated filter that combines historical knowledge with immediate sensory input to produce a high-quality plan in real-time, which is what we need for reliable control systems.
Taro: That combination of retrieval, refinement, and rejection gives the system a layer of self-correction that mimics how an experienced human operator might adjust their strategy mid-task when things don't go as expected.
Rosa: It really does give me hope for field robotics; if this can handle those dynamics implicitly, we could see robots operating in much more complex and less predictable environments than what we can currently simulate or even test in a lab.
Conclusion: Rosa: So, we’re wrapping up our chat on "Video Replanning via Latent Embedding Refinement and Rejection-Based Sampling," which essentially shows how using interaction videos to update an internal state embedding and reject bad plans can make video planning much more adaptive in the real world.
Dev: It’s a really neat way to handle the uncertainty that comes with real-time physical interaction; I'm still focused on those loop rate concerns, though I see this framework promises better stability during replanning.
Taro: For me, the big implication is that autonomy can become much more flexible because the AI isn't locked into a rigid plan from the start; it’s actively learning and pruning its own mistakes based on what it sees.
Rosa: And I think that flexibility is exactly what we need for field robotics; if this works reliably outside of a perfect simulation, we could see robots handling much more unpredictable physical tasks.
Dev: If the authors can keep those refinement steps fast enough, it solves a major headache in control systems—getting responsive action without introducing too much lag.
Taro: I'm still wondering about the long-term scalability; if this method holds up across different object types and long sequences of actions, we could see a real step toward general-purpose embodied AI.
Rosa: That generalization is what makes me really hopeful; being able to handle new objects without retraining is the kind of capability that moves robots from niche tasks into everyday utility.
Dev: I'm just waiting for the next paper to show us how they tackle the execution noise issue in a way that's more robust than just attribute failures solely to planning errors.
Taro: That’s definitely the next frontier; if they can integrate physical sensing data more deeply, it would make this framework truly bulletproof for messy real-world deployment.
Rosa: Well, that brings us to the end of our discussion on "Video Replanning via Latent Embedding Refinement and Rejection-Based Sampling." It’s a fascinating piece of research showing how we can use interaction data to build smarter, more adaptable planning systems.
Dev: I think it sets a very high bar for what we expect from video-based planning in the next few years, forcing us to demand better temporal consistency and faster inference times.
Taro: We should keep an eye on how this idea of implicit state estimation evolves; it has huge potential for making autonomous systems genuinely learn how to handle the world without explicit instruction sets.
National Taiwan University · MIT CSAIL (Computer Science and Artificial Intelligence Laboratory) · Cornell University · Harvard University
cs.RO
Submitted: 2025-10-20
Updated: 2026-10-05
Comments: 20 pages. Substantial revision with a new title, updated author list and order, expanded simulation and real-world experiments, and additional analyses. Previously titled "Implicit State Estimation via Video Replanning"; the method is now called RELIC (formerly ISE)
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 84/100
The gist: Video-based representations have gained prominence in planning and decision-making due to their ability to encode rich spatiotemporal dynamics and geometric relationships, but existing frameworks
Key concepts
- Implicit State Estimation (ISE)
- This technique estimates unknown system parameters, like mass or friction, without creating explicit mathematical models for them. It achieves this by using interaction data to refine a latent embedding that implicitly captures these hidden dynamics. The system adapts dynamically during planning by adjusting these embedded parameters based on real-time observations.
- Retrieval Module
- This module extracts state embeddings from past interactions by encoding previous videos. At inference time, it compares the current interaction video's features against stored embeddings from a dataset. A softmax function is used to sample the most relevant historical state embedding, effectively retrieving a representative hidden state for the current situation.
- Rejection-Based Replanning
- Instead of generating one plan, this module creates multiple candidate plans based on different sampled state embeddings. The rejection module then selects the candidate plan that is most dissimilar from previously failed plans. This strategy encourages the system to explore novel strategies rather than repeating unsuccessful approaches.
- Video Plan Generation
- This component uses a video diffusion model conditioned on the current observation and a state embedding to create future video plans. The state embedding is projected into the CLIP-Text token space, allowing the generator to produce scenario-specific rollouts tailored to hypothesized environment configurations encoded in that embedding.
Terminology
Summary
Video-based representations have gained prominence in planning and decision-making due to their ability to encode rich spatiotemporal dynamics and geometric relationships, but existing frameworks struggle with adaptation at interaction time due to their inability to reason about uncertainties in partially observed environments. This paper introduces a novel framework that integrates interaction-time data into the planning process by updating model parameters online and filtering out previously failed plans during generation, enabling implicit state estimation (ISE) without explicitly modeling unknown state variables.
The gist
Implicit state estimation allows the system to adapt dynamically without explicitly modeling unknown state variables by integrating interaction-time data into the planning process through online model parameter updates and plan rejection.
How it works
The core of the framework involves a pipeline that leverages prior interaction videos to refine an internal latent embedding, which implicitly captures hidden parameters like mass or friction. This is achieved through two main components:
-
The Retrieval Module: This module extracts state embeddings from prior interactions by encoding past videos and selecting a representative embedding for each object ID. At inference time, it computes distances between the current interaction video's features and stored embeddings from the dataset D, using a softmax function to sample the most relevant state embedding according to a probability distribution.
-
The Refining State Embedding: To improve alignment with the current interaction, an optimization-based refinement step is introduced. An Identification Module is trained with parameters shared with the Video Plan Generator to generate all possible outcomes (successful and unsuccessful). During replanning, this module's parameters are frozen, and the state embedding is optimized to
minimize the discrepancy between its generated rollouts and the observed interaction
using a standard denoising diffusion loss.
Video Plan Generation
The Video Plan Generator uses a video diffusion model conditioned on the current observation and a state embedding. To adapt this to unknown dynamics, the system projects the state embedding into the token space of CLIP-Text and treats it as an additional language token appended to the 3D U-Net backbone. This allows for scenario-specific rollouts
that are conditioned on hypothesized environment configurations encoded in the embedding. The generator is trained end-to-end with a denoising diffusion objective, generating a video plan of 7 future frames at 128x128 resolution, which are then concatenated with the first frame to form an 8-frame video plan.
Rejection-Based Replanning
To ensure robustness and encourage exploration, the framework employs a rejection module. Instead of generating a single rollout, the Video Plan Generator produces N candidate plans
conditioned on N independently sampled state embeddings from the Retrieval Module. The Rejection Module then selects the plan with the largest distance to past failures in a data buffer F:
The Rejection Module then selects the plan most different from past failures.
This procedure encourages novel strategies rather than repeating unsuccessful ones, as it computes a distance function (e.g., L2 distance on raw pixel space) between candidate plans and failed plans to select the one with the largest discrepancy.
Action Module
The final step converts the selected video plan into executable low-level actions. This is achieved by following the concept of dense object- or wrist-tracking, where predicted video frames are used to track either the object of interest or, in tasks like bar picking, the robot’s wrist when object tracking is insufficient. The resulting trajectory is then translated into control commands using simple heuristic policies such as grasp strategies or trajectory-following controllers, enabling a training-free approach for plan execution.
Evaluation and Results
The framework was evaluated on the Meta-World System Identification Suite, which tests online adaptation under unknown system parameters (e.g., mass, friction). Experiments show that the method significantly reduces replanning failures and generates more accurate video plans compared to baselines like AVDC. The performance scaling with data size is also demonstrated; in limited data regimes (28%), the proposed method surpasses baselines in 4 out of 5 tasks, achieving the best overall performance by a substantial margin. Furthermore, state embeddings constructed from interaction videos prove robust across unseen object appearances and long-horizon, multi-mode environments. The theoretical analysis shows that the rejection strategy strictly amplifies the probability of selecting the correct interaction mode
by eliminating inconsistent candidates. Finally, a comparison against latent-context adaptation methods confirms that the integrated formulation of planning, retrieval, and adaptation consistently outperforms prior latent-state estimation approaches.
Limitations
The work assumes a reliable action module and attributes failures solely to planning errors, omitting real-world execution noise. Additionally, a first-frame bias
in the video generator is observed under low-data regimes, which can be mitigated by increasing candidate plans or injecting Gaussian noise into the first frames. Future work could address visually ambiguous tasks by incorporating tactile or proprioceptive sensing.
Improvements for AI systems
Here are the specific improvements to AI systems derived from this research, categorized by capability:
The core improvement lies in developing an adaptive, uncertainty-aware video planning system that learns environmental dynamics implicitly through interaction history rather than relying on explicit model parameterization or external belief models.
Specifically, the improved AI system can perform the following:
-
[Implicit State Estimation via Video Replanning]
-
[Dynamic Adaptation to Unknown System Parameters]
-
[Robust Replanning in Partially Observable Environments]
-
[Improved Performance on Complex, Long-Horizon Tasks]
The specific improvements and what the improved AI system can do are:
-
[Implicit State Estimation via Video Replanning]: The system can infer hidden physical parameters (like mass, friction coefficients, or object centers of mass) directly from interaction videos (successful and failed trials) without requiring prior knowledge of these parameters.
-
[Dynamic Adaptation to Unknown System Parameters]: The AI can rapidly adapt its planning strategy online based on new interaction data by continuously updating an internal latent state embedding. This allows the system to adjust its understanding of the environment's dynamics as it gathers experience, overcoming limitations where existing planners fail due to lack of explicit parameter models.
-
[Robust Replanning in Partially Observable Environments]: The system maintains a buffer of failed plans and uses a rejection mechanism to actively avoid repeating previously unsuccessful strategies during replanning. This allows the system to explore novel actions and modes (e.g., choosing between
push
orpull
for a door) when faced with uncertainty, significantly reducing replanning failures in unstructured settings. -
[Improved Performance on Complex, Long-Horizon Tasks]: By combining retrieval (sampling from past successful interactions) and refinement (optimizing the state embedding against the current interaction), the system generates video plans that are physically more plausible and contextually accurate than those generated by methods relying only on first frames or fixed task tokens. This leads to higher success rates in tasks with complex dynamics, such as those requiring sequential maneuvers (e.g., opening a box with two modes: lift or slide) or long planning horizons (e.g., the triple-faucet task).
-
[Generalization Across Object Variations]: The system can generalize its learned interaction strategies to new objects by leveraging stored state embeddings associated with previously seen objects, even if the visual appearance (color, texture) is different. This allows the AI to correctly infer the interaction mode for a
new
object based purely on its underlying physical dynamics.
Abstract
Video planning has emerged as a flexible framework for robot manipulation, in which a generative model predicts a video of task completion, and a downstream module translates the predicted frames into actions. However, existing methods typically ignore information from past interactions, limiting their ability to adapt to latent physical properties that can only be revealed through trial and error, such as whether a door should be pushed or pulled, or how friction affects object dynamics. When a plan fails, these methods usually replan from scratch without leveraging the information revealed by the failure. To address this limitation, we introduce RELIC, REplanning with Latent embedding refInement and Candidate rejection, a video planning framework that adapts to hidden physical properties from test-time failures. RELIC optimizes a latent embedding that captures the environment's hidden physical properties from interaction videos and introduces a rejection-based sampling mechanism that filters out hypotheses inconsistent with prior failures. Across eight tasks in two simulation suites, RELIC consistently reduces the number of replanning steps required for success, and linear probes show that its embedding captures the hidden parameters from the interaction itself rather than from scene appearance. Across four challenging real-world robotic manipulation tasks involving hidden interaction modes, e.g., friction, center of mass, and object mass, RELIC raises the one-shot replanning success rate of a video planning baseline from 30.0% to 63.8% after a single physical interaction.
Sources
- RMA: Rapid Motor Adaptation for Legged Robots
- Estimating Mass Distribution of Articulated Objects using Non-prehensile Manipulation
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- DensePhysNet: Learning Dense Physical Object Representations via Multi-step Dynamic Interactions
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving