Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Multisensory Continual Learning".
Dev: Robot manipulation often depends on sensory data beyond vision, especially in contact-rich tasks where force, tactile, or audio feedback reveals interaction states not directly visible from images.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So, wrapping up our discussion on "Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force," the authors introduce MuSe as a way to adapt vision-only policies using limited multisensory data through multi-stage fusion, future prediction, and experience replay.
Dev: Essentially, they are showing that you can improve pretraining performance by incorporating force sensing without needing huge amounts of new task-specific data upfront.
Taro: The main implication I see is that this framework provides a pathway for building more versatile robots that can handle contact-rich tasks efficiently by leveraging existing visual skills and augmenting them with learned force predictions.
Rosa: That’s right; it suggests a path toward more general robot capabilities where the ability to interact physically isn't something you have to train from scratch every time a new sensor is introduced.
Dev: From an engineering standpoint, the framework addresses the practical hurdle of data scarcity by creating mechanisms like experience replay and unified representation training that make learning from limited multisensory inputs more effective.
Taro: The work points toward a future where robot autonomy isn't just about following preprogrammed visual paths but involves intelligently predicting and responding to physical interaction dynamics across different sensory inputs.
Rosa: I think the title itself, "Multisensory Continual Learning," really encapsulates the essence of what they did—it’s about continuous adaptation and learning new sensory skills while maintaining what you already know.
Dev: The authors are demonstrating that this approach offers concrete performance gains in both backward and forward transfer scenarios, which is significant evidence for its practical utility in improving robot learning pipelines.
Taro: It's a strong demonstration of how we can bridge the gap between purely visual policies and the complex physical reality encountered by robots in manipulation tasks.
Conclusion: Rosa: So, we've been diving into MuSe and its ability to teach robots new physical skills using force feedback, and now we need to wrap up by talking about what this whole piece actually means for us out here in the real world.
Dev: I think we should start with the title itself, Rosa; "Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force" is pretty descriptive of what they're doing, isn't it? It clearly signals that the core idea involves learning new ways to move when you add more senses.
Taro: I agree with Dev; the title sets expectations for a system that can handle continuous learning while integrating force information into existing vision-based policies. It suggests a progression from just seeing to truly interacting physically.
Rosa: Exactly, and when we look at the authors, they’re showing us how to make robots more adaptable without needing massive new datasets for every single task they throw at them. This is huge because in the lab, we always have those perfect datasets, but out in the field? That's a different story.
Dev: Right; from an engineering standpoint, the authors are tackling that data scarcity problem directly by showing how limited multisensory data can still yield significant improvements on tasks that require force sensing. It’s not about having every sensor forever; it’s about leveraging what you have effectively.
Taro: That's where the autonomy aspect gets interesting for me; if a robot can learn to predict and react to physical contact just by seeing and feeling, it opens up possibilities for navigating unstructured environments where precise force control is essential. What happens when the world throws something unexpected at it?
Rosa: That's a big question, Taro; what does this mean practically? I wonder how long these robots could operate reliably in those messy, real-world scenarios before they start losing their learned force skills or getting confused by novel interactions.
Dev: Latency and failure modes are key here, Rosa; if the prediction loop for force feedback is too slow or inaccurate, that entire system collapses quickly under dynamic conditions. We need to see how robust this architecture is when those predictions drift, especially in high-speed manipulation scenarios.
Taro: The paper suggests the world model component—predicting future observations and actions—is what gives it resilience; if the robot can anticipate physical consequences before they fully happen, it handles unexpected events much better than a purely reactive system.
Rosa: That sounds like a really promising direction for real-world deployment, provided those prediction capabilities hold up under stress. So, to summarize this part of the discussion: the authors have given us a framework that lets robots learn physical interaction skills by intelligently combining vision and force data without needing endless new training data.
Dev: And the implication is that we can start seeing much better performance on tasks like delicate manipulation or grasping in environments where contact is frequent, which is a step toward more capable autonomy.
Taro: It’s about moving beyond just visual navigation toward genuine physical engagement, and MuSe seems to provide a solid mechanism for bridging that gap in learning.
Stanford University
cs.RO
Submitted: 2026-06-29
Updated: 2026-10-06
Project page: https://jadenvc.github.io/multisensory-continual-learning
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: Robot manipulation often depends on sensory data beyond vision, especially in contact-rich tasks where force, tactile, or audio feedback reveals interaction states not directly visible from images.
Key concepts
- Multi-stage fusion
- This method integrates force-torque data with vision and robot proprioception immediately after encoding. It uses an 'early' pathway to concatenate new features with existing tokens, allowing the robot's brain to pay attention across different types of sensory information throughout its processing layers.
- Multi-sensory future prediction
- MuSe trains the model to simultaneously predict what will happen visually, in terms of force/torque readings, and what actions will be taken next. This unified prediction objective forces the model to create a shared representation across all modalities that generalizes well beyond the specific data it was trained on.
- Experience Replay
- To prevent the robot from forgetting its original vision skills when learning new tasks with F/T data, MuSe mixes training samples from both old and new datasets. When old data lacks F/T information, a special 'learned modality-mask token' is used to keep the model stable while it learns the new sensor.
- Backward Transfer
- This refers to the ability of the adapted robot policy to perform tasks on original pretraining tasks even without any additional force-torque supervision. MuSe shows that F/T sensing is surprisingly general and can be effectively transferred back to earlier vision-only skills.
Terminology
Summary
Robot manipulation often depends on sensory data beyond vision, especially in contact-rich tasks where force, tactile, or audio feedback reveals interaction states not directly visible from images. This study introduces MultiSensory World Model (MuSe), a framework that adapts pretrained robot policies to new tasks with newly introduced modalities while preserving performance under the original sensor suite.
The gist
MuSe enables improved performance on pretraining tasks with no additional task-specific data (backward transfer), zero-shot F/T prediction where no F/T supervision was collected (cross-modal generalization), and improved performance on finetuning tasks that benefit from F/T sensing (forward transfer).
How it works
MuSe is a general framework for multisensory continual learning that integrates limited multisensory data into pretrained vision-only policies using three critical components:
-
Multi-stage fusion: MuSe fuses force-torque (F/T) with vision and proprioception immediately after encoding by embedding all modalities into a shared token space. This includes an
early fusion pathway
where new modality features are concatenated with existing tokens, allowing the backbone to perform cross-modal attention throughout its layers. It also uses alate fusion pathway
injecting features through lightweight cross-attention adapters to strengthen modality-specific conditioning throughout the backbone. -
Multi-sensory future prediction (i.e., world modeling): MuSe is trained to predict future visual observations, F/T observations, and actions simultaneously. The objective is formalized as:
L = λaLact + λoLobs + λn+1Lnew, where Lnew supervises the prediction of the newly introduced modality on+1t. This encourages a unified representation across modalities that generalizes beyond the finetuning data and improves backward transfer.
- Experience Replay: To mitigate catastrophic forgetting when fine-tuning on new tasks with multisensory data, MuSe uses experience replay. Training batches are sampled from a mixture of the new multisensory dataset (Dnew) and the original pretraining dataset (Dpre). For samples from Dpre where F/T is unavailable, the missing modality input is replaced with a
learned modality-mask token
and Lnew is masked out, while retaining supervision for actions and original observation prediction.
Key Results
Experiments demonstrate that MuSe achieves significant performance gains across three key areas:
Backward Transfer:
MuSe shows performance gains on pretraining tasks, achieving Backward Transfer: Zero-shot Force Prediction on Pretraining Task +17.5% success on pretrain tasks.
It outperforms the pretrained model with no F/T and the No ER (No Experience Replay) baseline, suggesting that F/T sensing is surprisingly general
and that experience replay is important for preserving visual skills and transferring F/T prediction back to earlier tasks.
Forward Transfer:
MuSe improves performance on finetuning tasks by leveraging F/T sensing. For contact-rich tasks like Vase wiping,
MuSe achieves 11.5/15 success, compared to 5/15 for the No F/T baseline and 8/15 for the No Pretrain model.
Cross-modal Generalization:
MuSe demonstrates zero-shot F/T prediction on pretraining tasks where F/T was recorded but never used for supervision. The MuSe model obtains the lowest L2 error (8.424) compared to baselines, indicating that future image prediction and late fusion adapters help ground force prediction in visual dynamics shared across tasks, leading to stronger cross-modal generalization beyond the finetuning distribution.
Implementation Details
MuSe is instantiated by augmenting a Unified Video-Action (UVA) model with force-torque (F/T) sensing. The policy conditions on past four image frames, past 16 actions, robot proprioception, and synchronized force–torque histories from both fingers.
At deployment, MuSe uses its predicted future F/T trajectory to drive an adaptive compliance control,
allowing the robot to remain stiff in free space while becoming compliant when contact is anticipated. The architecture utilizes a Masked Autoregressive (MAR) framework with a MAR-Large backbone and incorporates specific encoding strategies for F/T measurements into the shared token sequence.
Conclusion
MuSe successfully demonstrates that pretrained robot policies can be expanded with new sensors after pretraining, without requiring large-scale multisensory data from the start. The combination of multi-stage fusion, world modeling via future prediction, and experience replay provides a robust mechanism for learning transferable F/T representations that support forward transfer to contact-rich tasks and backward transfer to original vision-only distributions. Future work will explore extending these principles to other modalities like tactile or audio.
Improvements for AI systems
Based on the provided research paper, here are specific improvements that can be made to existing or future robot manipulation AI systems by integrating the principles of MultiSensory Continual Learning (MuSe):
- Improve Generalization Beyond Training Distributions:
Individual policies are often brittle, performing well only in their training distribution. MuSe provides a mechanism for backward transfer
and cross-modal generalization.
By learning to predict force-torque (F/T) signals and visual observations simultaneously (via multi-sensory future prediction), the system learns a more fundamental, physically grounded representation of interaction dynamics that is transferable.
- Enhance Robustness to Novel Contact Dynamics:
Standard vision-only policies fail catastrophically when encountering unobserved contact forces or slippage (e.g., jamming during insertion). MuSe equips the policy with a predictive model for future F/T, allowing it to anticipate compliance requirements and adjust its motion proactively.
-
The improved system can now navigate high-friction environments like corrugated holes by predicting jams and adjusting compliance in real-time, rather than reacting post-failure.
-
It gains the ability to detect when a peg is
stuck
by observing anomalous F/T patterns during insertion attempts, allowing for corrective wriggling motions.
- Enable Efficient Sensor Expansion with Limited Data:
A major bottleneck in robotics is the scarcity of large, labeled multisensory datasets. MuSe demonstrates that a modest amount of new sensor data (like force-torque) can significantly enhance a policy's general capability on pretraining tasks (backward transfer).
- The improved system can be adapted to use new sensors without requiring massive retraining or collecting huge amounts of task-specific multisensory data for every deployment scenario.
- Improve Learning Efficiency Through Co-Training:
MuSe incorporates an Experience Replay
mechanism that co-finetunes on both the new multisensory data and the original pretraining data, using modality masking when F/T is unavailable. This prevents catastrophic forgetting of fundamental skills learned during vision-only pretraining.
- The system can learn complex contact behaviors (e.g., wiping) while simultaneously retaining its core object manipulation skills (e.g., pick-and-place), ensuring the robot doesn't forget how to move objects when force feedback is absent.
- Achieve Adaptive and Safe Manipulation Control:
The MuSe architecture explicitly separates conditioning
on F/T history from using predicted F/T trajectories to set an adaptive compliance profile for low-level controllers (Adaptive Compliance Policy).
- The improved system can achieve nuanced contact regulation: it can remain stiff during free-space motion for precision, but instantaneously transition to a compliant state upon anticipated contact or constraint, preventing excessive force application and reducing the risk of damaging objects or exceeding safety limits.
- Develop Unified Cross-Modal Representations:
By employing multi-stage fusion (early and late fusion) where new modalities are embedded into a shared token space, MuSe forces the model to learn cross-modal attention between vision, proprioception, and force signals deeply within the transformer backbone.
- This results in a representation that is inherently more physically meaningful than models relying solely on simple late-stage adapters, leading to superior performance in tasks requiring coordinated visual perception and physical interaction.
The improved AI system (MuSe) can perform the following:
-
Perform contact-rich manipulation tasks (wiping, peg insertion) with significantly higher success rates by leveraging force feedback for precise regulation.
-
Generalize its learned contact-rich skills to novel, unseen task variations and object geometries (e.g., different peg/hole sizes).
-
Maintain high performance on basic pick-and-place tasks even after incorporating new contact modalities, ensuring the robot doesn't
forget
how to move objects when force feedback is not relevant. -
Adapt its physical compliance dynamically during interaction to prevent jamming and safely manage forces within operational limits in real-world scenarios.
Abstract
Robot manipulation often relies on sensory feedback beyond vision, particularly in contact-rich settings where force, tactile, or audio signals reveal interaction states that are not directly observable from images. However, these modalities are often hardware- and task-specific, and large-scale multisensory robot datasets remain scarce. As a result, it is impractical to pretrain policies with every sensor they may encounter. We study multisensory continual learning: adapting a pretrained robot policy such as a World-Action Model or VLA to new tasks with newly introduced modalities while preserving performance under the original sensor suite. We propose MultiSensory World Model (MuSe), which incorporates limited multisensory data into pretrained vision-only policies through multi-stage fusion, multisensory future prediction, and experience replay over pretraining data. We instantiate MuSe by augmenting a pretrained vision-only World-Action Model with tactile force-torque sensing and evaluate it on real-world manipulation tasks. Our experiments show that MuSe performs strongly on contact-rich finetuning tasks while preserving, and in some cases improving, performance on the original pretraining tasks. These results suggest that a modest multisensory dataset can improve general robot capabilities beyond the finetuning distribution. Project website: https://jadenvc.github.io/multisensory-continual-learning/
Sources
- In-the-Wild Compliant Manipulation with UMI-FT
- ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation
- ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data
- Touch in the Wild: Learning Fine-Grained Manipulation with a Portable Visuo-Tactile Gripper
- OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
- VTAM: Video-Tactile-Action Models for Complex Physical Interaction Beyond VLAs
- Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning
- Breaking Lock-In: Preserving Steerability under Low-Data VLA Post-Training
- A Taxonomy for Evaluating Generalist Robot Manipulation Policies
- Compliant Residual DAgger: Improving Real-World Contact-Rich Manipulation with Human Corrections
- Unified Video Action Model
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving