ModPack: Extensible Teleoperation Interface for Bimanual Mobile Manipulation

arXiv:2607.19479 · cs.RO, cs.AI · Submitted 2026-07-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "ModPack: Extensible Teleoperation Interface for Bimanual Mobile Manipulation".

Rosa: ModPack introduces a modular and extensible teleoperation system designed to support diverse robot embodiments and task requirements within a unified framework,

Dev: First, who's behind it and why it matters.

Paper summary: Rosa: Thinking about the title, "ModPack: Extensible Teleoperation Interface for Bimanual Mobile Manipulation," it really captures the essence of what this system achieves by focusing on extensibility and bimanual control across mobile platforms.

Dev: I think the authors managed to clearly articulate how this backpack core serves as that common substrate, allowing them to decouple the fundamental system infrastructure from any specific robot or task needs.

Taro: The implication for autonomy research is significant because if we can create a standardized way to gather diverse data across varied hardware, it lowers the barrier for training general policies.

Rosa: Essentially, ModPack gives researchers a flexible and reusable framework specifically designed for collecting data needed to train imitation learning models effectively.

Dev: It’s about making sure the data collection interface doesn't become a bottleneck when you're trying to generalize control strategies across different robot types or manipulation environments.

Taro: If this design proves practical outside of controlled lab settings, it could mean that we can gather more diverse datasets quickly, which feeds directly into building more robust AI agents.

Rosa: The authors open-sourced the complete hardware design and software stack, which is a big step toward making this kind of flexible teleoperation accessible to a wider community for future work.

Dev: I'm looking at the long-term potential here; if this modular approach scales well, it could enable rapid iteration in complex manipulation tasks that currently require bespoke solutions for every robot.

Taro: It suggests a path forward where the focus shifts from building one perfect system to building a highly adaptable ecosystem of components and policies.

Rosa: So, the core message of "ModPack: Extensible Teleoperation Interface for Bimanual Mobile Manipulation" is that modularity is the way to support diverse robot embodiments without sacrificing unified control capabilities.

Conclusion: Rosa: So, we've seen how ModPack uses that modular backpack to handle everything from precise joint control to mobile movement across different robot setups, and now we’re wrapping up with some final thoughts on the paper itself.

Dev: Yeah, I gotta say, that title really nails what they accomplished by focusing on extensibility and bimanual control across varied platforms. It sounds like they built a flexible backbone for teleoperation rather than just one specific robot solution.

Taro: I think it's important to remember the authors are pushing for a system that supports diverse embodiments, which is key because if you can get a unified interface to work across different hardware, the possibilities for general autonomy training really expand.

Rosa: Exactly! If we can create this kind of standardized way to collect data from different robots without having to completely redesign our entire software stack every time we switch platforms, that makes the whole learning process much more scalable.

Dev: From an engineering standpoint, I’m thinking about how robust that unified interface needs to be; it has to maintain a stable loop rate even when you swap out those leader arms or add new perception modules mid-task.

Taro: That robustness is what interests me most—if the system can handle unexpected situations in the physical world, like an object slipping or occlusion happening suddenly, that’s where the real value for autonomy lies.

Rosa: That brings us to the bigger picture; this isn't just about making one specific robot better; it’s about creating a data collection pipeline that lets researchers explore complex manipulation scenarios much faster than before.

Dev: And I wonder how long this setup can actually run reliably in a real-world, messy environment before we start seeing those latency issues creep in or the hardware starts failing under sustained stress.

Taro: That’s exactly the question—can it handle the unpredictability of real-world interaction long enough to generate high-quality training data for sophisticated AI models?

Rosa: So, while ModPack shows incredible flexibility in its design, we still need to figure out how much time and physical durability it has before we can confidently deploy it outside of a controlled lab setting.

Dev: That’s the crucial next step; proving that the system maintains those low-latency connections and doesn't have hidden failure modes under real load is what separates a proof-of-concept from a reliable tool.

Stanford University

cs.RO, cs.AI

Submitted: 2026-07-21

Updated: 2026-10-07

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: ModPack introduces a modular and extensible teleoperation system designed to support diverse robot embodiments and task requirements within a unified framework, addressing the limitations of existing

Key concepts

Leader Arms
These are swappable robotic arms that are kinematically identical to the follower arm. They enable direct joint-space teleoperation, allowing the operator to control different robot arm kinematics precisely, which is crucial for adapting to various robot platforms.
Mobile Base
This feature allows the operator's base motion (tracked via an iPhone) to be mapped directly onto the robot's movement. This enables mobile manipulation for tasks requiring large workspaces, ensuring that when the operator walks forward, the robot moves forward in its local frame.
Active Perception
This module uses a Vision Pro to stream real-time robot head-camera views while tracking operator motion. It allows operators to naturally walk and look around during data collection, enabling behaviors like object search and viewpoint selection while compensating for base movement.

Terminology

Summary

ModPack introduces a modular and extensible teleoperation system designed to support diverse robot embodiments and task requirements within a unified framework, addressing the limitations of existing systems that are often tailored to specific hardware. The core concept is a self-contained wearable “backpack” that integrates onboard computation, power, communication, and data storage.

How it works

ModPack is built on a shared interface where the backpack acts as a common substrate, decoupling core system infrastructure from task- or robot-specific functionalities. This foundation supports plug-and-play capability modules that allow operators to flexibly add or remove extensions based on task needs. Key capabilities enabled by these modules include:

  1. Cross-robot joint-level control through swappable leader arms, enabling adaptation to diverse robot arm kinematics and supporting precise joint-level control across different robot platforms.

  2. Mobility, achieved by tracking the operator’s base motion via an iPhone mounted on ModPack, which enables mobile manipulation for tasks that require large workspace coverage.

  3. Active perception, where ModPack streams real-time robot head-camera observations while allowing the operator to control the camera pose, enabling demonstrations of behaviors such as object search, occlusion handling, and viewpoint selection.

  4. Haptic feedback provided by leveraging force feedback from the robot and active motors on leader arms to enable safer and more precise teleoperation of contact-rich tasks.

System Architecture

The core system is organized around a portable backpack that provides the necessary hardware for modular extensions. The internal volume is organized into five removable shelves housing the onboard compute, power, and optional accessories. The software stack abstracts communication using a lightweight message queue, which decouples the core backpack stack from embodiment-specific implementations. This allows the same system to support different robots and arbitrary subsets of extended modules.

Extensible Modules

The system is extensible through three primary modules: Leader Arms, Mobile Base, and Active Perception.

Leader Arms:

These modules consist of 6-DoF ARX5 or 7-DoF leader arms constructed to be kinematically equivalent to its corresponding follower arm, enabling direct joint-space teleoperation. To enhance stability, active gravity compensation is implemented, where the gravity torque vector is computed using the Orocos Kinematics and Dynamics Library (KDL), followed by filtering using two cascaded exponential moving average (EMA) filters to obtain the final commanded torque. Haptic feedback is integrated by mapping net wrench from force/torque sensors to joint-space torques via the Jacobian transpose, resulting in a final control torque that is the superposition of gravity compensation and haptic feedback.

Mobile Base:

The operator’s SE(2) pose is tracked using an iPhone running a modified WebXR session. An egocentric mapping procedure couples the robot’s base translation directly to the operator’s displacement within their local frame, ensuring that when the operator walks forward, the robot moves forward in its own local frame. A sequence of offsets is applied to recover the user’s true center of rotation, accounting for physical displacement between the iPhone and the operator. This tracked pose is then mapped into world coordinates to issue absolute target setpoints to the robot base controller.

Active Perception:

This module utilizes an Apple Vision Pro running a custom visionOS app to stream real-time robot head-camera observations. The system tracks the operator’s head and body motion, and compensates for base motion before sending adjusted robot head commands. This decoupling enables operators to naturally walk, look around, and control the arms during data collection.

Policy Learning

For policy training, a standard transformer-based Diffusion Policy model is trained for each task using the collected teleoperation demonstrations. The input modalities vary by robot: for the custom mobile robot, the policy takes RGB-D observations from the head camera and RGB observations from the left and right wrist cameras. For the RB-Y1m robot, it takes RGB observations from two head cameras and the left and right wrist cameras, and optionally left and right arm joint torques. All visual inputs are encoded using a CLIP-pretrained ViT, with depth images repeated across three channels before encoding. Torque observations are encoded with sequential CausalConv layers.

Experimental Evaluation

ModPack was evaluated through real-world bimanual mobile manipulation experiments across two distinct robot platforms: the customized mobile robot with ARX5 arms and holonomic base for Cloth Placement, and an RB-Y1m robot for Box Transfer with Haptic Feedback. The results demonstrated the flexibility of the design, showing that policies trained on this data achieve strong deployment performance, validating ModPack's utility as a practical data collection interface for robot learning. For instance, in the Cloth Placement task, policies conditioned on All Cam + Torque achieved higher success rates compared to those using only head or wrist cameras.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the ModPack paper. The core innovation lies in creating a highly modular, unified teleoperation framework that decouples operator input from robot embodiment and task requirements.

Here are the specific improvements to AI systems derived from this work, detailing what these improved systems can achieve:


  1. The system can support the development of a generalized Teleoperation Policy Learning Framework capable of rapidly adapting to novel robot hardware and task domains without requiring complete retraining.

  2. This framework allows for the creation of high-fidelity imitation learning policies (e.g., Diffusion Policies) by leveraging heterogeneous, multi-modal sensory inputs (RGB, RGB-D, proprioception, force/torque signals).

  3. The system enables the creation of robust Cross-Modal Perception Policies that can dynamically select the optimal visual input modality—such as switching between head-camera only views for long-horizon search and full camera/torque views for contact-rich manipulation—based on real-time task context (as demonstrated by the success of 'Head Cam' vs. 'All Cam' policies).

  4. The system allows for the development of Haptic-Conditioned Control Policies where arm joint torques are explicitly conditioned on force/torque feedback signals, enabling the robot to learn precise contact-rich behaviors like stable grasping and controlled object placement (as evidenced by the superior performance of 'All Cam + Torque' policies in Box Transfer).

  5. The system facilitates Simultaneous Multi-Objective Control Policies that explicitly disentangle operator motion into component-wise commands (arm, base, perception), allowing the learned policy to manage complex, coupled dynamics—such as mobile manipulation while simultaneously tracking a dynamic visual target—without operator ambiguity.

  6. The framework can be used to generate Active Perception Policies that automatically optimize viewpoint selection during data collection or execution based on task progress (e.g., automatically shifting from a search pattern to a detailed inspection view when an object is detected).

  7. The system enables the creation of Ergonomic Control Policies that incorporate gravity compensation and haptic feedback into the learned control loop, resulting in policies that are inherently safer and more robust for bimanual manipulation under physical load constraints.

This improved AI system can perform:

  1. Perform complex, bimanual mobile manipulation tasks (e.g., Cloth Placement, Box Transfer) across a wide variety of robot embodiments (different leader arms/kinematics).

  2. Execute dexterous contact-rich manipulation by accurately controlling force application and object placement using real-time haptic feedback.

  3. Navigate large workspaces while maintaining precise manipulation through coordinated base motion and arm control.

  4. Learn generalized skills by ingesting diverse, multi-modal human demonstrations (visual, proprioceptive, force) and producing high-performance policies that are robust to variations in robot hardware and environmental conditions.

Abstract

Existing teleoperation systems are often tailored to specific robot hardware and task domains, limiting their scalability and adaptability. We present ModPack, a modular and extensible teleoperation system designed to support diverse robot embodiments and task requirements within a unified framework. At the core of ModPack is a self-contained wearable "backpack" that integrates onboard computation, power, communication, and data storage. Built on top of this shared interface, the system supports plug-and-play capability modules including joint-level teleoperation with haptic feedback, mobile manipulation, and active perception. Experiments across two distinct robot platforms and real-world mobile manipulation tasks demonstrate that ModPack provides a flexible and reusable framework for data collection and policy learning and collects higher quality data than a robot-free alternative. To support future research, we open-source the complete hardware design and software stack. Project website: https://modpack-robotics.github.io/

Sources

Related papers