Bimanual Robot Manipulation via Multi-Agent In-Context Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Bimanual Robot Manipulation via Multi-Agent In-Context Learning".
Dev: Language Models (LLMs) are emerging as powerful reasoning engines for embodied control, and this paper introduces BiCICLe, the first framework enabling standard LLMs to perform few-shot bimanual manipulation without fine-tuning.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: Now that we have the basic idea of the framework, let’s look at exactly what "Bimanual Robot Manipulation via Multi-Agent In-Context Learning" summarizes as its main contribution. The authors are really focusing on how they frame bimanual control as this multi-agent leader-follower problem to tackle the inherent complexity.
Dev: They summarize it by stating that they decouple the action space into sequential, conditioned single-arm predictions, which is their core mechanism for enforcing inter-arm consistency while managing the reasoning burden per agent. This structure is what makes it distinct from previous monolithic approaches.
Taro: That sounds like a solid technical summary for the mechanism, but I wonder if they fully capture the nature of the input observation representation that they use to bridge those two agents? How do we know how well this works when we move from text descriptions to actual physical perception?
Rosa: They represent observations as text-based dictionaries mapping object names to discretized voxel coordinates, which they argue effectively encodes the spatial relationships between objects and the robot’s end-effectors that are critical for bimanual coordination. This is a clever way to encode spatial context without requiring raw pixel data.
Dev: That text encoding allows the LLMs to reason about the scene structure, which is essential for understanding where both arms need to be relative to each other in a coordinated movement, instead of just looking at isolated end-effector poses. That’s how they reduce the reasoning burden per agent by providing richer contextual information.
Taro: So, the summary highlights this text-based encoding as a bridge; it moves us away from needing perfectly aligned visual input for every step and toward a semantic understanding that the LLM can process effectively. This feels like a necessary abstraction for scaling up.
Rosa: It really is an abstraction, and that’s where I see the potential impact—it means we don't need perfect sensor fusion immediately; we can rely on language models to infer those crucial spatial relationships from structured text prompts provided in the demonstration sequences.
Dev: Precisely; it trades raw perceptual fidelity for zero-shot generalizability, which is what they highlight as a worthwhile exchange when compared to methods that rely on massive paired image-action datasets. It’s about finding the right trade-off for deployment speed and generalization power.
Taro: If this holds up, it could mean that future embodied agents don't need incredibly complex perception pipelines just to handle basic coordination; they can rely on a powerful language model to bridge the gap between what the robot sees and what it needs to do next.
Rosa: That’s a huge implication for deployment simplicity; if we can get that level of reliable coordination from text prompts, we significantly lower the barrier for deploying complex manipulation policies onto different robotic platforms.
The paper's summary: Dev: Moving on to what they actually improved, the paper points out that their improvement lies in explicitly conditioning the follower agent on the leader’s plan; this explicit trajectory conditioning is what enables them to enforce inter-arm consistency.
Taro: So, specifically regarding coordination improvements, they claim this leads to gains like twenty-two point six percent over methods like Sequential Arms and eleven point zero percent over Dual Agent methods on tasks like "Straighten Rope." That quantitative evidence shows a measurable benefit in synchronization for tightly coupled movements.
Rosa: That’s a tangible result; seeing those percentage improvements against established baselines gives us concrete proof that the leader-follower structure is better at handling synchronized lifting or precise handovers than simpler methods. It proves the explicit conditioning helps beyond just factorization alone.
Dev: And they also demonstrated this capability in a real-world setting, showing a success rate of fifty-three point three percent across three tasks on a physical bimanual platform, which is impressive given the complexity of real-world variables. That moves it from theoretical proof to practical applicability.
Taro: I’m interested in the generalization aspect they touched on; they showed performance on out-of-distribution tasks like "Close Jar" and "Take Item Out of Box," achieving success rates around fifty-four point five percent compared to less than ten percent for fine-tuned supervised methods. That zero-shot capability is a big deal for true autonomy.
Rosa: The fact that it handles those novel scenarios without needing task-specific fine-tuning is significant because it shows the model has learned general manipulation principles, not just memorized a specific set of actions. It’s about learning how to manipulate things fundamentally.
Dev: However, I also have to point out the limitations they flag; they noted that including object rotations in observations generally degraded performance, reducing success from seventy point five percent down to sixty-five point two percent, suggesting that position-only observations are a better default for this specific interface right now.
Taro: That limitation is important because it tells us exactly what we need to address next; if the system struggles with rotations, then improving the observation representation itself, maybe by incorporating rotation information in a more robust way, becomes the next research frontier.
Rosa: So, the paper suggests that while position-only observations work well for this ICL interface right now, it points toward future work needing richer sensory input to handle more complex manipulation geometries effectively.
The paper's improvements: Rosa: So, to wrap up our discussion on "Bimanual Robot Manipulation via Multi-Agent In-Context Learning," the central implication is that we’ve established a training-free paradigm for bimanual manipulation by bridging the dimensionality–coordination trade-off through explicit trajectory conditioning.
Dev: That structure proves that structured prompting strategies are robust enough to transfer from simulation environments to physical Franka Panda systems, showing it works in practice across different model scales. The leader-follower decomposition is effective at managing the complexity of inter-arm dependencies without overwhelming the agent with too much simultaneous action prediction.
Taro: For me, I think the most important implication is that this method validates using LLMs as generalist planners for embodied control tasks, showing they can infer complex manipulation goals from just a few demonstrations without needing massive, task-specific training sets.
Rosa: That’s right; it shows that we can achieve high performance on bimanual tasks by leveraging the inherent reasoning capabilities of language models in a way that is highly accessible and transferable to new applications. It's about making complex coordination achievable through this structured prompting strategy.
Dev: We should watch how they handle the efficiency trade-offs going forward, specifically with those scaling techniques; the paper showed that while "Best-of-N" improves performance slightly, it comes with a significant increase in token budget that we need to manage for real-time deployment.
Taro: I just want to add one final thought: while this framework is powerful for learning coordination patterns, we still have to rigorously test its resilience when the world misbehaves in unpredictable ways, because the current structure relies heavily on the leader’s initial trajectory being correct.
Rosa: That’s a very fair caution; we need more work on making this system truly resilient to unexpected environmental disturbances before we can deploy it for high-stakes tasks. But overall, "Bimanual Robot Manipulation via Multi-Agent In-Context Learning" gives us a clear roadmap for how LLMs can tackle complex embodied control problems.
Dev: It certainly provides a clear direction; the path forward involves refining those conditioning channels and optimizing the inference pipeline to ensure we get high performance without sacrificing the low latency our control systems demand.
Conclusion: Rosa: So, we've seen how BiCICLe tackles bimanual manipulation by framing it as a multi-agent leader-follower problem using in-context learning, and now we get to the conclusion of the paper itself.
Dev: Yeah, they summarize their findings by showing that explicit trajectory conditioning between the leader and follower agents is what allows them to successfully decouple the action space and enforce consistency. It really hammers home how that decomposition works for reducing reasoning burden.
Rosa: Exactly; it seems this method offers a solid path forward for training-free control, suggesting that we can achieve high-quality coordination without needing massive datasets or complex reward functions.
Taro: I gotta say, the generalization capability they showed on out-of-distribution tasks is what really gets me excited about this work; it suggests these models are learning underlying manipulation principles rather than just memorizing specific trajectories.
Dev: That’s a key point, Taro; when you see them handle novel scenarios like those in "Close Jar," it really shows the model has learned something more fundamental about spatial relationships than just imitation. However, I still have to ask about deployment time; how fast can we expect this loop rate to be on a real robot versus the simulation speed they used?
Rosa: That’s a big question, Dev; I'm wondering if this works reliably outside the lab for long periods without constant supervision.
Dev: Well, based on their results, the success rates across those three tasks in real-world trials were pretty respectable at fifty-three point three percent, which is encouraging for deployment, but we still need to check robustness against sensor noise and sudden occlusions.
Taro: That resilience under uncertainty is definitely where my focus lies; if the world misbehaves and things change unexpectedly, how well does this leader-follower structure handle that disruption?
Rosa: It seems the leader's initial plan sets a strong trend, but I hope the follower agent can adapt quickly enough when that trend gets derailed by something unexpected in its observation.
Dev: That adaptation is what we need to monitor closely; if the latency in processing the leader’s plan or updating it with new observation data causes drift, those gains disappear instantly.
Taro: I think the way they structure the conditioning helps mitigate that drift, but it still leaves a lot of room for improvement in how quickly an agent can correct its trajectory when faced with true novel situations.
Rosa: So, to wrap up on "Bimanual Robot Manipulation via Multi-Agent In-Context Learning," this paper lays out a really compelling framework for using LLMs to handle the inherent complexity of coordinating multiple robotic arms.
Dev: It’s a solid contribution that proves structured prompting can effectively manage high-dimensional action spaces, and I’m looking forward to seeing how their next iterations handle real-time constraints on the hardware side.
Taro: I’m just curious to see if this approach can scale up to more complex, multi-task manipulation scenarios beyond what they tested in the TWIN benchmark.
Rosa: Well, that's where we'll be looking next; keeping an eye on how this framework evolves is going to be crucial for the future of embodied AI.
Alessio Palma, Indro Spinelli, Vignesh Prasad, Luca Scofano, Yufeng Jin, Georgia Chalvatzaki, Fabio Galasso
Sapienza University of Rome, Italy
cs.RO, cs.AI, cs.MA
Submitted: 2026-04-22
Updated: 2026-09-29
Comments: Accepted at CoRL 2026
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 89/100
The gist: Language Models (LLMs) are emerging as powerful reasoning engines for embodied control, and this paper introduces BiCICLe, the first framework enabling standard LLMs to perform few-shot bimanual
Key concepts
- BiCICLe Framework
- A novel approach treating bimanual control as a multi-agent leader-follower problem. It decouples the complex action space into sequential, single-arm predictions, allowing two specialized LLM agents to coordinate their movements.
- Leader Agent
- The first LLM agent responsible for predicting its full trajectory based only on single-arm demonstrations and the current scene observation. It establishes the overall plan for one arm.
- Follower Agent
- The second LLM agent that predicts its actions by conditioning on both the scene observation and the complete plan generated by the Leader. This forces inter-arm consistency in coordination.
Terminology
Summary
Language Models (LLMs) are emerging as powerful reasoning engines for embodied control, and this paper introduces BiCICLe, the first framework enabling standard LLMs to perform few-shot bimanual manipulation without fine-tuning. This matters because bimanual manipulation—requiring precise inter-arm coordination—is challenging for existing ICL methods due to high-dimensional action spaces and tight constraints. BiCICLe addresses this by framing control as a multi-agent leader-follower problem, allowing LLMs to learn complex coordination patterns directly from demonstrations through structured prompting.
BiCICLe Framework and Architecture
BiCICLe frames bimanual control as a multi-agent leader-follower problem, decoupling the action space into sequential, conditioned single-arm predictions.
It utilizes two distinct LLM agents: a Leader and a Follower. The Leader agent predicts its full trajectory first from the scene observation using only single-arm ICL demonstrations. The Follower agent then predicts its actions conditioned on both the observation and the Leader’s complete plan. This factorization is designed to enforce inter-arm consistency while reducing the reasoning burden per agent.
Action and Observation Representation
The framework operates within a discretized action space where each arm has a 7-dimensional end-effector action: an SE(3) pose plus a binary gripper command, resulting in a joint bimanual space of Z14 at each keyframe. Observations are represented as text-based dictionaries mapping object names to discretized voxel coordinates, which encode the spatial relationships between objects and the robot’s end-effectors that are critical for bimanual coordination.
Leader-Follower Decomposition
The process is structured into two sequential phases:
-
Phase 1 (Leader prediction): The Leader receives a prompt containing single-arm ICL demonstrations and the test observation, generating the predicted leader trajectory, denoted as
Aˆ L = [aˆ L 1,..., aˆ L KˆL].
-
Phase 2 (Follower prediction): The Follower is conditioned on the Leader’s plan. This is achieved by augmenting its observation dictionary with the leader's predicted actions, creating an
augmented observation
and generating its own trajectory, Aˆ F.
Performance and Evaluation
BiCICLe was validated on 13 tasks from the TWIN benchmark, achieving a 70.5% average success rate,
which outperforms the best training-free baseline by 6.1 percentage points and surpasses most supervised methods. Furthermore, the framework demonstrated superior real-world performance on three tasks without hardware-specific retraining, achieving a success rate of 53.3% across these trials compared to baselines like KAT-DA (40.0%) and RoboPrompt-DA (30.0%).
Ablation Studies and Insights
Ablation studies revealed the necessity of the leader-follower conditioning channel; Sequential Arms, which removes this mechanism, only reaches 65.0%, remaining 5.5 percentage points below BiCICLe (70.5%).
Additionally, analysis showed that test-time scaling techniques like Best-of-N
improves performance slightly to 71.2% but incurs a significant cost increase, whereas the Conversation
variant degrades performance substantially due to issues like Gripper-state inversion
and Spatial coordinate drift.
Finally, including object rotations in observations generally degraded performance (reducing success from 70.5% to 65.2%), suggesting that position-only observations are the better default for this ICL interface.
Conclusion
BiCICLe establishes a training-free paradigm for bimanual manipulation by effectively bridging the dimensionality–coordination trade-off through explicit trajectory conditioning, proving that explicit conditioning helps beyond factorization alone.
It demonstrates that this structured prompting strategy is robust across different model scales and successfully transfers from simulation to physical Franka Panda systems.
Prompt Templates
The framework relies on specific prompt templates for inference:
(Leader Arm System Prompt)
"You are the 〈rightleft〉 arm of a bimanual Franka Panda robot with parallel grippers. We provide you with some demos in the format of observation>[action 1, action 2,...]. Then you will receive a new observation and you need to output a list of actions that matches the trend in the demos. Do not output anything else."
(Follower Arm System Prompt)
"You are the 〈leftright〉 arm of a bimanual Franka Panda robot with parallel grippers. We provide you with some demos in the format of observation>[action 1, action 2,...]. Then you will receive a new observation and you need to output a list of actions that matches the trend in the demos. Do not output anything else.
Improvements for AI systems
Here are specific improvements to AI systems based on the BiCICLe framework, detailing what these improved systems can achieve:
-
Improved Bimanual Coordination via Leader-Follower Decomposition:
-
Enhanced Generalization to Out-of-Distribution (OOD) Tasks via Prompt Learning:
-
Robust Real-World Deployment with Low Hardware Dependency:
-
Optimized Efficiency for Edge/Real-Time Robotics:
- Improved Bimanual Coordination via Leader-Follower Decomposition:
The system can now handle complex, time-sensitive bimanual tasks (like synchronized lifting or precise handover) by explicitly modeling inter-arm dependency rather than relying on monolithic joint action space prediction.
It achieves this by decoupling the problem into a Leader agent and a Follower agent. The Leader predicts its trajectory first based on the scene observation, and the Follower then predicts its actions conditioned specifically on that leader's plan.
The system can now perform primary-secondary
manipulation where one arm dictates the necessary temporal sequence for the other (e.g., opening a door while another arm reaches for an object), leading to significantly higher success rates in tasks requiring tight synchronization, as demonstrated by gains like +22.6% over SA and +11.0% over DA on Straighten Rope.
- Enhanced Generalization to Out-of-Distribution (OOD) Tasks via Prompt Learning:
The system can adapt to entirely new bimanual manipulation tasks without requiring any task-specific fine-tuning or massive datasets.
By serializing demonstrations into a text prompt and using Large Language Models (LLMs) as generalist planners, the system leverages the LLM's inherent reasoning capabilities to infer the objective from few demonstrations. This is superior to traditional Imitation Learning (IL) or Supervised Learning (SL), which fail when encountering novel object geometries or scene layouts.
The improved AI can successfully tackle out-of-distribution
challenges like the Close Jar
and Take Item Out of Box
tasks, where it achieves 54.5% success rates compared to <10% for fine-tuned SL methods, proving its zero-shot generalization capability.
- Robust Real-World Deployment with Low Hardware Dependency:
The system can be deployed directly on physical bimanual robotic platforms (e.g., Franka Panda) using only a pose-based ICL interface, eliminating the need for expensive hardware-specific retraining or massive paired image-action datasets required by Reinforcement Learning or Imitation Learning.
The framework handles real-world complexities such as varying lighting, occlusions, and object poses by relying on discretized semantic object coordinates in the observation dictionary rather than raw pixel data (as shown by the superior performance of text-based ICL over appearance-based keypoint methods like KAT). This makes the system more robust to real-world sensory noise and significantly reduces the barrier to entry for deploying complex manipulation policies.
- Optimized Efficiency for Edge/Real-Time Robotics:
The system offers a strong trade-off between performance and computational cost, allowing for deployment on systems with constrained resources while maintaining high coordination quality.
Unlike monolithic joint prediction (SA), BiCICLe maintains a per-arm action space of only 7 dimensions. Furthermore, the Leader-Follower decomposition reduces the reasoning burden per agent call compared to Dual Agent (DA) methods, leading to a more efficient inference pipeline. While Best-of-N
scaling increases token budget significantly (up to 180.7k tokens), it provides a measurable performance gain (+1.3% average success) over the base BiCICLe pipeline, proving that the leader-follower structure is the optimal efficiency–accuracy operating point for training-free ICL policies.
Abstract
Large Language Models (LLMs) have emerged as powerful reasoning engines for embodied control. In particular, In-Context Learning (ICL) enables off-the-shelf, text-only LLMs to predict robot actions without any task-specific training while preserving their generalization capabilities. Applying ICL to bimanual manipulation remains challenging as the high-dimensional joint action space and tight inter-arm coordination constraints rapidly overwhelm standard context windows. To address this, we introduce BiCICLe (Bimanual Coordinated In-Context Learning), the first framework that enables standard LLMs to perform few-shot bimanual manipulation without fine-tuning. BiCICLe frames bimanual control as a multi-agent leader-follower problem, decoupling the action space into sequential, conditioned single-arm predictions. Evaluated on 13 tasks from the TWIN benchmark, BiCICLe achieves 70.5% average success rate, outperforming the best training-free baseline by 6.1 percentage points and surpassing most supervised methods. We also demonstrate superior real-world performance on 3 tasks without hardware-specific retraining. The project page is available at https://alesspalma.github.io/bicicle
Sources
- Qwen2.5-VL Technical Report
- CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
- What Matters in Building Vision-Language-Action Models for Generalist Robots
- AnyBimanual: Transferring Unimanual Policy for General Bimanual Manipulation
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos
- OpenAI GPT-5 System Card
- Qwen2.5 Technical Report
- Open3D: A Modern Library for 3D Data Processing
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving