PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation

arXiv:2601.17885 · cs.CV, cs.AI, cs.RO · Submitted 2026-08-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation".

Jane: The paper was written by Authors not found in provided excerpt. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: So, we’ve talked about the title and where this research is headed, but let’s dive into the summary of "PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation."

Jane: The authors highlight that previous vision-language action models often fail because they treat multiple camera views too loosely.

Lu: They say existing systems just concatenate the visual tokens from different cameras without actually enforcing three dee consistency between those views, which is a major problem in real-world scenarios.

Meng: That makes sense, because if you’re trying to grasp an object with two hands, having weak spatial understanding means your robot might reach for the wrong part of the object.

Lalam: The summary points out that language instructions were also often just injected globally, meaning they weren't really helping pinpoint specific objects or locations.

Tom: And PEAfowl solves these two major bottlenecks: better geometric reasoning across all the cameras and a much smarter way to understand what parts of the instruction are relevant.

Jane: It seems like they’ are building a perception system that is both spatially aware and instruction-aware, which is quite a big leap.

Lu: I see it as moving from a generalist vision model to creating a highly targeted spatial reasoning engine for complex tasks.

Meng: From an engineering view, this means we aren't just throwing all the visual data at the policy; we're filtering and structuring it.

Lalam: The summary suggests that AI is getting much better at reading intent and executing it, not just reacting to visual input.

Tom: That sets us up perfectly for understanding how they actually achieve this, which leads into our next segment.

Improvements: Tom: Now that we understand the problem and the high-level solution, let's look at the specific improvements in "PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation."

Jane: The authors introduce two main architectural innovations to solve those problems. First, there is this "Geometry-Guided Multi-View Fusion" module.

Lu: That module takes the raw RGB and depth data from all the cameras and uses three dee lifting to create shared spatial representations across views.

Meng: It's not just merging pictures; they are projecting tokens into a shared base frame, which allows them to aggregate information from different cameras that see the same physical region.

Lalam: This is crucial because it means the robot knows exactly where an object is regardless of which camera captures it, which helps us achieve more consistent movement.

Tom: It sounds like they are solving the "where is this?" problem by creating a set of three dee anchors for every single visual token.

Jane: And to make that process robust, they also use "Pairwise RGB-D Token Fusion," which only looks at co-located pairs, ignoring noisy depth data.

Lu: That's a brilliant way to maintain the semantic richness of the RGB image while still leveraging the geometric guidance from the depth information.

Meng: It prevents noise from corrupting the core visual understanding, which is vital when using commodity sensors in real life.

Lalam: The text-side improvement is equally important: they replace global conditioning with a "Perceiver-Style Text-as-Query Readout."

Tom: That means the AI isn't just receiving the whole sentence; it’s actively querying the images to find exactly what the words refer to.

Jane: Exactly, it’s like having a highly focused attention mechanism that makes sure the AI only focuses on task-relevant objects.

Lu: It’s a sophisticated way of saying "pay attention to these things" instead of just feeding it all raw data and hoping it pays attention itself.

Meng: This approach ensures the policy is grounded in the instruction, making the whole system much more predictable for real-world deployment.

Results: Tom: The theory is fascinating, but how does it perform in practice? Let's look at the results presented in "PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation."

Jane: The simulation results on RoboTwin two point zero show that PEAfowl achieves a massive average success rate of forty-seven point one percent.

Lu: That’s a huge jump from the baselines, and the data suggests that even under heavily domain-randomized settings, we can be much more stable.

Meng: Stability is key here; if the environment changes—the lighting shifts or the background becomes cluttered—a robust system is required to keep working.

Lalam: The performance in-distribution also shows that AI systems are becoming less reliant on just a few examples and more generalizable.

Tom: We are seeing big gains in long-horizon tasks, which means the consistent spatial awareness they achieve really matters when things get complicated.

Jane: And we can't forget the real-world results, where they show consistent sim-to-real transfer on the dual-arm platform.

Lu: That confirms that these models aren' not just working in a perfect simulation but are ready for deployment in messy, physical reality.

Meng: The fact that it performs well when we intentionally narrow the cameras’ field of view shows how vital those multiple perspectives truly are for robust operation.

Lalam: It suggests a future where AI can handle tasks that look very different from the ones it was trained on us.

Conclusion: Tom: We've covered so much ground today, from the conceptual title to seeing how these systems perform in the real world, but let's wrap things up by summarizing "PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation."

Jane: This paper provides a powerful framework that merges geometry and language guidance into one cohesive system for bimanual manipulation.

Lu: It is a major milestone because it proves we can achieve strong spatial reasoning combined with accurate instruction grounding simultaneously.

Meng: I think the most impactful takeaway is that this approach, allowing for stable three dee-aware learning without adding test-time overhead, makes it highly practical for real manufacturing use.

Lalam: It allows AI to learn and execute tasks with a level of spatial coherence that feels genuinely intelligent and dependable in the world.

Tom: We've seen the numbers, we've seen the methods, and we’ve seen what this system can do; it’s clear that "PEAfowl" is a significant contribution to advancing robotic AI.

Authors not found in provided excerpt.

cs.CV, cs.AI, cs.RO

Submitted: 2026-08-24

Updated: 2026-08-25

Comments: Accepted by IEEE Robotics and Automation Letters (RA-L), 2026. This version includes an extended appendix with additional implementation details, deployment analysis, robustness evaluation, and real-world evaluation protocol. DOI: 10.1109/LRA.2026.3726379

DOI: 10.1109/LRA.2026.3726379

Project page: https://peafowlvla.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 88/100

The gist: This paper introduces PEAfowl, a perception-enhanced multi-view vision-language-action (VLA) policy designed for bimanual manipulation in cluttered, changing environments.

Key concepts

Bimanual Manipulation
This refers to complex tasks requiring a robot to use two hands simultaneously. The paper focuses on improving the system's ability to perform these coordinated actions by achieving consistent spatial awareness and executing precise movements.
Geometry-Guided Multi-View Fusion
This architectural module takes raw RGB and depth data from multiple cameras. It uses 3D lifting to create shared spatial representations, allowing the robot to aggregate information about a physical region regardless of which camera captures it.

Terminology

Summary

This paper introduces PEAfowl, a perception-enhanced multi-view vision-language-action (VLA) policy designed for bimanual manipulation in cluttered, changing environments. It addresses critical failures in existing VLA models—specifically weak 3D spatial understanding due to view-agnostic token concatenation and coarse instruction grounding caused by global language conditioning—to enable stable robot control under occlusions and viewpoint variations.

The Core Problem

Existing bimanual VLA policies often struggle with generalization in cluttered scenes with distractors and under appearance and viewpoint variations. The authors identify two primary technical bottlenecks:

  1. Multi-view features are typically fused via view-agnostic token concatenation, which fails to model explicit cross-view geometric correspondences or enforce 3D consistency.

  2. Language is often injected as global conditioning, leading to unfocused, instruction-agnostic attention that fails to identify relevant objects in multi-task, multi-object scenes.

How it works

PEAfowl implements a dual-pathway architecture that integrates geometry-guided multi-view fusion with language-guided readout. To achieve geometrically grounded, cross-view consistent representations, the spatial module performs:

((

) Per-patch RGB–D fusion:

) Depth distribution prediction:

) Differentiable 3D lifting:

) Cross-view neighbor aggregation:

The system predicts a discrete depth distribution for each visual token, allowing it to softly lift each 2D token into the robot base frame. It then performs local 3D neighborhood aggregation in a shared base frame, which aligns tokens that correspond to the same physical region across cameras, thereby improving robustness to occlusions.

To address instruction grounding, the model replaces global conditioning with a Perceiver-style text-as-query readout over frozen CLIP visual features. In this mechanism, text tokens act as latent queries that iteratively cross-attend to per-view patch features, producing a compact set of language-conditioned visual tokens. This process enables iterative evidence accumulation and sharpens attention over task-relevant objects and spatial relations.

Depth Distillation for Real-World Use

To handle the noisy and incomplete commodity depth provided by standard sensors without increasing inference overhead, the authors utilize a training-only depth distillation scheme. A pretrained Camera Depth Model (CDM) acts as a depth teacher to supervise the predicted per-token depth distributions during training. This process transfers refined geometric priors into the policy while allowing the model to use only raw depth images during both training and test time, ensuring no added computational cost at deployment.

Experimental Results

The effectiveness of PEAfowl was validated through extensive simulation on RoboTwin 2.0 and real-robot experiments using a dual-arm AgileX Piper platform. Key findings include:

In Simulation:

) Under domain-randomized settings, PEAfowl improves the strongest baseline (SEM) by 23.0 pp in success rate.

) It demonstrates superior performance on long-horizon and occlusion-heavy tasks compared to models like π0 and SEM.

In Real-World Deployment:

) The model demonstrates reliable sim-to-real transfer, particularly on tasks requiring precise bimanual coordination.

) Depth distillation provides consistent improvements, especially on depth-sensitive tasks such as Put Bottles Dustbin and Hanging Mug.

Despite having only 300M trainable parameters, PEAfowl outperforms existing bimanual VLA models and visuomotor baselines across a variety of novel scene appearances and language instructions.

Improvements for AI systems

To improve existing Vision-Language-Action (VLA) systems, I would implement the following architectural upgrades based on the PEAfowl framework:

  1. Integrate a Geometry-Guided Multi-View Fusion (GGMVF) module to replace view-agnostic token concatenation. This involves:

Reads RGB and depth features into multi-scale pyramids;

Predicts discrete per-token depth distributions rather than single scalar values;

Performs differentiable 3D lifting of tokens into a shared robot base frame;

Uses 3D proximity (K-nearest neighbors) to aggregate features across cameras.

  1. Replace global text conditioning with a Perceiver-style Text-as-Query Readout. This involves:

Using frozen CLIP vision/text encoders to extract high-fidelity features;

Implementing latent blocks where text tokens act as queries that iteratively cross-attend to visual patch tokens;

Pooling these into compact, instruction-conditioned context tokens.

  1. Implement Training-Only Depth Distillation. This involves:

Using a heavy pretrained Camera Depth Model (CDM) as a teacher during training to supervise the per-token depth distribution head;

Maintaining raw, noisy commodity depth inputs during inference to ensure zero added latency.

  1. Adopt a Joint-Centric Diffusion Action Decoder. This involves:

Encoding robot proprioception (joint angles and FK-derived poses) into joint-graph attention tokens;

Using these tokens as the primary state representation for the diffusion process to ensure bimanual coordination is grounded in the robot's physical structure.


By implementing these specific improvements, the enhanced AI system will be able to:

  1. Execute stable, high-precision bimanual manipulation in cluttered environments with frequent self-occlusions and inter-object occlusions by maintaining a 3D-consistent spatial understanding.

  2. Generalize robustly to Domain-Randomized settings, including drastic changes in lighting, background textures, tabletop heights, and camera viewpoints/poses.

  3. Perform precise instruction grounding in multi-object scenes (e.g., distinguishing between the clear mineral water bottle and the green canned tea) by focusing attention specifically on task-relevant visual evidence through text-aware retrieval.

  4. Maintain high success rates in long-horizon tasks that require maintaining spatial awareness through repeated viewpoint shifts and object displacements.

  5. Achieve reliable sim-to-real transfer using low-cost, noisy commodity depth sensors without requiring expensive high-fidelity depth hardware at inference time.

Sources

Related papers