Communication-Efficient Relative Pose Estimation with Vision Foundation Models for Ephemeral Collaborative Perception
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Communication-Efficient Relative Pose Estimation with Vision Foundation Models for Ephemeral Collaborative Perception".
Dev: Relative pose estimation is crucial for collaborative perception in multi-robot systems, especially during ephemeral encounters where communication bandwidth and visual overlap are limited.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: To summarize this work, "Communication-Efficient Relative Pose Estimation with Vision Foundation Models for Ephemeral Collaborative Perception" presents a system called CERPE designed to estimate the relative pose between robots ri and rj at every time step t.
Dev: The authors claim that existing methods struggle with ephemeral encounters because they usually require sustained view overlap or incur too much communication cost, which limits their use in real-world collaborative perception scenarios.
Taro: So, the central thesis is that CERPE coordinates vision foundation models to jointly estimate ego-motion and inter-robot relative pose in a training-free pipeline under communication constraints.
Rosa: They achieve this by using SALAD to encode observations into fixed-size descriptors which robots then share with their local pose information in lightweight messages.
Dev: These shared descriptors are used as an online proxy for sufficient visual overlap; specifically, a pair with cosine similarity score greater than τ is treated as a candidate for direct raw-observation exchange for relative pose estimation.
Taro: The way it decomposes the problem into ego-motion and relative pose estimation, and then handles the non-overlapping cases via metrically scaled ego-motion propagation, seems like a practical solution to those real-world constraints.
Rosa: It matters because it moves away from methods that assume constant data flow, making multi-robot coordination viable even when visual overlap is intermittent or missing due to occlusions.
Dev: The primary contribution is this conditional execution pipeline that gates raw-observation requests based on descriptor similarity and uses metric depth to calibrate the monocular scale ambiguity during ego-motion calculation.
Conclusion: Rosa: Looking at the paper, "Communication-Efficient Relative Pose Estimation with Vision Foundation Models for Ephemeral Collaborative Perception," it really highlights how important reliable relative pose is for any multi-robot team working together in the real world.
Dev: The authors, including Qihang Li, Jo-Hao Huang, Jiewen Liu, Suyoung Kang, Hao Zhang, and Peng Gao1, developed a framework that tackles the communication bottlenecks of collaborative perception.
Taro: What this means in simple terms is that we can build systems where robots don't have to constantly stream high-resolution video data back and forth just to know their positions relative to one another during brief meetings.
Rosa: Instead, they use these compact descriptors as a smart filter; if the overlap isn't good enough based on similarity, they skip the heavy raw image requests and instead rely on calculated movement models.
Dev: That’s a significant implication because it makes complex coordination possible in environments with limited bandwidth or unpredictable visual conditions where sustained tracking is impossible.
Taro: I think the biggest impact will be in areas like search and rescue or disaster response, where robots have to navigate close together in chaotic settings without losing their relative positioning.
Rosa: It suggests a new way for autonomous systems to maintain awareness of their peers even when the visual channel is unreliable, which opens up possibilities for much more robust collaborative missions.
Dev: We're seeing a shift towards systems that can handle intermittent information flow effectively, and this CERPE paper provides a concrete structure for how vision foundation models can be adapted for this kind of constrained operational reality.
North Carolina State University · University of Massachusetts Amherst
cs.RO
Submitted: 2026-07-16
Updated: 2026-10-05
Comments: 8 pages, 6 figures
Project page: https://rokeeeto-li.github.io/cerpe.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 69/100
The gist: Relative pose estimation is crucial for collaborative perception in multi-robot systems, especially during ephemeral encounters where communication bandwidth and visual overlap are limited.
Key concepts
- Ego-motion Estimation
- Each robot estimates its own movement (ego-motion) by processing current observations against a memory of past keyframes. This is done using a backbone model (VGGT) and selecting new frames based on temporal spacing and viewpoint change criteria to track local motion.
- Metrically Scaled Ego-Motion
- To resolve scale ambiguity in monocular vision, the framework aligns predicted depth with metric depth data. A scale factor is computed from all keyframes in memory, which is then used to correct the ego-motion estimate, ensuring accurate 6-DoF pose recovery.
- Descriptor Gate
- Instead of constantly exchanging raw images, robots share compact global descriptors derived from vision models. When the similarity between a pair's descriptors exceeds a threshold, it signals sufficient visual overlap, triggering an exchange for direct relative pose estimation.
- Non-Overlapping Propagation
- When visual overlap is absent, the relative pose estimate is maintained by combining previous relative poses with each robot's accumulated ego-motion. This propagation uses metrically scaled ego-motion to ensure the relative pose remains accurate even without new visual data.
Terminology
Summary
Relative pose estimation is crucial for collaborative perception in multi-robot systems, especially during ephemeral encounters where communication bandwidth and visual overlap are limited. This work introduces CERPE, a system-level framework that coordinates vision foundation models to jointly estimate ego-motion and inter-robot relative pose efficiently.
The gist: CERPE reduces unnecessary raw-observation exchange by using continuously shared fixed-size descriptors to gate event-triggered raw-image requests independently of pose estimation, while handling non-overlapping encounters by propagating inter-robot relative poses through metrically scaled ego-motion.
Problem Formulation
The goal is to estimate the relative pose, denoted as the 6-DoF pose Ttij in SE(3), between two robots ri and rj at every time step t, given image sequences Ii and Ij. This estimation is decomposed into two sub-problems: Ego-motion estimation and Relative pose estimation. For ego-motion, each robot maintains a bounded keyframe memory Mt i and estimates its local pose Xt i by processing the current observation along with this memory. Incremental ego-motion is recovered by composing the corresponding local poses, e.g., Tt i,t-1 i = (Xt i)−1Xt i-1 i.
Metrically Scaled Ego-Motion Estimation
To estimate the ego-motion Tt i,t-1 i = (Xt it)−1Xt(t-1) it, the framework adopts VGGT [15] as the backbone. The memory is updated using both temporal spacing and descriptor-based viewpoint change criteria to select new keyframes:
c ti
denote the cosine similarity between the current descriptor and that of the most recent local keyframe, and ∆t i denote their index gap. The current observation is selected as a new keyframe when (c ti ∆k) ∨ c ti < τk, where τm selects frames with moderate viewpoint change after sufficient temporal spacing, and the stricter threshold τk immediately accepts frames with large viewpoint change.
The scale ambiguity inherent in monocular vision is resolved by aligning the predicted depth with metric depth from Metric3Dv2 [20]. The scale factor s(Mt i) is computed as the median over all pixels across all keyframes in memory: s(Mt i) = 1/√K medianp D∗ k(p)/Dk(p). This scale factor is then applied to recover the metrically scaled ego-motion: Tt i,t-1 i = Si · (Xt it)−1Xt(t-1) it, where Si = diag(si, si, si, 1).
Communication-Efficient Relative Pose Estimation
To estimate the relative pose Ttij between robot ri and rj when visual overlap exists at time t, the direct estimation is performed using vision foundation models: Tij Di Dj = fvggt(Ii, Ij) (6). However, to handle intermittent or missing visual overlap, CERPE uses descriptor similarity as an online proxy for sufficient visual overlap.
-
A dense feature map Fi = fenc vggt(Ii) is extracted from the input image Ii.
-
This feature map is compressed into a compact global descriptor di = ψ(Fi) ∈ R D using the visual aggregation head proposed in SALAD [19], which uses Sinkhorn-based optimal transport to aggregate local patch features and a learned dustbin mechanism to discard non-informative regions.
-
Robots frequently exchange these fixed-size descriptors (mt i = (i, Xt i, d ti)).
-
A pair with cosine similarity score = d⊤ idj / (di dj) ≥ τ is treated as a candidate overlap and triggers raw-observation exchange for direct inter-robot pose estimation.
Handling Non-Overlapping Encounters
If the descriptor gate does not activate, the ego robot does not request the raw observation, thereby avoiding unnecessary communication and computation.
In this non-overlapping period, the relative pose is propagated using a previously established relative pose estimate and each robot’s accumulated ego-motion. The propagation formula is Ttij = (Xt j)−1Xt-k j Tt-k ij (Xt-k i)−1Xt i, where k denotes the most recent time step when overlapping observations are available. This mechanism maintains relative pose estimates even in the absence of visual overlap by using "metrically scaled ego-motion.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems by implementing the CERPE framework, along with what these improved systems will be able to do:
-
The AI system will achieve reliable relative 6-DoF pose estimation between robots even in ephemeral collaborative perception scenarios characterized by intermittent visual overlap and limited communication bandwidth.
-
The system can operate effectively in real-world, unconstrained environments (e.g., driving, indoor navigation) without relying on unreliable global references like GPS or prebuilt maps.
-
The system will maintain temporal continuity of robot trajectories during periods of non-overlapping observations by propagating the latest relative pose anchor using metrically scaled ego-motion, preventing abrupt jumps in estimated states.
-
The system can efficiently utilize large vision foundation models (like VGGT) for pose estimation while managing high communication costs by replacing raw observation exchange with fixed-size, continuously shared visual descriptors.
-
The AI system will dynamically gate high-bandwidth raw image requests based on descriptor similarity scores, ensuring that communication is only triggered when sufficient visual overlap is detected, significantly reducing network traffic and latency compared to methods that transmit full images or rely on constant communication.
-
The system can maintain metric-scale accuracy in ego-motion estimation by aligning predicted rotations/translations with metric depth references derived from a dedicated metric depth foundation model (Metric3Dv2), overcoming the inherent scale ambiguity of monocular vision models.
-
The system will be more robust to challenging visual conditions, such as large viewpoint changes and partial occlusions, by leveraging bounded memory queues that retain keyframes selected based on both temporal spacing and significant viewpoint change detection.
Sources
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving