AGT-CV: An Aerial-Ground Team Cross-View Dataset for Heterogeneous Robot Teams in Unstructured Environments

arXiv:2605.06478 · cs.RO · Submitted 2026-05-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "AGT-CV: An Aerial-Ground Team Cross-View Dataset for Heterogeneous Robot Teams in Unstructured Environments".

Rosa: Heterogeneous air-ground robot teams combine complementary sensing modalities, mobility characteristics, and spatial viewpoints that can significantly enhance perception in complex outdoor environments.

Dev: First, who's behind it and why it matters.

Paper summary: Rosa: So we're diving into this paper today about AGT-CV: An Aerial-Ground Team Cross-View Dataset for Heterogeneous Robot Teams in Unstructured Environments. It sounds like they’re tackling a real problem out there where robots need to see things from different angles to understand complex outdoor situations.

Dev: Exactly, Rosa. The core idea seems to be that combining the sensors and viewpoints of a ground robot and an aerial robot can really boost perception when you're dealing with unstructured settings, which is something most prior research hasn't really focused on much <ref:2605.06478#pg1>.

Taro: I think the real importance here is how this addresses that lack of real-world data; it moves beyond lab simulations into actual field conditions where things get messy. It’s about seeing what happens when different sensors overlap in a way that a single robot couldn't capture.

Rosa: That makes sense, Taro. The paper claims they collected this dataset using a Clearpath Husky UGV and an Autel EVO II UAV across several very different environments like forest trails and muddy terrain <ref:2605.06478#pg0>. It seems the thesis is that this collaborative approach offers better robustness and spatial coverage than any single-platform solution, which is what they claim makes it attractive for challenging field settings <ref:2605.06478#pg1>.

Dev: From an engineering standpoint, I'm interested in how they managed to synchronize all that data from two different platforms running in the real world; getting the loop rates and managing the latency between the ground and aerial sensors must have been a huge hurdle <ref:2605.06478#pg2>.

Taro: Well, they address that by recording synchronized sensor streams alongside joystick control commands, which supports future studies on demonstration-based driving policies <ref:2605.06478#pg2>. That means the data isn't just static images; it includes the actual actions taken to get there.

Rosa: That’s really interesting, Dev. So, what are the specific claims they make about why this dataset is needed for research in autonomous systems? What gap were they trying to fill with AGT-CV?

Dev: They specifically highlight that existing perception research has mostly focused on single-robot systems in structured settings like urban roads, and AGT-CV provides the counterpoint by focusing on heterogeneous teams in unstructured environments <ref:2605.06478#pg1>. They emphasize that this combination of Unmanned Ground Vehicles and Unmanned Aerial Vehicles is particularly compelling for improving robustness <ref:2605.06478#pg1>.

Paper summary: Taro: And they show how this helps with things like collaborative traversability estimation, which is crucial when the terrain itself is uncertain <ref:2605.06478#pg2>. When the ground robot encounters something difficult, like deep mud trenches from vehicle treads, having an aerial view can give you context about the overall situation <ref:2605.06478#pg2>.

Rosa: It sounds like they’re not just collecting data for data's sake; they’re targeting specific application areas that are hard to study otherwise. They mention that this setup allows for rich cross-modal and cross-view perception <ref:2605.06478#pg0>.

Dev: I see the technical setup is pretty comprehensive, involving three dee LiDAR from the ground platform paired with RGB imagery and thermal observations from the UAV <ref:2605.06478#pg2>. That kind of multi-modal input is what makes their data set so rich for cross-view perception research.

Taro: The annotation pipeline they used, involving a foundation model like SAM three assisted by human refinement, shows they’re thinking about how to handle the scale of labeling required for this kind of complex scene understanding <ref:2605.06478#pg2>. It’s an iterative process where the labels actually help improve the model itself.

Rosa: That iterative refinement aspect is significant because it suggests they are building a dataset that can evolve with the research needs, rather than just a static collection of images <ref:2605.06478#pg2>. It really speaks to creating something useful for long-term development in this area.

Dev: And concerning the operational aspect, Rosa, how long were these robots actually running in the field? I need to know if this is something that holds up under sustained operation outside of a controlled test bed <ref:2605.06478#pg0>.

Taro: The paper mentions they collected over thirteen thousand synchronized frames across approximately twenty-nine minutes of operation, which gives us a solid measure of real-world endurance <ref:2605.06478#pg2>. That duration is important because it shows the system's ability to maintain data integrity during a significant period in diverse conditions.

Rosa: Twenty-nine minutes sounds like a substantial amount of time for field testing, especially across those varied terrains; I wonder if they encountered any major failures or unexpected environmental challenges during that run <ref:2605.06478#pg2>.

Dev: The synchronization process itself is key here; they refined the UGV trajectory by aligning LiDAR odometry with GPS measurements using KISS-ICP to get a globally consistent ground reference <ref:2605.06478#pg1>. That level of spatial alignment is essential for making sense of the cross-view data later on.

Paper summary: Taro: The methodology for aligning the aerial and ground streams, using a gradient-domain matching pipeline with CLAHE and Sobel gradient magnitude extraction, was clever because it specifically targeted shared structural boundaries <ref:2605.06478#pg2>. That technique helps them create those thermal overlays to identify things like the UGV's thermal signature even when it's partially occluded.

Rosa: That sounds like a very practical approach to overcoming the occlusion problem they mentioned earlier, which is a big deal for real-world perception <ref:2605.06478#pg0>. So, what about the results or benchmark evaluations they did to prove this dataset is actually useful?

Dev: They conducted a terrain segmentation benchmark using SAM three on a subset of those eight thousand manually annotated images <ref:2605.06478#pg2>. The key finding there was that adapting the SAM three model on GA3T improved performance on both views, with the largest gains actually seen on the UAV data <ref:2605.06478#pg2>.

Taro: That result strongly suggests that GA3T captures domain characteristics that generic priors simply don't cover well, which validates the entire effort of collecting this specific type of data <ref:2605.06478#pg1>. It’s showing that the heterogeneous data is valuable precisely because it introduces new information.

Rosa: I think what this means for the broader field is that we can start training perception models on a dataset that genuinely reflects the messy reality of off-road navigation, which is something simulation often struggles to capture accurately <ref:2605.06478#pg1>. The paper really emphasizes its utility for cross-view semantic prediction and collaborative traversability estimation <ref:2605.06478#pg2>.

Dev: And the downstream applications they suggest, like path planning with synchronized UAV–UGV context, point toward a future where robots can make decisions based on a richer environmental awareness than current single-sensor systems allow <ref:2605.06478#pg2>. That level of contextual understanding is what we need for truly autonomous operation.

Taro: If this dataset helps us move toward broader collaborative scene understanding beyond the common assumptions made in existing cooperative perception research, that’s a big step forward for multi-agent systems <ref:2605.06478#pg1>. It opens up avenues where the robots can coordinate their sensing capabilities more effectively in complex scenarios.

Rosa: So, to wrap up on the findings and what these authors suggest about where this goes next, what do they say is important for future work with AGT-CV?

Dev: They point toward supporting emerging directions like learning from demonstration and visuomotor policy learning for off-road robot navigation by recording synchronized teleoperation commands with perception data <ref:2605.06478#pg2>. That connection between action and observation is a powerful training signal.

Paper summary: Taro: I think the implication is that we can start benchmarking algorithms specifically designed to handle this level of heterogeneous, cross-view context in unstructured outdoor environments <ref:2605.06478#pg1>. It provides a concrete resource for developing these kinds of algorithms.

Rosa: To summarize what we’ve covered about AGT-CV: it’s a real-world dataset combining ground and air views across difficult terrain, focusing on collaborative perception to improve robustness <ref:2605.06478#pg0>. It shows how domain-specific data can significantly boost model performance when dealing with complex outdoor environments <ref:2605.06478#pg2>.

Dev: And the technical aspects, like the synchronization and alignment techniques, are what make it a viable resource for engineers working on low-latency perception systems <ref:2605.06478#pg1>. We have to be careful about those processing stages to ensure reliable operation in real-time applications.

Taro: Ultimately, this paper provides the necessary data foundation for moving autonomous navigation research into more complex, multi-robot collaborative settings that operate in genuinely unstructured outdoor conditions <ref:2605.06478#pg1>. It’s a resource for testing those advanced coordination algorithms.

Rosa: It sounds like AGT-CV is setting a new standard for what we consider a rich dataset when it comes to heterogeneous aerial and ground sensing, moving us closer to systems that can truly perceive the world collaboratively in the field <ref:2605.06478#pg0>.

Dev: Yeah, the combination of three dee LiDAR geometry from the ground and thermal/RGB context from above is a specific capability that’s hard to get otherwise, which makes this dataset very targeted for those kinds of applications <ref:2605.06478#pg2>.

Taro: It’s exciting because it directly addresses the challenges of real-world environmental uncertainty by providing the complementary views needed for robust decision-making in those uncertain settings <ref:2605.06478#pg1>.

Rosa: So, if you're listening and want to see how this data translates into actual navigation systems, you can look into the downstream tasks they propose, like path planning with that synchronized context <ref:2605.06478#pg2>.

Dev: Exactly, and remember that because of the synchronization requirements mentioned in the paper, any system built on this will need to be very efficient at handling those temporal and spatial alignments <ref:2605.06478#pg1>.

Taro: It’s a resource for developing and benchmarking algorithms for heterogeneous robot teams operating in challenging outdoor environments, which is where the real testing happens <ref:2605.06478#pg1>.

Rosa: That’s a lot of exciting work on the AGT-CV dataset today, really showing how combining different robot types can provide such a significant enhancement to perception in complex outdoor situations <ref:2605.06478#pg0>.

Conclusion: Rosa: So we're wrapping up our discussion on AGT-CV, which is this new dataset for aerial and ground robot teams in unstructured settings <ref:2605.06478#pg1>.

Dev: Yeah, Rosa, it's a comprehensive resource because it brings together those very different sensing modalities we talked about earlier <ref:2605.06478#pg2>.

Taro: I think the authors really hammered home that this isn't just another collection of images; it’s a structured way to look at how robots coordinate perception in tough situations <ref:2605.06478#pg1>.

Rosa: Exactly, Taro, and thinking about the title, "AGT-CV: An Aerial-Ground Team Cross-View Dataset for Heterogeneous Robot Teams in Unstructured Environments," it really sums up what they've done <ref:2605.06478#pg1>.

Dev: It highlights the core challenge they set out to solve, which is getting different robot types to share a common understanding of the environment <ref:2605.06478#pg2>.

Taro: And the authors clearly show how this helps move us beyond just single-robot perception by providing that cross-view perspective we need for complex navigation <ref:2605.06478#pg1>.

Rosa: I'm really excited about what this means for the practical application of robotics because it’s all about building systems that can actually function reliably when conditions are messy and unpredictable <ref:2605.06478#pg1>.

Dev: That reliability is key, Rosa, and the synchronization techniques they used to align the data streams suggest a path toward more robust real-time perception systems <ref:2605.06478#pg1>.

Taro: And I think their focus on things like collaborative traversability estimation shows that this dataset has real potential for autonomy research when the terrain itself is uncertain <ref:2605.06478#pg2>.

Rosa: So, this paper really lays the groundwork for a future where we can test more sophisticated multi-robot coordination algorithms in real-world conditions <ref:2605.06478#pg1>.

Dev: And if they're successful in providing data that captures these specific environmental challenges, it could significantly impact how we design collaborative perception stacks across different robot platforms <ref:2605.06478#pg2>.

Taro: It really opens up new avenues for research into broader collaborative scene understanding that goes beyond what we currently see in most cooperative perception studies <ref:2605.06478#pg1>.

Rosa: So, this paper isn't just a data release; it's a framework for testing how heterogeneous robot teams can genuinely perceive and act together in the wild <ref:2605.06478#pg1>.

Dev: And I’m keen to see how their methods for handling those complex alignments translate into low-latency, reliable systems that can operate in those real-world scenarios we discussed earlier <ref:2605.06478#pg1>.

Drexel University

cs.RO

Submitted: 2026-05-07

Updated: 2026-10-02

Comments: CDEL -- ECCV 2026 Workshops

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 84/100

The gist: Heterogeneous air-ground robot teams combine complementary sensing modalities, mobility characteristics, and spatial viewpoints that can significantly enhance perception in complex outdoor

Key concepts

Heterogeneous Air-Ground Robot Teams
This refers to using two different types of robots—one on the ground and one in the air—working together. They each have unique sensors (like LiDAR and cameras) and movement capabilities, which provides a richer set of information than a single robot could gather alone.
Cross-View Perception
This is the ability to understand an environment by looking at it from multiple angles simultaneously. In this dataset, the ground robot sees terrain details directly, while the aerial robot sees overhead context or occlusion-free views, allowing for a more complete picture of what's happening.
Thermal/Infrared Observations
These are images that capture heat signatures rather than visible light. This sensor is crucial because it helps identify objects like the ground robot even when they are hidden behind obstacles like trees or in shadows, providing valuable information for occlusion-aware perception.

Terminology

Summary

Heterogeneous air-ground robot teams combine complementary sensing modalities, mobility characteristics, and spatial viewpoints that can significantly enhance perception in complex outdoor environments. The GA3T dataset addresses a critical gap in research by providing real-world multi-robot collaborative perception data collected using a ground robot and an aerial robot across diverse unstructured settings.

The gist: GA3T is a real-world heterogeneous air-ground multi-robot dataset for collaborative perception in diverse off-road unstructured environments.

Dataset Composition and Collection

GA3T is designed to support research on cross-view perception, air-ground viewpoint fusion, traversability estimation, and collaborative scene understanding in real off-road environments. The dataset was collected using a Clearpath Husky UGV equipped with 3D LiDAR, stereo camera, IMU, and GPS data alongside an Autel EVO II UAV carrying RGB imagery and thermal/infrared observations. The collection spanned four unique environments: forest trails, rocky paths, muddy terrain, snow piles, and grass-covered fields. This setup allows the ground platform to provide 3D LiDAR, while the aerial platform contributes RGB imagery and thermal/infrared observations, enabling a rich cross-modal and cross-view perception.

Data Characteristics and Uniqueness

A unique aspect of the GA3T dataset is its collection period during early spring, which allowed sparse tree canopies to permit the aerial robot to partially observe the ground robot and terrain through the trees, allowing for occlusion-aware collaborative perception. The dataset includes over 13,000 synchronized frames across approximately 29 minutes of operation and features both SAM 3-based zero-shot segmentation and over 8,000 manually labeled images. The environments captured include uneven terrain, dense vegetation, and substantial uncertainty in traversability, specifically including scenarios like deep mud trenches from vehicle treads.

Sensor Fusion and Data Processing

The dataset records synchronized sensor streams, including joystick control commands alongside perception data, which supports future study of demonstration-based driving policies. To align the aerial and ground data streams in both time and space, several processing stages are employed:

  1. The UGV trajectory is refined by aligning LiDAR odometry with GPS measurements using KISS-ICP [21] to obtain a globally consistent ground reference.

  2. UAV RGB and thermal image streams are aligned using a gradient-domain matching pipeline, applying CLAHE and Sobel gradient magnitude extraction to emphasize shared structural boundaries. This alignment enables thermal overlays such as white-hot and inferno renderings to help identify the UGV thermal signature under occlusion.

Annotation Pipeline

Semantic segmentation annotations are generated through a foundation-model-assisted human-in-the-loop pipeline. This process involves:

  1. Generating mask proposals using SAM 3, which is noted for its ability to generate high-quality segmentation masks and support of interactive refinement.

  2. Annotators primarily assign semantic classes and correct local mask errors, which substantially reduces annotation effort compared to drawing pixel-accurate segmentations from scratch.

  3. The refined labels are used to further train an updated SAM3 model adapted to the off-road domain, creating an iterative process that improves quality.

Benchmark Evaluation

To evaluate the dataset's utility, a terrain segmentation benchmark was conducted using the SAM3 foundation model on a subset of 8,000 manually annotated images. The results compared zero-shot SAM3 baseline models against models fine-tuned on GA3T using a two-stage training strategy (head-only adaptation and backbone fine-tuning). The discussion indicated that adaptation on GA3T improves performance on both views, with the largest gains on the UAV data, suggesting the dataset captures domain characteristics not sufficiently covered by generic priors. The overall score for class-conditioned mask IoU was reported across UGV View, UAV View, and Overall metrics.

Downstream Applications

GA3T is designed to support concrete downstream tasks enabled by heterogeneous air-ground sensing, including:

cross-view semantic prediction

collaborative traversability estimation

path planning with synchronized UAV–UGV context

broader collaborative scene understanding beyond the BEV assumptions common in existing cooperative perception research.

Furthermore, by recording synchronized teleoperation commands alongside perception data, the dataset supports emerging directions such as learning from demonstration and visuomotor policy learning for off-road robot navigation. GA3T provides a resource for developing and benchmarking collaborative perception algorithms for heterogeneous robot teams operating in challenging unstructured outdoor environments.

Conclusion

GA3T is a comprehensive collaborative perception dataset tailored for UAVUGV teams navigating diverse off-road environments, offering complementary aerial and ground viewpoints with 3D LiDAR, stereo camera, GPS, and RGB+IR imagery. It stands out as a comprehensive resource for future work in heterogeneous multi-robot coordination.

Improvements for AI systems

Here are specific improvements to AI systems that can be made by leveraging the GA3T dataset and methodology:

  1. Improve cross-view semantic segmentation accuracy for off-road environments by training models on GA3T data. The improved system will achieve superior performance in distinguishing complex, cluttered features (like mud boundaries, vegetation edges, and snow patches) when fusing ground-level LiDAR/stereo data with aerial RGB/thermal imagery.

  2. Develop a collaborative traversability estimation system that integrates air-ground viewpoints for real-time path planning. The improved system will be able to predict the traversability of a terrain segment by simultaneously analyzing surface characteristics from both the ground robot's close-range LiDAR and the UAV's broader context, leading to safer and more robust off-road navigation policies.

  3. Enhance perception robustness under partial visibility through occlusion-aware modeling. The improved system will be capable of maintaining accurate scene understanding even when the aerial view is partially occluded by sparse tree canopies (as captured in the early-spring collection period), by effectively fusing complementary data streams to infer hidden ground structures.

  4. Create a perception-conditioned visuomotor learning system for heterogeneous robot teams. The improved system will learn optimal control policies for off-road navigation directly from synchronized teleoperation commands and fused multimodal sensor streams, enabling robots to adapt their driving behavior based on real-time context provided by the complementary air-ground views.

  5. Improve pose estimation accuracy in GPS-denied or challenging outdoor settings using robust fusion techniques. The improved system will utilize the methodology of fusing KISS-ICP LiDAR odometry with GPS measurements for ground platform localization, resulting in more accurate trajectory refinement and better spatial matching between aerial and ground observations during collaborative tasks.

  6. Develop a scalable, foundation-model-assisted human-in-the-loop annotation pipeline for off-road data labeling. The improved system will allow researchers to efficiently generate high-quality semantic segmentation masks (using SAM3) by leveraging automated pre-labeling followed by expert correction, significantly reducing the manual effort required to build large, domain-specific perception datasets for challenging unstructured environments.

Sources

Related papers