G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection

arXiv:2607.19942 · cs.CV, cs.AI · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection".

Jane: The paper was written by Yechan Kim, JongHyun Park, Dongho Yoon, Namhoon Jung and Moongu Jeon from LIG Defense & Aerospace and Gwangju Institute of Science and Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. Today we’re looking at a paper that’s got a mouthful of a title: “G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection.” Jane, I’m going to need you to unpack that acronym for our listeners.

Jane: Happy to, Tom. So G-MAD stands for a data Generator for Multi-view Aerial object Detection, and they built it on top of Arma3, which is a military simulation game. The big idea is that they’re using a video game engine to create training data for AI systems that need to spot objects from the sky, like tanks or planes, using both regular visible light cameras and thermal cameras.

Lu: And that’s the part that really caught my attention. Getting real-world aerial footage with perfectly aligned visible and thermal images is incredibly expensive. You need to fly drones with specialized sensors, coordinate the timing, and then pay people to label every single object in every frame. This paper basically says, why not just generate all of that inside a game where the computer already knows exactly where everything is?

Tom: Right, and that’s the clever part. The game engine gives them perfect ground truth for free. If you place a tank at a certain coordinate, the engine knows its exact three dee position, so the annotation is mathematically perfect, not drawn by hand.

Meng: But hold on, Tom. I’ve seen synthetic data pipelines before. The usual problem is that the models trained on game graphics don’t transfer well to the real world. The colors are off, the textures are too clean, the shadows don’t behave right. What makes this framework different?

Jane: That’s a fair challenge, Meng. And the paper actually addresses that head-on. They don’t just dump the raw game images into a model. They use a style-transfer technique to make the synthetic images look more like real satellite or drone photos. And they also test whether pretraining on their synthetic data helps when you fine-tune on real-world benchmarks. The numbers show a solid improvement, like a six-point jump in mean average precision on one of the real-world visible datasets.

Lu: Which is a big deal, because that suggests the synthetic data is capturing the geometry and the structure of the problem, not just the pixels. The model is learning what a tank looks like from different angles, and that knowledge transfers.

Tom: So we’re not just talking about a fun tech demo. This could actually make aerial surveillance and disaster response systems more robust, and cheaper to build. Jane, what’s the catch?

Jane: The catch is that Arma3 is a military game, so the assets are mostly military vehicles. The dataset they release, AMOD, is specialized for that. But the framework itself is modular, so you could plug in civilian assets if you wanted to.

Meng: So the tool is general, even if the first dataset isn’t. That’s useful.

Tom: And that’s the hook for our next segment. We’re going to dig into how they actually built this pipeline and what makes the multi-view part so special. Stay with us.

Summary: Tom: So we’ve established that G-MAD uses Arma3 to generate synthetic aerial data. Jane, walk us through the actual pipeline, because I found the architecture surprisingly elegant.

Jane: It is. The whole thing is broken into three stages. First, you specify a scenario. That means you tell the system things like how many objects you want, what categories, what weather, what time of day, and which map to use. Second, the system samples that scenario and automatically flies virtual cameras around the area of interest. And third, it uses the game engine’s internal geometry to draw the bounding boxes automatically.

Lu: And the key insight is that they separate the scenario from the scene. A scenario is the fixed setup, like a tank sitting in a field at two PM. A scene is a specific camera angle looking at that tank. So from one scenario, they can generate six different scenes by moving the camera from straight overhead to a fifty-degree angle. That gives you synchronized multi-view data of the exact same object.

Meng: Synchronized is the word that matters to me. In real life, if you want to see a tank from six different angles at the same moment, you need six drones flying in perfect formation. That’s nearly impossible to coordinate. In the game, it’s just a script. You get perfect alignment for free.

Tom: And the annotation quality is perfect too. They’re not using image differencing like some older game-based pipelines. They’re querying the engine for the three dee bounding box of the object and projecting it onto the 2D image plane. That gives them both horizontal boxes and rotated boxes, which are important for aerial imagery because objects don’t always sit upright.

Jane: Exactly. And they also handle occlusion automatically. If a hill blocks the camera’s line of sight to the tank, the system knows that and filters out the annotation. That reduces label noise, which is a huge problem in manually labeled datasets.

Lu: What I find exciting is the control this gives researchers. You can systematically vary the viewing angle and see exactly how a detector’s performance degrades as you move from nadir to off-nadir. The paper shows that a model trained on all six angles generalizes much better than a model trained on just one angle, even when you control for the amount of data.

Meng: So the multi-view training isn’t just adding more data, it’s adding diversity that matters. That’s a strong result.

Tom: And it’s a result that would be really hard to get with real data, because you’d need to fly six drones simultaneously. Jane, what about the practical side? How hard is this to actually run?

Jane: That’s where they made some smart engineering choices. The framework runs entirely through Arma3’s scripting language, SQF, so it doesn’t need any GUI automation. You can launch it headlessly, which makes it scalable. And it cleans up temporary files automatically, which sounds minor but is a huge quality-of-life improvement for researchers running thousands of scenarios.

Meng: Headless execution is a big deal. The older Arma3 pipelines I’ve seen rely on clicking buttons through the UI, which is fragile and slow. This is a proper engineering solution.

Tom: So we have a scalable, controllable, and precise data generator. But the real question is, does it actually work on real-world images? That’s what we’re going to tackle in the next segment.

Improvements: Tom: We’re back with G-MAD, and I want to get into the experimental results, because that’s where the paper really proves its value. Jane, what did they actually test?

Jane: They ran two main experiments. The first was about multi-view learning. They trained a detector on images from a single viewing angle, say ten degrees, and then tested it on all the other angles. The performance dropped significantly, especially at the more extreme angles. But when they trained on all six angles together, the average performance jumped by more than six points in mean average precision.

Meng: That’s a clean demonstration that viewpoint diversity is a real factor in aerial detection. It’s not just about having more images, it’s about having images from different perspectives.

Lu: And the second experiment is where it gets really interesting for the broader community. They took their synthetic data, translated it to look like real satellite imagery, and then used it to pretrain a detector. Then they fine-tuned that detector on a real-world benchmark called DIOR-R. The pretraining gave them a nearly seven-point improvement over starting from ImageNet weights.

Jane: And they did the same thing for thermal imagery. They translated their synthetic thermal images to look like real thermal drone footage, pretrained on that, and then fine-tuned on a real thermal benchmark called HIT-UAV. That gave them a two-and-a-half-point boost.

Meng: So the synthetic data is genuinely transferable. That’s the result that matters for practical deployment. But I’m curious about the thermal side. Thermal images are notoriously hard to simulate because heat signatures depend on so many factors.

Jane: That’s a great point, Meng. And the paper addresses it with a clever trick. They noticed that translating visible images to thermal is inherently ambiguous. The same visible scene could have many different thermal signatures depending on how hot the objects are. So instead of trying to make a perfect translation, they add stochasticity during pretraining. They randomly blur the boundaries of objects in the translated thermal images to simulate the edge attenuation you see in real thermal sensors.

Lu: That’s a really thoughtful approach. They’re acknowledging that the translation is one-to-many, and they’re injecting that uncertainty into the training process rather than pretending there’s a single correct answer.

Tom: And the numbers back it up. Adding that stochastic boundary smoothing gives them another point and a half of improvement on the thermal benchmark.

Meng: So the framework isn’t just about generating data, it’s about generating data that’s actually useful for real-world tasks. That’s a meaningful contribution.

Jane: And they even show a qualitative example where a detector trained only on their synthetic data can identify military vehicles in real satellite photos from news coverage of the Russia-Ukraine war. No fine-tuning at all, just straight inference.

Tom: That’s a striking demonstration. It shows the synthetic data is capturing the essential visual features of these vehicles, not just the game’s rendering style.

Lu: It does make me wonder about the ethical dimensions, though. This technology could be used for surveillance in ways that aren’t always benign.

Jane: That’s a conversation worth having, and the authors do note that the dataset is specialized for military targets. But the framework itself is neutral. You could use it to generate data for search-and-rescue or wildlife monitoring just as easily.

Tom: Let’s bring in Lalam to give us a broader perspective on what this means for the field.

Lalam: Thank you, Tom. What strikes me about G-MAD is that it democratizes access to high-quality training data. In many parts of the world, researchers don’t have the budget to fly drones with expensive thermal cameras. This framework lets anyone with a copy of Arma3 generate a large-scale, precisely annotated dataset for a fraction of the cost. That could accelerate research in countries and institutions that have been excluded from this area. And beyond that, the multi-view capability could improve systems that need to understand scenes from multiple perspectives, like disaster response coordination or environmental monitoring. The cultural impact is that we’re moving toward a world where the bottleneck isn’t data collection, but the creativity to design the right scenarios.

Tom: That’s a powerful way to frame it. We’ll carry that thought into our conclusion.

Conclusion: Tom: So let’s wrap up our discussion on G-MAD. Jane, give us the final takeaway.

Jane: The paper presents a complete framework for generating synthetic aerial detection data using the game Arma3. It solves three big problems: controlling viewpoints, aligning RGB and thermal images, and getting perfect annotations. And it does all of that automatically, without any manual labeling.

Lu: And the experimental results show that this synthetic data is genuinely useful. Pretraining on it improves performance on real-world benchmarks for both visible and thermal imagery. That’s the proof that the approach works, not just in theory but in practice.

Meng: From an engineering standpoint, I appreciate that they built it to be headless and modular. You can run it at scale, and you can customize the camera trajectories to fit your specific task. That makes it a practical tool, not just a research prototype.

Tom: And as Lalam pointed out, it lowers the barrier to entry for researchers who don’t have access to expensive drone hardware. That could have a real impact on who gets to do research in this field.

Jane: There are limitations, of course. The dataset is focused on military targets, and the game’s rendering isn’t a perfect match for real-world imagery. But the framework is designed to be extended, and the authors are clear about those constraints.

Lu: I’d say the biggest contribution is showing that a commercial game can be repurposed into a rigorous scientific instrument. That’s a pattern we’re likely to see more of as synthetic data becomes more central to computer vision.

Tom: Well said. We’ve covered the architecture, the experiments, and the implications. G-MAD is a solid contribution to the field of aerial object detection, and we’re excited to see what the community builds with it.

Jane: And with that, we’ll say goodbye to this paper and get ready to dive into the next one. Thanks for listening, everyone.

Tom: See you on the next episode.

Yechan Kim, JongHyun Park, Dongho Yoon, Namhoon Jung, Moongu Jeon

LIG Defense & Aerospace · Gwangju Institute of Science and Technology

cs.CV, cs.AI

Submitted: 2026-08-15

Updated: 2026-08-18

Comments: ACM Multimedia 2026 (Supplementary Material Included)

Code: https://github.com/cloftus96/Synthetic-Data-Generation

Project page: https://unique-chan.github.io/G-MAD-Project

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 81/100

The gist: This paper introduces G-MAD, an open-source framework that uses the game Arma3 to generate synchronized multi-view RGB-T (visible and thermal) data for aerial object detection.

Key concepts

G-MAD
G-MAD stands for a Data Generator for Multi-view Aerial object Detection. It uses the military simulation game Arma3 to create synthetic aerial training data for AI systems that need to spot objects from above, using both visible light and thermal cameras.
Multi-view Data Generation
The framework separates a fixed scenario from a scene (camera angle). This allows researchers to generate synchronized multi-view data of the exact same object from different angles automatically, solving the difficulty of coordinating multiple real drones.
Synthetic Data Transferability
The paper shows that synthetic data is genuinely useful. By translating game images to look like real satellite or drone photos and adding stochasticity during thermal simulation, pretraining on this synthetic data significantly improves performance on real-world benchmarks for both visible and thermal imagery.

Terminology

Summary

This paper introduces G-MAD, an open-source framework that uses the game Arma3 to generate synchronized multi-view RGB-T (visible and thermal) data for aerial object detection. The authors state that G-MAD addresses key limitations of real-world aerial dataset construction, including limited viewpoint control, imperfect RGB-T alignment and high annotation cost. The framework supports structured scenario specification, controllable multi-view camera placement, simultaneous visible/thermal capture, and automatic bounding box annotation using engine-level geometric metadata.

The paper describes the architecture of G-MAD as an end-to-end data generation pipeline that transforms a compact set of user-specified parameters into structured multi-modal (RGB-T) data for aerial detection. The pipeline is decomposed into three stages: (1) scenario specification, (2) scenario sampling and synchronized multi-view, multi-modal capture, and (3) automatic annotation and export. The overall pipeline is implemented in SQF, the native scripting language of Arma3, allowing direct procedural control over scene composition, environmental configuration, in-game sensor control, and data export within the simulator.

In the scenario specification stage, the proposed framework generates scenarios automatically from a compact set of high-level constraints provided by the user. These constraints include "① asset categories, ② the expected minimum and maximum numbers of objects to place, ③ weather, ④ time (day/night), ⑤ map, ⑥ camera sensor configuration (⑥-a. Altitude, ⑥-b. FOV, ⑥-c. RGB-thermal modality, ⑥-d. captured image size (width/height), ⑥-e. camera view sampling strategy per each scene), ⑦ number of scenarios."

For scenario sampling and capture, the paper explains that G-MAD formulates data generation as a two-stage sampling process over scenarios and viewpoints, expressed as s ∼ P (S constraints), v ∼ P (V s), where each sampled pair (s, v) produces a single annotated scene. This formulation decouples scenario-level variables (object arrangement, environment) from scene-level variables (camera viewpoint and sensing modality), enabling structured data generation. The framework offers two default camera motion modes: air-to-air, where the camera moves laterally while maintaining a relatively stable altitude, producing sweeping aerial observations, and air-to-land, where both altitude and lateral position vary following a parabolic trajectory, resulting in oblique views directed toward the ground. The viewpoint sampler is designed to be modular and extensible, allowing users to incorporate custom motion patterns.

For automatic annotation, G-MAD implements a direct annotation pipeline by querying object-level geometric information from the Arma3 engine and projecting it into the image plane. The paper notes that "unlike prior approaches that rely on foreground-background image differencing, our framework retrieves precise 3D bounding box information from the Arma3's engine and transforms it into 2D annotations aligned with each captured scene, ensuring high-quality labels. Both horizontal and oriented bounding boxes (HBB/OBB) are generated. The framework also incorporates automatic filtering of invalid annotations, such as objects fully occluded by other scene elements, as determined by the absence of a direct line-of-sight between the camera and the object surface."

The paper compares G-MAD with two existing open-source Arma3-based frameworks, noting that prior approaches assume at most one object per image, rely on manual or inflexible camera setup, and provide little support for research-oriented, controlled dataset construction. Additionally, existing methods depend on GUI-level automation via click-based actions to launch and operate Arma3, making the pipeline brittle to UI state changes and user interference, and their annotations are obtained through image differencing, which is cumbersome and error-prone for multi objects. In contrast, G-MAD enables constraint-aware multi-object generation, GUI-independent execution, and Arma3 engine-native annotation of both HBBs and OBBs.

Using G-MAD, the authors construct and release AMOD (Aerial Military Object Detection), a new multi-view RGB-T aerial detection benchmark for studying viewpoint robustness, RGB-T learning, and synthetic-to-real transfer. The AMOD dataset consists of synchronized RGB-T multi-view aerial observations with OBB annotations of following categories: Armored, Artillery, Helicopter, LCU, MLRS, Plane, RADAR, SAM, Self-propelled Artillery, Support Vehicle, Tank, and TEL. The dataset contains 73,920 total frames, 383.2k total instances, and 12 categories, with an image size of 1920x1440. It is described as the only RGB-T aerial detection benchmark with simultaneous multiple viewpoints.

The paper reports three main experimental findings. First, for multi-view learning, while single-view models degrade under unseen viewing angles, multi-view training significantly improves generalization, achieving a mean AP50:95 of 50.96 compared to 44.57 for the best single-view model (+6.39). Second, for pretraining for real-world detection, AMOD pretraining consistently outperforms ImageNet initialization, improving mean AP50 from 63.14 to 69.83 on DIOR-R (+6.69) and from 74.57 to 77.08 on HIT-UAV (+2.51). Third, for unseen real-world inference, the AMOD-trained detector still identifies plausible military targets on previously unseen real images from public news coverage of the Russia–Ukraine war.

The paper also notes a limitation: "Although AMOD provides a large-scale, synchronized multi-view RGB-T benchmark for aerial object detection, it is primarily specialized for military targets. This specialization stems from the use of Arma3, whose built-in assets and terrains are particularly well suited for military scenarios. However, since Arma3 supports user-defined asset integration, G-MAD can also be used to construct datasets for non-military general equipment."

Regarding licensing, the paper states that "G-MAD relies on Arma3 as an off-the-shelf rendering and scenario-execution environment. Accordingly, the framework is designed to avoid redistributing Bohemia Interactive intellectual property including Arma3 executable game files, original asset packages, or modified copies of the game. G-MAD only provides external scripts that users may run with their own legitimately obtained local installation of Arma3. The framework is released solely for non-commercial academic or prototype research purposes, and users are responsible for ensuring that their use complies with the Arma3 End User License Agreement."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with what the improved systems can do:


Improvement: Implement a training strategy that uses synchronized multi-view data (6 angles: 0°–50°) from the same underlying scene, as described in Section 3 and Table 3 of the paper.

What the improved system can do:

  • Achieve +6.39 mean AP50:95 improvement (from 44.57 to 50.96) over single-view training when evaluated across all viewing angles.

  • Maintain robust detection performance under unseen camera angles (e.g., a model trained on 0°–30° still performs well at 40°–50°).

  • Reduce performance variance across viewpoints, making the detector more reliable in real-world UAV/satellite applications where the viewing angle is not controlled.

Bottom line: The most impactful improvements are (1) multi-view training for viewpoint robustness, (2) AMOD-based pretraining with domain adaptation for real-world transfer, and (3) SBS augmentation for thermal detection. These are directly validated with quantitative gains in the paper and can be immediately integrated into existing detection pipelines (e.g., Oriented R-CNN with Swin-S in MMRotate).

Abstract

This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detection. G-MAD addresses key limitations of real-world aerial dataset construction, including limited viewpoint control, imperfect RGB-T alignment and high annotation cost. The framework supports structured scenario specification, controllable multi-view camera placement, simultaneous visible/thermal capture, and automatic bounding box annotation using engine-level geometric metadata. These capabilities enable controlled studies of viewpoint variation, multi-modal fusion, and synthetic-to-real transfer in aerial object detection. Besides, using G-MAD, we construct and release AMOD, a new large-scale multi-view aerial RGB-T object detection benchmark. The source code and the dataset are available at https://unique-chan.github.io/G-MAD-Project.

Sources

Related papers