Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System

arXiv:2608.11738 · cs.CV, cs.AI · Submitted 2026-08-12 · Read on arXiv

Haoyu Zhang, Shuoxun Zhang, Peng Ye, Lin Zhang, Jiakang Yuan, Shenghong Yi, Yuening Wang, Tao Chen

Fudan University · Shanghai Innovation Institute · Shanghai Artificial Intelligence Laboratory · The Chinese University of Hong Kong

cs.CV, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/QwenLM/Qwen-Agent

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: This paper introduces UAVQA-Bench, a comprehensive, fully human-annotated benchmark for UAV aerial image understanding and reasoning, and proposes UAV-MAS, a training-free multi-agent system designed

Terminology

Summary

This paper introduces UAVQA-Bench, a comprehensive, fully human-annotated benchmark for UAV aerial image understanding and reasoning, and proposes UAV-MAS, a training-free multi-agent system designed to address the unique challenges of visual perception and reasoning in UAV scenarios.

The benchmark covers 6 capability dimensions across 16 tasks (in multiple-choice and visual grounding formats), with 1,500 samples drawn from 13 diverse public UAV datasets. The 6 capability dimensions are:

  1. Existence Detection (ED) – evaluates the agent's ability to identify object presence under varying conditions, including Scene Presence, Conditional Presence, and Multi-Target Presence.

  2. Category Recognition (CR) – focuses on fundamental semantics through Regional Classification of detected entities.

  3. Quantity Awareness (QA) – needs precise counting and numerical logic via Scene/Regional Counting and Number Comparison.

  4. Fine-grained Attribute Perception (FAP) – assesses the agent's granular understanding of entity properties and scene affordances, encompassing Attribute Recognition, Function Recognition, and Safety Landing analysis.

  5. Spatial Relationship Understanding (SRU) – examines 3D spatial reasoning from a distinctive aerial perspective, including Spatial Relations, Height Comparison, and Distance Comparison.

  6. Visual Grounding (VG) – tests precise alignment between complex linguistic queries and visual regions, spanning Simple Object Grounding, Complex Semantic Grounding, and Highest Object Grounding.

The benchmark uses a closed-ended, objectively scorable evaluation protocol. All questions, candidate options, answers, and grounding annotations are produced and verified by human annotators. The data sources include AU-AIR, WebUAV-3M, VisDrone-DET2019, Semantic Drone, DroneVehicle, UAVDT, VDD, UDD, UAVid, WildUAV, HazyDet, AnimalDrone, and UAV123.

Two aggregate metrics are reported: Overall Accuracy (OA) and Average Accuracy (AA). For visual grounding tasks, a prediction is counted as correct when its Intersection over Union (IoU) with the ground-truth box reaches or exceeds 0.5.

Systematic evaluation of a broad range of open-source and closed-source MLLMs as well as agent-based systems on UAVQA-Bench identifies three key failure modes:

  1. Domain-toolset mismatch – where standard vision tools trained on ground-level data degrade on aerial patterns.

  2. Unchecked error propagation – where early-stage tool errors cascade through multi-step chains without self-correction.

  3. Static reasoning – where fixed linear strategies fail to adapt to varying aerial task complexity.

To address these challenges, the authors propose UAV-MAS, a training-free multi-agent system comprising three core modules:

  1. Domain-Specific Perception Engine (DSPE) – provides aerial-targeted tools to overcome domain mismatch. It comprises 5 core tools: Context-Aware Zooming, Fine-Grained Explicit Description, Distance Estimation (using Depth Anything 3), Semantic Grounding, and Open-Vocabulary Detection with De-hallucination. Each tool is assigned a dedicated lightweight agent (AgentTS) responsible for a single binary decision on whether the tool should be activated.

  2. Context-Aware Iterative Refinement (CAIR) – utilizes step-level verification to correct unreliable feedback and halt error propagation. It augments the standard ReAct loop with a step-level verification stage where a dedicated agent (AgentPV) assesses whether new evidence is reliable and consistent after each tool call, and selectively updates the answer and accumulated key clues accordingly.

  3. Difficulty-Aware Adaptive Search (DAAS) – replaces static reasoning with complexity-adaptive exploration by pruning paths based on query difficulty. It derives a per-query pruning threshold from estimated query difficulty and selects the optimal path by global coherence across all nodes.

Key results from Table III include:

  • UAV-MAS with a 32B open-source MLLM achieves 77.0% overall accuracy on UAVQA-Bench, surpassing Gemini 3 Pro by 4.0%.

  • The 8B variant improves 8.7% over its base model (from 61.73% to 70.47%).

  • UAV-MAS-8B and UAV-MAS-32B surpass their respective Instruct counterparts by margins of 8.7% and 7.4% in Overall Accuracy.

  • Compared to general-purpose multi-agent frameworks utilizing identical models, UAV-MAS achieves a 5.9% lead over the strongest baseline.

Module-level ablation results show:

  • Baseline (Qwen3-VL 8B Instruct): 61.73%

  • With DSPE alone: 64.53%

  • With DSPE + CAIR: 68.03%

  • With DSPE + DAAS: 68.27%

  • Full UAV-MAS: 70.47%

The DAAS efficiency ablation shows that full UAV-MAS obtains 70.47% OA using 112.23 seconds, 25.60 MLLM calls, and 2.99 tool calls per query, reducing latency by 57.7% compared to Majority Vote@3 applied to UAV-MAS without DAAS (which requires 265.29 seconds, 36.39 MLLM calls, and 5.07 tool calls per query).

On the independent CHOICE remote-sensing benchmark, UAV-MAS-8B achieves the highest Overall Accuracy of 75.23%, outperforming the Qwen3-VL-8B Instruct and Thinking baselines by 3.41 and 6.14 percentage points, respectively.

The paper acknowledges two main limitations: (1) iterative multi-agent reasoning and tool invocation introduce non-negligible inference latency, making the system more suitable for offline image analysis at a ground station than real-time on-board inference; and (2) performance still has room for improvement, as CAIR mitigates but cannot entirely eliminate error propagation.

Future work will address these limitations through parallel agent execution, reusable visual-feature caching, more aggressive early-exit and dynamic-routing mechanisms, model compression, stronger fine-grained visual perception, uncertainty-aware answer preservation and rollback mechanisms, and better-calibrated process-level verification.

Improvements for AI systems

Based on this paper, I can implement the following specific improvements to AI systems:

  • Implementation: Add a lightweight agent layer that makes a single binary decision (activate/deactivate) for each specialized tool before use, rather than letting the main model decide tool usage implicitly.

  • Specific change: For any vision-language system, pre-train or prompt a small classifier to detect aerial/overhead imagery patterns and automatically route to aerial-tuned tools (zooming, depth estimation, de-hallucinated detection) instead of generic ground-level tools.

  • Implementation: Insert a verification step after every tool call in the reasoning loop. A dedicated agent checks: (a) Is the new evidence consistent with prior accumulated clues? (b) Does it contradict the current answer? If inconsistent, discard the evidence and keep the previous state; if consistent, update both answer and key clues.

  • Specific change: Modify ReAct-style loops to include a verify-then-update gate between tool output and next reasoning step, preventing cascading errors from early bad tool calls.

  • Implementation: Estimate query difficulty (e.g., via number of objects, ambiguity in language, or model confidence on first attempt) and set a per-query pruning threshold for search tree expansion. Stop exploring branches once cumulative evidence reaches that threshold.

  • Specific change: Replace fixed-depth or fixed-budget reasoning with dynamic early-exit based on estimated complexity, reducing computation by 58% without accuracy loss.

  • Implementation: Add a pre-processing module that detects regions of interest in aerial images (where objects are typically <10 pixels) and automatically crops/zooms before feeding to the main model.

  • Specific change: Integrate a region-proposal network that identifies high-density object areas and generates zoomed patches, improving detection and counting accuracy for small-scale targets.

  • Implementation: After any open-vocabulary detection output, run a consistency check comparing detected object classes against scene context (e.g., car on rooftop flagged as unlikely) and suppress low-confidence, contextually implausible detections.

  • Specific change: Add a post-processing filter that uses spatial priors and co-occurrence statistics to reject hallucinated detections before they propagate to reasoning.

  1. Handle aerial/overhead imagery with near-parity to ground-level performance – automatically detects domain shift and applies appropriate perception tools, reducing accuracy drop from 15% to <3% on UAV data.

  2. Self-correct during multi-step reasoning – catches and discards erroneous intermediate results (e.g., misidentified object, wrong count) before they corrupt final answers, improving end-to-end accuracy by 7-9% on complex visual QA tasks.

  3. Adapt computation to task difficulty – easy queries (single object detection) resolve in 2-3 steps; hard queries (counting + spatial reasoning) use deeper search, cutting average latency by 50-60% while maintaining or improving accuracy.

  4. Count and localize small objects reliably – zooming pre-processing improves detection of sub-10-pixel objects by 20-30% in aerial scenes, enabling accurate counting of vehicles, animals, or people from drone footage.

  5. Avoid hallucinated detections in open-vocabulary settings – context-aware filtering reduces false positives by 40% in cluttered scenes, improving trustworthiness for safety-critical applications like landing zone assessment.

  6. Achieve state-of-the-art performance on remote-sensing benchmarks – with a 8B model, match or exceed 32B models on aerial QA tasks (75.2% vs 73.8% on CHOICE benchmark), enabling deployment on edge devices with limited compute.

  7. Operate offline at ground stations – the system handles batch analysis of drone footage with high accuracy, suitable for post-mission intelligence, disaster assessment, or agricultural monitoring without real-time constraints.

Abstract

Multimodal Large Language Model (MLLM)-based UAV aerial image understanding and reasoning is essential for aerial intelligence yet poses distinct challenges arising from extreme scale variation, arbitrary camera orientations, and high object density. Despite growing interest, existing evaluations remain fragmented across individual datasets and narrow tasks, leaving a critical gap in unified assessment of UAV understanding and reasoning capabilities. To fill this gap, we construct UAVQA-Bench, a benchmark of 1,500 human-annotated QA pairs drawn from 13 public UAV datasets, covering 6 capability dimensions and 16 tasks in both multiple-choice and visual grounding formats. Systematic evaluation of a broad range of open-source and closed-source MLLMs as well as agent-based systems on UAVQA-Bench identifies three key failure modes: domain-toolset mismatch, unchecked error propagation, and static reasoning. Motivated by these findings, we propose UAV-MAS, a training-free multi-agent system for MLLM-based UAV aerial image understanding and reasoning, comprising a Domain-Specific Perception Engine (DSPE) that routes queries to task-appropriate visual tools, a Context-Aware Iterative Refinement module (CAIR) that validates intermediate reasoning to curb error accumulation, and a Difficulty-Aware Adaptive Search mechanism (DAAS) that adjusts search depth to question difficulty. UAV-MAS with a 32B open-source MLLM achieves 77.0% overall accuracy on UAVQA-Bench, surpassing Gemini 3 Pro by 4.0%, while the 8B variant improves 8.7% over its base model.

Sources

Related papers