Online Planning for Sparse Ground Target Search from a High-Altitude UAV under Partial Observability

arXiv:2610.01067 · cs.RO · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Online Planning for Sparse Ground Target Search from a High-Altitude UAV under Partial Observability".

Dev: Unmanned aerial vehicles (UAVs) searching for sparse ground targets from high altitudes face a unique challenge when targets are small and unobservable, necessitating novel sensing strategies.

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: Let's talk about the paper "Online Planning for Sparse Ground Target Search from a High-Altitude UAV under Partial Observability" and who put it together. Ashik E Rasul is one of the main contributors, working out of Tennessee Technological University in the USA, alongside Hyung-Jin Yoon from the same department.

Dev: It’s interesting to see researchers from a single institution tackle this kind of complex problem; usually, you see different teams collaborating on these high-level autonomy challenges.

Taro: The focus on UAVs searching for sparse ground targets is relevant because that's exactly where we need robust systems—finding something tiny against a huge background is hard for standard sensors.

Rosa: The paper tackles this by framing the search as a partially observable Markov decision process, which is a formal way to describe making sequential decisions under uncertainty about the true state of the target.

Dev: That sounds mathematically sound, but I wonder if translating that into something that runs reliably on actual flight hardware without introducing too much computational drag is going to be a hurdle for implementation.

The paper's summary: Rosa: In essence, the paper summarizes their approach by defining the search space using windows parameterized by coordinates and side lengths, and then modeling observation uncertainty using a Beta-Bernoulli conjugate prior based on whether the target is present or not.

Dev: That modeling of uncertainty is key; if they can properly quantify how much their sensors might be fooled by scale variations, that should give them a much better picture of what they are actually seeing.

Taro: The summary also highlights using Partially Observable Monte Carlo Planning, or POMCP, to solve this process sequentially and the idea of deploying selective ensembles based on how concentrated the belief state is.

Rosa: That selective ensemble deployment is where they try to manage computational load; they decide when it's worth using a more complex set of models versus just one simple model.

Dev: So, instead of always running a heavy detection system, the system scales its computational effort based on its current confidence in the target's location; that seems like a smart way to balance performance and processing power.

The paper's improvements: Rosa: The authors suggest several key improvements to their method, primarily focusing on refining how they handle the observation model uncertainty by explicitly conditioning it on the object’s apparent scale ratio rho, which is defined as l a / l.

Dev: Conditioning the observation model uncertainty on that scale ratio is a nice addition because it directly links the camera's zoom level and field of view to how reliably they can detect something.

Taro: I see them also propose deploying an ensemble of Deep Neural Networks at a fixed scale when the belief concentration gets high, suggesting that when the system is reasonably sure, it can leverage multiple models for better prediction accuracy.

Rosa: That selective deployment mechanism is designed to be efficient; they swap out a default scan for this ensemble scan involving five models trained on different subsets of data if the belief concentration exceeds a threshold tau e.

Dev: That makes sense from an engineering standpoint; it means when the system is highly confident, it uses more resources for better results, but during early exploration, they stick to simpler operations.

Conclusion: Rosa: To wrap things up on "Online Planning for Sparse Ground Target Search from a High-Altitude UAV under Partial Observability," the paper shows that framing the PTZ operation as a POMDP allows for sequential action selection that optimizes target discovery by managing observation uncertainty through scale-conditioned priors.

Dev: From my side, I think what they've demonstrated is a framework that accounts for physical constraints and sensor limitations by making the planning process explicit about what it can and cannot observe given the UAV's altitude and camera settings.

Taro: I think the implication here for autonomy is that we can develop search policies that are explicitly designed to be adaptive, not just reactive, especially when dealing with very sparse targets where traditional methods fail completely.

Rosa: That’s a big picture idea; it moves the field toward systems that can intelligently manage their sensing capabilities rather than just blindly sweeping the area.

Dev: It's certainly a lot of work to get that kind of planning loop running smoothly, but if they can prove its effectiveness in simulation and then deploy it reliably, it could be really useful for real-world scenarios where detection is tricky.

Taro: I think the long-term impact is in creating search agents capable of operating in complex, partially observable environments where the initial assumptions about target visibility are constantly being tested and refined during the mission.

Ashik E Rasul, Hyung-Jin Yoon

Department of Mechanical and Nuclear Engineering, Tennessee Technological University

cs.RO

Submitted: 2026-10-01

Updated: 2026-10-01

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 81/100

The gist: Unmanned aerial vehicles (UAVs) searching for sparse ground targets from high altitudes face a unique challenge when targets are small and unobservable, necessitating novel sensing strategies.

Key concepts

Partially Observable Markov Decision Process (POMDP)
A framework used when an agent doesn't know the true state of its environment, like the exact location of a small target. The agent maintains a 'belief state'—a probability distribution over all possible locations—which it updates as it gathers noisy observations. The goal is to choose actions (like moving the camera) to maximize reward despite this uncertainty.
Belief State
This represents the agent's current best guess about where the target is located. Instead of knowing for sure, the agent keeps a set of hypotheses (particles) representing various possible locations and their probabilities. This belief state is crucial because it dictates which PTZ action to take next to best reduce uncertainty and find the target.
Partially Observable Monte Carlo Planning (POMCP)
This is a planning algorithm designed specifically for POMDPs. It works by simulating many possible future scenarios starting from the current belief state. It uses an Upper Confidence Bound (UCB) criterion to select actions that balance exploring new areas with exploiting known good options, helping the agent navigate the search space effectively.
Scale Ratio ($ ho$)
This parameter describes how large or small a detected target appears on the sensor, relating its actual size to the camera's zoom level. The observation model explicitly links detection reliability to this ratio; larger targets are easier to detect reliably than very small ones, allowing the planning system to adjust its search strategy accordingly.

Terminology

Summary

Unmanned aerial vehicles (UAVs) searching for sparse ground targets from high altitudes face a unique challenge when targets are small and unobservable, necessitating novel sensing strategies. The core contribution of this work is formulating sequential exploration with Pan-Tilt-Zoom (PTZ) operation as a partially observable Markov decision process (POMDP) to optimize target discovery while explicitly modeling observation uncertainty conditioned on the target's scale.

The gist: This work formulates the sequential exploration with PTZ operation as a partially observable Markov decision process (POMDP), in which the agent maintains a belief state over the target’s true location. To solve this POMDP, we deploy partially observable Monte Carlo planning (POMCP), where we condition the sensing reliability on target object scale and deploy selective ensemble detection as an additional reasoning step.

Problem Formulation and Modeling Uncertainty

The search problem is defined by discretizing the search space using a family of windows, parameterized by top-left corner coordinates (r, c) and side length l, denoted as W(L). The state space (S), action space (A), and observation space (O) are instantiated from these windows. Crucially, the transition model is simplified to the identity: T(s′s, a) = 1[s′ = s] for all s, s′ ∈ S, a ∈ A, as the target is assumed static. The observation model uncertainty is explicitly modeled using a Beta-Bernoulli conjugate prior. This uncertainty depends on both the presence indicator of the target (bp) and its apparent scale ratio ρ, where ρ = la / l. The probability of detection is decomposed into two cases: when bp = 1, the observation follows a Gaussian distribution conditioned on the scale ratio, while when bp = 0, it follows a Bernoulli distribution.

POMDP and Planning Framework

The control policy operates directly on the belief space to sequentially actuate PTZ operations to optimize long-term objectives. The belief state (bt) is maintained as an unweighted particle set of K hypotheses and updated via rejection sampling after each observation, approximating the true Bayesian posterior. To solve this POMDP, Partially Observable Monte Carlo Planning (POMCP) is deployed. The planning process involves running NSIM simulations from the current belief node, denoted as root node, traversing the tree using a UCB action selection criterion to balance exploitation and exploration: a∗ = arg max a∈A 'V (ha) + c s ln N(h) / N(ha) for action selection.

Observation Model and Selective Ensemble Detection

The observation model is decomposed based on the presence indicator bp. When bp = 1, the observation likelihood is modeled using: Z(o s, a) = θ+(ρ) · N ([ro, co, so]; [r, c, l], Σ) o ≠ ∅ / 1 − θ+(ρ) o = ∅ for non-null observations. When bp = 0, the model uses: Z(o s, a) = θ−(ρ) Nfp o ≠ ∅ / 1 − θ−(ρ) o = ∅. To enhance reasoning when belief concentration is high, the methodology includes selective ensemble detection. This is triggered when the belief concentration exceeds a threshold τe, where POMCP-PLAN suggests a scan action covering the MAP with scale ratio ρ = 1, we swap the default scan for an ensemble scan involving five models trained on different subsets of the dataset.

Reward Structure and Termination Policy

The reward structure is decomposed into in-tree step rewards and a terminal declaration reward. The in-tree reward function is defined as: R(s, a) = (Rstep + Rdet − Rfp o ≠ ∅ / Rstep o = ∅) where "Rstep < 0 is the per-step cost encouraging minimization of detector inferences. The terminal reward, applied outside the POMCP tree when termination criteria are met, is defined as: RT(ˆst) = (R+ sˆt = s∗ / R− sˆt ≠ s∗), where R+ ≫ R− determines the terminal reward, s∗ denotes the true target state. Target declaration occurs when belief concentration exceeds a rational threshold τ: The agent declares the MAP state sˆt when belief concentration exceeds the rational threshold τ."

Validation and Results

The methodology was validated in a photorealistic simulator built in Unreal Engine under various environmental conditions and vehicle states. Experiments compared POMCP-PTZ against baselines like Fixed Scan, Exhaustive-Crop, and Coarse-to-Fine methods across different altitudes (100m, 200–350m) and camera resolutions (640 × 640).

Improvements for AI systems

Here are the specific improvements that can be made to existing AI systems by applying the methodology described in this paper, along with what these improved systems can achieve:


Improvements for AI Systems

  1. Transition from Static/Fixed-Scale Detection to Dynamic, Uncertainty-Aware Active Sensing:

  2. Integration of Belief State Planning (POMDP) into Perception Pipelines:

  3. Adaptive Detection Strategy via Selective Ensemble Deployment:

  4. Robustness to Scale and Resolution Variation in Aerial Target Search:

What the Improved AI System Can Do

The improved AI system, based on the POMCP-PTZ framework, can perform the following specific functions:

  1. Efficient Sparse Target Localization Under Partial Observability:

Specifically, it can successfully locate small ground targets (like helipads) from high altitudes where standard detection methods fail due to resolution downscaling and limited context. It achieves this by dynamically deciding the optimal Pan-Tilt-Zoom (PTZ) action—balancing wide-scan exploration for initial saliency against precise scans for reliable detection—to navigate the partially observable environment.

  1. Sequential, Information-Maximizing Search Policies:

Instead of relying on fixed scanning patterns (like Fixed Scan or Exhaustive Crop), the system can execute a sequence of actions (PTZ adjustments) that are explicitly optimized to maximize the probability of target detection within a minimum number of steps, effectively modeling human-like spatial attention and search efficiency.

  1. Adaptive Inference Resource Management:

The system intelligently manages computational load by selectively deploying computationally expensive multi-model ensembles only when the belief state is sufficiently concentrated (i.e., when the confidence threshold is met). This prevents unnecessary high-cost inference during early, uncertain exploration phases, leading to significant reduction in overall computational overhead compared to always running large ensembles.

  1. Uncalibrated Confidence Reasoning:

By explicitly modeling observation uncertainty using a Beta-Bernoulli conjugate model conditioned on the target object's apparent scale (derived from PTZ settings), the system can generate more reliable and calibrated confidence scores for detections, even when sensor resolution is severely degraded or highly variable. This allows it to distinguish between true positives and false positives with greater statistical rigor than methods that rely solely on fixed detection thresholds.

  1. Robustness Across Varying Environmental Conditions (Altitude/Resolution):

The system demonstrates superior performance in high-altitude scenarios (e.g., 350m) where traditional methods fail due to severe resolution downscaling, and it maintains competitive performance even when camera resolutions are varied, showing a resilience that fixed-slicing or coarse-to-fine methods lack.

Abstract

Unmanned aerial vehicles (UAVs) searching for ground targets from a high altitude face a unique challenge, particularly when the target is already within the field of view but effectively unobservable because of its small apparent scale. Standard object detectors often underperform in such scenarios because of resolution downscaling and limited context. In contrast, active object search frameworks address this challenge by directing the agent to a suitable pose to gather richer visual information. However, flight regulations in urban airspace often restrict such physical movements for UAVs. As an alternative active sensing approach, the UAV can leverage the pan-tilt-zoom (PTZ) mechanism of the onboard camera to dynamically adjust its field of view and sequentially gather enhanced visual information from specific regions of interest. Once the candidate locations of the target are identified, it can deploy more expensive object detection schemes, such as an ensemble of multiple models, to get better reasoning at a fixed scale. In this work, we formulate the sequential exploration with PTZ operation as a partially observable Markov decision process (POMDP), in which the agent maintains a belief state over the target's true location. To solve the POMDP, we deploy partially observable Monte Carlo planning (POMCP), where we condition the sensing reliability on target object scale and deploy selective ensemble detection as an additional reasoning step. We validate our methodology with experiments in a photorealistic simulator under different environmental conditions and vehicle states, showing detection of ground targets at variable scales with significantly fewer steps and minimal dependence on sensor resolution compared to baseline methods.

Sources

Related papers