DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation

arXiv:2608.12308 · cs.CV, cs.AI · Submitted 2026-08-12 · Read on arXiv

Yan Deng, Fei Xu

Xi'an Technological University

cs.CV, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 24 pages, 6 figures, 3 tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 50/100

The gist: DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation Abstract Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual

Terminology

Summary

DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation

Abstract

Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observability. Although recent VLA models offer a promising perception-to-action paradigm, adapting them to aerial navigation remains challenging. Policies conditioned primarily on the current observation may lose track of previously observed landmarks; single-step action prediction provides limited lookahead; and implicit termination through action generation may be unreliable.

To address these challenges, we propose DreamFly, a diffusion-based aerial VLN framework built on Dream-VLA. DreamFly first introduces a causally aligned historical memory that augments the current visual representation with evidence drawn exclusively from observations preceding the current decision step, thereby enabling temporal reasoning without leaking future information. Leveraging the bidirectional diffusion backbone, we further formulate navigation as receding-horizon diffusion planning, where the policy jointly predicts a K-step action chunk but executes only the first action before replanning from the next observation. This plan-K, execute-one strategy treats predicted future actions as auxiliary planning targets while preserving closed-loop visual feedback. Finally, we introduce LiteStop, a lightweight termination module that estimates the stop probability directly from the action logits at the initial all-mask state, thereby decoupling explicit termination from action generation.

Together, these components establish a causal closed-loop cycle of observation, memory retrieval, diffusion planning, termination assessment, execution, and memory update. Experiments on the OpenFly benchmark demonstrate consistent improvements in both seen and unseen environments. DreamFly achieves 32.04%/29.46% SR and 28.22%/23.54% SPL on the test-seen/test-unseen splits, respectively, outperforming all compared methods on both metrics while also attaining the lowest navigation error. These results highlight the effectiveness of jointly modeling historical context, future action structure, and explicit termination for aerial vision-language navigation.

Introduction

Vision-and-Language Navigation (VLN) requires embodied agents to interpret natural-language instructions and navigate to specified goals using sequentially acquired visual observations. In contrast to static visual recognition and single-step action classification, VLN is inherently a partially observable, closed-loop sequential decision problem: an action taken at the current step affects the agent's subsequent state and visual observation, while newly acquired environmental evidence updates its estimates of task progress, spatial context, and direction of travel. A reliable VLN policy therefore requires more than vision–language grounding at an isolated decision step. It must maintain temporal consistency across historical visual evidence, the current observation, planned future actions, and executed actions throughout its interaction with the environment.

Aerial vision-language navigation extends VLN to unmanned aerial vehicles (UAVs), requiring agents to navigate autonomously through large-scale 3D environments by following natural-language instructions. However, aerial VLN cannot be treated as a direct transfer of conventional ground-based VLN to a flying platform. Ground agents typically move on approximately two-dimensional traversable surfaces, whereas UAVs must jointly coordinate horizontal displacement, vertical motion, and viewpoint adjustment. Changes in altitude affect not only the UAV's spatial position but also the spatial extent of the visible scene, the apparent scale of landmarks, the level of visible detail, and obstacle visibility, thereby tightly coupling horizontal and vertical navigation decisions. Moreover, urban aerial environments often span large areas and contain complex 3D structures and substantial visual occlusion, such that a single egocentric observation captures only a limited portion of the surrounding scene. Consequently, local decision errors can compound over time. These characteristics impose stringent requirements on historical context modeling, 3D spatial reasoning, cross-modal grounding, and online error correction in aerial VLN.

Early VLN methods commonly relied on sequence-to-sequence architectures, typically implemented as recurrent policies, to autoregressively predict the next navigation action from the language instruction, current visual observation, and recurrent hidden state. With advances in vision–language pretraining and Transformer-based architectures, subsequent studies developed recurrent vision–language Transformers, history-aware representations, and explicit environmental memories to support reasoning across successive decision steps. For example, Recurrent VLN-BERT maintains a recurrent state representation, whereas HAMT employs a hierarchical Transformer to encode navigation history, thereby facilitating multimodal reasoning over extended trajectories. In aerial VLN, recent approaches have further explored multidirectional view selection, bird's-eye-view feature maps, keyframe selection, and compression of historical visual observations to mitigate restricted egocentric visibility, redundant observations, and difficulties in tracking navigation state. Together, these studies underscore the importance of historical information for navigation. However, historical modeling depends not only on how past observations are represented, but also on precisely when they become available to the policy during closed-loop interaction.

More recently, vision–language–action (VLA) models have emerged as a promising paradigm for language-conditioned end-to-end control. By formulating actions as discrete tokens or continuous values, VLA models adapt pretrained vision–language models to generate robot actions conditioned on language instructions and visual observations. Action chunking captures short-horizon temporal structure by jointly predicting a sequence of actions, while diffusion-based policies model multimodal action distributions through iterative denoising and can be deployed with receding-horizon replanning. Dream-VLA, for example, employs a diffusion Transformer to jointly predict action chunks, thereby capturing dependencies among actions over a short future horizon. Recent studies of aerial navigation have likewise explored end-to-end VLA policies, diffusion-based action modeling, and joint world–action prediction, highlighting the potential of explicitly predicting future states or action sequences for navigation in complex 3D environments. Several of these approaches seek to enhance planning by forecasting future world states or visual observations. However, how to exploit dependencies among predicted future actions over a short horizon while retaining closed-loop replanning from newly acquired observations remains underexplored.

Despite recent advances in historical modeling, 3D reasoning, and multi-step action prediction, three temporal aspects of aerial VLN remain insufficiently addressed within a unified closed-loop decision process. First, an online historical memory requires an explicit temporal boundary. Here, causal refers to temporally ordered information access rather than causal relations between variables: at decision step t, historical information is restricted to observations acquired before step t. Accordingly, the historical memory M<t is read before the current observation is written to it, preventing current-step information from entering the historical branch. Applying this read-before-write ordering during both training and deployment gives M<t a well-defined temporal boundary and ensures consistent information availability across the two phases. This temporal notion of causality should not be conflated with covariate shift between expert and policy-induced state distributions in imitation learning. It does not mitigate the covariate shift arising from policy execution; rather, it constrains when observations enter the historical memory.

Second, multi-step action prediction should provide lookahead for the current decision without sacrificing closed-loop execution. Replanning after each executed action allows the agent to exploit dependencies among predicted future actions without committing to an open-loop action chunk, thereby retaining responsiveness to newly acquired observations. Third, termination carries markedly different consequences from ordinary motion actions. Motion errors can often be corrected through subsequent actions, whereas a premature stop irreversibly ends the navigation episode. Treating termination and motion actions under the same prediction and calibration mechanism can therefore obscure the asymmetric risks of premature and delayed stopping. Together, these considerations motivate the explicit coordination of historical information, future-action planning, and termination within a temporally consistent closed-loop decision process.

To address these issues, we propose DreamFly, a temporally consistent, closed-loop aerial VLN framework built on discrete diffusion-based action generation. First, DreamFly introduces a causally aligned historical memory constructed exclusively from visual observations acquired before the current decision step. Under a read-before-write protocol, M<t is held fixed during decision step t and updated with the current observation only after the current decision, for use at subsequent steps. The resulting historical representation is fused with the current visual features through a lightweight memory-conditioning module, allowing the policy to reason jointly over the language instruction, historical visual evidence, and current observation.

Building on the memory-conditioned representation, DreamFly performs receding-horizon diffusion planning by jointly predicting a chunk of K discrete actions that comprises the current action and captures dependencies among actions over a short future horizon. During online inference, DreamFly follows a plan-K, execute-one strategy: it executes only the first action and replans upon receiving the next observation. In this way, the future positions in the action chunk serve as auxiliary prediction targets for the current decision, while execution remains closed loop.

Finally, we introduce LiteStop, a lightweight module that explicitly predicts termination. Rather than treating termination as an ordinary motion action, LiteStop estimates the probability of stopping from the action-token logit grid Ht extracted from the diffusion policy at its initial all-mask state. LiteStop is trained separately while the navigation policy remains frozen, providing a dedicated termination objective without modifying the learned motion policy. Together, causally aligned historical memory, receding-horizon diffusion planning, and explicit termination form a temporally consistent, closed-loop navigation process.

We make the following three main contributions:

  1. We introduce a causally aligned historical memory for aerial VLN with an explicit temporal information boundary. DreamFly constructs its historical memory exclusively from observations acquired before the current decision step and adopts a read-before-write protocol that prevents current-step information from entering the historical branch. A lightweight gated cross-attention module selectively integrates task-relevant historical evidence into the current visual representation, thereby supporting temporally consistent reasoning under partial observability.

  2. We propose receding-horizon diffusion planning that exploits dependencies among predicted future actions while preserving closed-loop control. DreamFly jointly predicts the current action and a short sequence of future actions as a chunk of K discrete actions, using valid-prefix supervision and horizon-aware loss weighting to emphasize the immediately executable action. During inference, DreamFly executes only the first action and replans upon receiving the next observation, allowing the predicted future actions to serve as auxiliary planning targets while preserving closed-loop execution.

  3. We develop LiteStop, an explicit and decoupled termination module for aerial VLN. Rather than coupling termination with motion-action generation, LiteStop estimates the probability of stopping from the action logits produced by the diffusion policy at its initial all-mask state and is trained separately while the navigation policy remains frozen. This decoupling provides a dedicated termination objective without modifying the learned motion policy.

Method

Problem Formulation

Given a natural-language instruction I = (w1, w2,..., wNI), an aerial agent starts from an initial pose s0 = (p0, ψ0) in a 3D environment, where p0 = (x0, y0, z0) denotes its position and ψ0 denotes its heading. At each decision step t, the agent receives an egocentric RGB observation Ot and selects an action at from a discrete action space A. Executing at moves the agent to a new pose st+1 and produces the next observation Ot+1. The action space consists of Stop, forward motion with different step sizes, left and right turns, vertical movement, and lateral movement. The navigation episode terminates when the agent selects Stop or reaches the maximum number of decision steps. Given the ground-truth destination p*, an episode is considered successful if the Euclidean distance between the agent's final position pT and the destination is less than 20 m: ∥pT − p*∥2 < 20 m.

Framework Overview

DreamFly is an aerial vision-language navigation framework built upon Dream-VLA. We adopt Dream-VLA as the policy backbone because its perception-to-action formulation aligns naturally with aerial VLN, directly mapping visual observations and language instructions to executable actions. Compared with autoregressive VLA backbones, Dream-VLA employs bidirectional diffusion modeling to support action chunking and parallel action prediction, providing a suitable foundation for receding-horizon planning.

Unlike the original VLA setting, aerial navigation requires retaining historical visual evidence and continually replanning as new observations arrive. DreamFly therefore introduces three key designs. First, a causally aligned historical memory M<t is constructed only from observations acquired before the current decision step and conditions the current visual representation, preserving a well-defined temporal information boundary while retaining previously observed landmarks. Second, DreamFly performs receding-horizon diffusion planning by jointly predicting a K-step action chunk. The predicted future actions provide short-horizon planning structure without being committed to consecutive execution. If termination is not triggered, only the first action is executed before the policy replans from the next observation. Third, LiteStop estimates termination from the initial all-mask action logits and provides a separate pre-action termination decision. Together, these components establish a temporally consistent closed-loop process that integrates historical reasoning, receding-horizon future-action planning, and explicit termination control.

Causally Aligned Historical Memory

During continuous aerial navigation, previously observed landmarks and scene cues may leave the current field of view while remaining relevant to subsequent decisions. Retaining all past observations would preserve such information, but would also introduce substantial redundancy and cause the visual input to grow with trajectory length. To maintain useful historical evidence under a bounded input budget, DreamFly compresses instruction-relevant information from past observations into a fixed-capacity historical memory. A short candidate FIFO and active tracks are maintained internally to accumulate evidence across recent observations, while only the long-term memory slots are exposed to the navigation policy.

At decision step t, let Oτ denote the observation acquired at step τ. The historical memory available to the policy is defined as M<t = Fmem(I, (Oτ)τ<t), where Fmem denotes the fixed memory-construction operator. This definition constrains only the historical branch: the current observation Ot remains directly available for action prediction, while information derived from Ot can enter the historical memory only from subsequent decision steps.

Instruction-conditioned candidate construction: For each observation, a frozen CLIPSeg dense router and a frozen OWLv2 region router extract complementary visual candidates conditioned on the complete navigation instruction I. To accommodate the finite text context of the routers without truncating the instruction, the full instruction is covered by fixed overlapping token windows. CLIPSeg provides dense instruction-relevant responses together with a visual feature grid, whereas OWLv2 proposes spatially localized regions. Rather than combining features from the two models, OWLv2 regions are represented directly in the frozen CLIPSeg visual feature space. For an OWLv2 region b, its visual representation is obtained as f(b) = Norm(Σg area(Gg ∩ b) vg), where g indexes the CLIPSeg grid cells, Gg denotes the spatial region of the g-th cell, vg is its normalized visual feature, and Norm(·) denotes L2 normalization. CLIPSeg candidates directly use their corresponding grid features. Candidates are subsequently consolidated only when they exhibit sufficient spatial overlap and visual similarity.

Evidence-driven long-term memory: The resulting candidates are associated with active tracks over recent observations according to visual and spatial consistency, allowing evidence for the same visual content to accumulate across observations. Candidates receiving repeated and stable support become eligible for persistent promotion, whereas candidates without sufficient repeated support may enter memory through single-observation promotion when they satisfy the confidence, region-validity, score-separation, and novelty criteria. These complementary mechanisms preserve both stable cross-observation evidence and informative content that may be visible only briefly.

Eligible candidates are ranked by their write utility, and at most two long-term write candidates are selected at each decision step. For each selected candidate, if a compatible slot is found, that slot is updated with the candidate; otherwise, the candidate is assigned to an empty slot or, when the memory is full, replaces an existing slot according to the retention criterion. Each valid slot maintains an anchor and an optional prototype. The anchor stores the visual feature of the concrete instance with the highest historical write utility, whereas the prototype represents accumulated compatible evidence once stable repeated support has been established. A candidate admitted through single-observation promotion may therefore form a valid anchor-only slot; such a slot is promoted to include a prototype when it is subsequently matched by a persistent write supported by stable cross-observation evidence.

At decision step t, the memory adapter represents the policy-facing memory as M<t = rj<t, µj<t 16j=1, where j indexes the memory slots, µj<t ∈ 0, 1 indicates slot validity, rj<t = [ej,anc<t; ej,pro<t; ρj<t; log(1 + δj<t)], and ej,anc<t and ej,pro<t denote the anchor and prototype features, respectively, ρj<t ∈ 0, 1 indicates prototype presence, and δj<t is the number of decision steps since the slot was last updated. An anchor-only slot remains valid; when no prototype is available, its prototype component is zero-filled and ρj<t = 0. For an invalid slot with µj<t = 0, the complete input representation rj<t is zero-filled and is excluded from the key/value dimension of cross-attention by the slot-validity mask.

Memory-conditioned visual representation: DreamFly conditions the current visual representation on historical memory rather than appending memory slots as additional multimodal tokens. Let Zt = Projector(VisionEncoder(Ot)) denote the current image tokens. Each slot representation rj<t is mapped by a memory slot encoder φ(·), and the resulting slot embeddings are stacked in slot order: E<t = [φ(r1<t);...; φ(r16<t)]. Using the current image tokens as queries and the historical slot embeddings as keys and values, DreamFly retrieves historical context through Ct = MHA(Zt WQ, E<t WK, E<t WV; µ<t), where WQ, WK, and WV are learned projection matrices, µ<t = (µ1<t,..., µ16<t) is the slot-validity mask applied along the key/value slot dimension, and Ct denotes the retrieved historical context. All valid slots participate in the masked cross-attention, allowing the current visual tokens to adaptively retrieve relevant historical evidence without hand-crafted top-k selection.

The retrieved context is incorporated into the current visual tokens through a gated residual connection: Z̃t = Zt + Mimg ⊙ Gt ⊙ (Ct WO), Gt = 1 + tanh(Zt WG + bG), where Gt is the learned gate, WG and bG are its parameters, WO is the residual output projection, Mimg is the valid-image-token mask, and ⊙ denotes element-wise multiplication. The residual output projection WO and the gating parameters WG and bG are zero-initialized. Since WO = 0 at initialization, the adapter initially implements the identity mapping Z̃t = Zt, while the initial gate value is Gt = 1. The resulting memory-conditioned tokens Z̃t are then processed by the Dream-VLA backbone together with the navigation instruction for action prediction.

Causal alignment across training and deployment: Training and closed-loop deployment share the same frozen memory-construction mechanism and temporal visibility boundary, but operate on different observation histories. Training uses prefix memories constructed along expert trajectories, whereas deployment initializes a fresh online memory for each rollout and updates it from observations actually acquired by the agent. Consequently, the resulting memory states may differ, while the causal constraint remains unchanged: the memory M<t used at step t contains only information acquired before that decision step.

Receding-Horizon Diffusion Planning

Predicting only the next action provides an immediate control decision but does not explicitly represent the short-horizon action structure surrounding that decision. DreamFly therefore extends single-step prediction to a fixed-horizon discrete action chunk and exploits the bidirectional Dream-VLA backbone to jointly model the current action together with K − 1 near-future actions. The predicted chunk serves as a structured short-horizon planning representation rather than an open-loop command sequence. At each decision step, DreamFly generates the complete chunk. LiteStop may terminate the episode before action execution; otherwise, at most the leading action is handled, and the policy replans after a non-terminal motion produces new visual feedback.

Discrete action-chunk formulation: Given the memory-conditioned visual representation Z̃t and the navigation instruction I, DreamFly predicts a length-K discrete action chunk ât = [â0t, â1t,..., âK−1t], âht ∈ A, where A is the discrete aerial action space and K is the planning horizon. The first slot â0t corresponds to the current control decision, whereas the remaining slots â1:K−1t represent short-horizon future actions predicted under the same current context. These future predictions form part of the structured short-horizon plan but are not committed commands to be executed consecutively. Let V denote the full Dream-VLA vocabulary, and let χ: A ↪ V denote the injective mapping from each discrete action to its dedicated action token. Its image χ(A) ⊂ V forms the dedicated action-token set. This distinction allows us to use environment-level actions in the planning formulation while explicitly identifying their corresponding vocabulary targets during model training and diffusion generation.

Joint masked action learning: During training, the action chunk is supervised by the suffix beginning at decision step t of the stored action sequence. Let (a⋆0,..., a⋆T−1) denote the length-T action sequence stored for a training trajectory. At a decision step t ∈ 0,..., T − 1, the number of available actions within the prediction horizon is Lt = min(K, T − t), vt,h = I[h < Lt], where vt,h ∈ 0, 1 indicates whether the h-th action slot belongs to the available supervision prefix. To preserve a fixed-length template near the end of a trajectory, positions beyond this prefix are padded with the Stop action. We denote the padded target by ā⋆t,h = a⋆t+h if h < Lt, and ā⋆t,h = astop if h ≥ Lt, where astop ∈ A denotes the Stop action. The validity mask ensures that Stop tokens appended beyond the length-T supervision sequence solely to complete the fixed K-slot template contribute no training loss. Any Stop token already contained in a⋆ = (a⋆0,..., a⋆T−1) remains a valid supervised target.

Importantly, extending the prediction horizon does not expose future visual observations to the policy. At decision step t, the model is conditioned on the current visual representation Z̃t and instruction I; future observations Ot+1,..., Ot+K−1 are not provided. Supervision for future action slots is provided only by the corresponding action targets.

During training, all K action positions supplied to the bidirectional backbone are replaced with [MASK]. Only the valid action prefix is used for supervision, while synthetic tail padding is excluded by the valid-prefix mask. DreamFly then jointly predicts all action slots in a single bidirectional forward pass. Thus, action-policy training does not unroll iterative denoising; during inference, iterative discrete diffusion progressively resolves the masked action slots.

At decision step t, let zt,h ∈ RV, h = 0,..., K − 1, denote the shifted full-vocabulary logit vector at the sequence position that predicts the target token χ(ā⋆t,h) for the h-th action slot. Let Ut denote the valid non-action and non-padding positions in the shifted sequence, and let qt,h denote the shifted sequence position whose logits form zt,h. We define the deterministic geometric kernel κij = 0 if i = j, and κij = (βcar/2)(1 − βcar)i−j−1 if i ≠ j, and compute the context coefficient as ct,h = max(10−6, Σi∈Ut κi,qt,h), with βcar = 0.1. All action slots, including unsupervised tail slots, are excluded from Ut. Let B denote a minibatch of decision-step samples, with trajectory identities omitted for clarity. The action objective is Lact = [Σt∈B Σh=0K−1 vt,h ct,h γh CEV(zt,h, χ(ā⋆t,h))] / [Σt∈B Σh=0K−1 vt,h ct,h γh], where 0 < γ ≤ 1 controls horizon-dependent weighting; values below one progressively emphasize near-term action targets. The validity mask vt,h restricts the objective to positions contained in the length-T supervision sequence. Following Dream-VLA, the cross-entropy is evaluated over the complete vocabulary V. The restriction to the dedicated action-token set χ(A) is applied only during diffusion generation to guarantee executable action outputs.

Discrete diffusion generation: At inference time, DreamFly generates the action chunk through iterative discrete diffusion. Let m(s)t denote the abstract state of the K action slots at denoising step s; fixed formatting tokens in the assistant scaffold are omitted from this notation. Generation begins from m(0)t = [[MASK],..., [MASK]], and proceeds through m(0)t → m(1)t →... → m(S)t, with ΠA(m(S)t) = ât, where S denotes the number of diffusion steps and ΠA(·) decodes the K action slots through χ−1. This notation distinguishes the predicted action chunk from the complete assistant scaffold, which also contains fixed formatting tokens.

We use the monotonic origin sampler inherited from Dream-VLA. Let ξs Ss=0 be linearly spaced from 1 to εdiff = 10−3. At denoising step s ∈ 0,..., S − 1, each unresolved action slot is independently transferred from [MASK] to a sampled action token with probability ωs = 1 − ξs+1/ξs if s < S − 1, and ωs = 1 if s = S − 1. Conditional on transfer, the replacement token is sampled from the model distribution after suppressing all logits outside χ(A). Resolved action slots are not remasked, and the final denoising step resolves all remaining masks.

Unlike fixed left-to-right autoregressive decoding, the bidirectional backbone allows unresolved action slots to share the available context, and multiple slots may be resolved within the same denoising step. Resolved action tokens can subsequently provide context for slots that remain unresolved. Before sampling at each unresolved position, logits outside χ(A) are suppressed; consequently, every transferred token belongs to the dedicated action-token set.

The initial all-mask forward also provides the planning representation consumed by LiteStop. Let z(0)t,h ∈ RV denote the shifted full-vocabulary logit vector aligned with the h-th action position at the initial denoising step. We retain the corresponding action-token slices as H(0)t = [Sliceχ(A)(z(0)t,0);...; Sliceχ(A)(z(0)t,K−1)] ∈ RK×A. Here, Sliceχ(A) extracts the action-token coordinates in the fixed action-ID order. The representation is extracted during the initial denoising forward, while all K action positions remain masked and before any action token is transferred. It therefore records the policy responses at all K action positions under the same all-mask planning state. Importantly, retaining H(0)t does not terminate or shorten diffusion generation; the complete action chunk is generated before LiteStop is evaluated.

Receding-horizon execution: After the complete action chunk has been generated, LiteStop determines whether navigation should terminate before any chunk action is executed. If LiteStop is triggered, the episode terminates without executing an action from ât. Otherwise, the leading action â0t is handled according to the original action-level semantics. If â0t = astop, the episode terminates through the action-level Stop condition; otherwise, the selected motion is executed. The remaining predictions â1t,..., âK−1t are discarded rather than cached for subsequent execution.

After a non-terminal motion is executed, the agent receives the next observation Ot+1 and replans a complete action chunk. The historical memory available at the next decision step follows the same causal definition introduced in Sec. 3.3: M<t+1 = Fmem(I, Oτ τ≤t). Thus, information derived from Ot may influence the historical context at step t+1 only if it is retained by the memory-construction mechanism, whereas the newly acquired observation Ot+1 remains a separate current visual input and is not simultaneously exposed through the historical-memory branch. DreamFly therefore follows a receding-horizon strategy that plans K actions at each decision step and executes at most the leading action before replanning from newly observed visual evidence.

LiteStop: Decoupled Termination Control

Termination has a qualitatively different consequence from ordinary motion decisions: an erroneous motion may be corrected through subsequent replanning, whereas premature termination immediately ends the episode. We therefore introduce LiteStop, a lightweight auxiliary head that derives an additional decision-level termination signal from the frozen policy's initial all-mask planning response. LiteStop is optimized separately from the action policy: it consumes the frozen policy's planning logits, but does not update policy parameters or modify the original action-level Stop semantics.

All-mask termination representation: We reuse the initial planning representation H(0)t ∈ RK×A defined in Sec. 3.4. It contains the shift-aligned raw logits over the dedicated action-token set χ(A) for all K planning positions. LiteStop maps this complete response grid to a scalar Stop logit: stopt = gstop(H(0)t) = W2 SiLU(W1 LN(vec(H(0)t)) + b1) + b2, pstopt = σ(stopt). Here, vec(·) vectorizes the logit grid, LN(·) denotes layer normalization, and gstop contains the learnable LayerNorm and MLP parameters. Using the complete K × A grid allows LiteStop to exploit the policy's short-horizon planning state rather than relying only on the first-position Stop logit.

Frozen-policy supervision: Let a⋆t denote the expert action at decision step t. The binary Stop target is ystopt = I[a⋆t = astop]. A Stop appearing only at a future chunk position or in synthetic tail padding therefore does not affect the current label. The supervision uses neither geometric success nor terminal metadata, and thus calibrates the frozen policy's action-level Stop tendency rather than learning an independent goal-reached classifier.

For a training batch B, we minimize Lstop = −(1/B) Σt∈B [λ+ ystopt log pstopt + (1 − ystopt) log(1 − pstopt)], where λ+ = 4.0 is the fixed positive-class weight. The complete navigation policy, including its visual, memory, and action-planning components, remains frozen, and only LiteStop is optimized. During training, H(0)t is extracted by a single all-mask bidirectional forward pass without unfolding iterative diffusion, matching the all-mask representation semantics used at deployment.

Pre-action termination: During inference, H(0)t is cached from the initial all-mask denoising forward. The frozen policy nevertheless completes the full diffusion process and obtains the complete decoded action chunk before LiteStop is evaluated. LiteStop therefore requires no additional backbone forward pass, but it is not a diffusion early-exit mechanism. This ordering preserves the frozen policy's complete generation path and confines LiteStop to pre-action termination control; computational early exit is outside the scope of the present design.

Given a fixed threshold ηstop, the LiteStop decision is dstopt = I[pstopt ≥ ηstop]. The final termination decision combines LiteStop with the frozen policy's action-level Stop condition: dtermt = dstopt ∨ I[â0t = astop]. If dstopt = 1, the episode terminates before any action from the current chunk is executed. Otherwise, the first action â0t is handled under the receding-horizon protocol from Sec. 3.4. If â0t = astop, the episode terminates through the retained action-level Stop condition; if neither termination condition is satisfied, the selected motion is executed and the agent acquires a new observation before replanning the complete action chunk. LiteStop therefore provides an additional pre-action termination pathway rather than vetoing the frozen policy's action-level Stop condition.

Overall closed-loop operation: Taken together, the three components form a causal receding-horizon control loop. At decision step t, DreamFly combines the current observation Ot with historical memory M<t to obtain the memory-enhanced visual representation Z̃t. Together with the navigation instruction I, this representation conditions the action policy, which generates a complete K-step action chunk ât while retaining the initial all-mask planning representation H(0)t for LiteStop. After the complete action chunk has been generated and before any action from the chunk is executed, LiteStop determines whether the episode should terminate. If termination is not triggered, the leading action â0t is handled according to the original action-level semantics.

The historical memory available to the current decision contains only information derived from observations preceding Ot. Information extracted from Ot may affect the historical memory only from subsequent decisions onward and can therefore first become accessible through M<t+1 if it is retained by the memory-construction mechanism. After a non-terminal motion produces a new observation Ot+1, DreamFly repeats the same process using Ot+1 as the current visual input together with the corresponding causal memory prefix.

The components are optimized in stages rather than through a single joint objective. Historical memory construction remains fixed while the memory-conditioned action policy is trained with the horizon-aware action objective Lact. After the navigation policy has been trained, the complete policy is frozen and only LiteStop is optimized with Lstop. Consequently, LiteStop training does not back-propagate into the visual, memory-adaptation, or action-planning components. This staged design integrates causal historical context, short-horizon diffusion planning, and decoupled termination calibration within a unified closed-loop navigation policy.

Experiments

Experimental Setup

Datasets: We conduct our experiments based on the OpenFly dataset. Before training, we apply four standardization and augmentation steps to the released data. First, we correct the 8-D action vector associated with Forward 6m to its canonical encoding. Second, approximately 190,000 decision steps with non-standard action labels are remapped to the corresponding canonical actions, with −1 mapped to Go Up and −2 to Go Down. Third, we remove the pre-packaged historical keyframes and retain only the current RGB observation at each decision step. Fourth, We additionally store [∆x, ∆y, ∆z, yaw] as provenance metadata. This metadata is not used as an input to either historical-memory construction or policy inference. These relative poses are not used as additional policy inputs during closed-loop inference. The resulting training set contains 20 subsets with 85,785 trajectories and 1,356,622 decision steps.

Our evaluation covers eight AirSim/UE environments with 1,796 trajectories. The test-seen split contains UE BigCity and six AirSim urban environments, totaling 1,392 trajectories, whereas the test-unseen split contains 404 trajectories from UE SmallCity. This evaluation therefore covers both Unreal Engine and AirSim environments and assesses navigation performance in seen environments as well as generalization to an unseen urban scene. As shown in Fig. 3, the training data exhibit a pronounced imbalance toward forward actions, while the test-seen and test-unseen splits show different distributions of initial goal distances.

Evaluation Metrics: Following the evaluation protocol of OpenFly and prior aerial VLN studies, we adopt four standard metrics: navigation error (NE), success rate (SR), oracle success rate (OSR), and success weighted by path length (SPL). NE measures the average Euclidean distance between the UAV's final stopping position and the ground-truth target position, with a lower value indicating more accurate goal localization. SR measures the proportion of successful trajectories, where a trajectory is considered successful if the UAV stops within 20 m of the target. OSR measures the proportion of trajectories that come within 20 m of the target at any point during navigation, regardless of the final stopping position, and therefore reflects whether the agent successfully reaches the target vicinity. SPL jointly evaluates navigation success and path efficiency by weighting each successful trajectory according to the ratio between the shortest-path distance and the actual traveled path length. Higher values of SR, OSR, and SPL indicate better performance.

Implementation Details: The Dream-VLA backbone is fine-tuned with all-linear LoRA (r = 32, α = 16), while the memory-fusion adapter is trained jointly and the base projector remains frozen. We use an action-chunk length of K = 4, horizon decay γ = 0.7, and CAR reweighting probability p = 0.1. Training uses AdamW with a learning rate of 1 × 10−4 and batch size 8 for up to 10,000 optimization steps. The navigation checkpoint at step 5,000 is used to train LiteStop and for closed-loop evaluation. Inference uses 12 discrete diffusion steps. The historical memory contains 16 long-term slots with 512-dimensional features. Memory is fused with the current visual tokens using a gated cross-attention module with dimension 512 and eight attention heads.

LiteStop operating-point selection: We evaluate ηstop ∈ 0.50, 0.65, 0.80 using the step-500 LiteStop checkpoint on a balanced 64-trajectory calibration set (8 per environment), disjoint from the final evaluation split. Based on the overall trade-off in Table 1, We select ηstop = 0.50 because it yields the smallest OSR–SR gap and a slightly lower NE than ηstop = 0.80 while preserving the same SR. Unlike ηstop = 0.80, it also produces nontrivial LiteStop interventions.

Experiment Result

As shown in Table 2, we compare DreamFly with six baselines: Random, Action Sampling, Seq2Seq, CMA, AerialVLN, and OpenFly-Agent. Random uniformly samples an action from the discrete action space at each step and executes it until either Stoptok is sampled or the maximum number of navigation steps is reached. Action Sampling follows the same procedure but samples actions according to their empirical distribution in the training set, providing a stronger stochastic baseline that reflects the action prior of the dataset. For OpenFly-Agent, we directly evaluate the officially released checkpoint in the same test environment to avoid introducing discrepancies by reproducing its original training and historical-observation pipeline under our standardized data protocol. All other learning-based baselines are trained on our processed training data for the same number of optimization steps.

As shown in Figure 4, the non-zero success rates of the stochastic baselines should not be interpreted as evidence of effective goal-directed navigation. To examine the influence of the initial state, we further partition the test trajectories according to whether their initial goal distance is within the 20 m success radius and report the conditional success rates for the two groups. Both Random and Action Sampling exhibit substantially higher success rates when the agent is initialized within the success radius than when it starts outside this region. This pronounced gap indicates that their non-zero overall success rates are strongly influenced by favorable initial configurations, rather than reflecting navigation policies that consistently drive the agent toward the target.

Ablation Studies

To evaluate the contribution of each component, we conduct both progressive and leave-one-out ablation studies, as reported in Table 3. Starting from the Dream-VLA baseline, introducing the causally aligned historical memory consistently improves navigation performance, while receding-horizon diffusion planning further improves the results by jointly modeling the current action and short-horizon future actions. Incorporating LiteStop yields the strongest overall performance. Conversely, removing Memory, Action Chunk, or LiteStop from the complete framework results in performance degradation. These consistent trends indicate that historical context modeling, future-action planning, and explicit termination provide complementary benefits to DreamFly.

To further examine how these components contribute under different navigation distances, we partition the test trajectories into three groups according to their initial shortest-path distance. As shown in Fig. 5, LiteStop provides its largest gain in the shortest-distance group, where successful navigation depends more directly on recognizing the appropriate termination point. Historical memory contributes more prominently at intermediate distances, while for larger initial distances, memory and action-chunk planning provide complementary benefits by maintaining cross-step visual context and introducing short-horizon future-action structure into each replanning step.

The contribution of LiteStop decreases as the initial distance increases, which is consistent with termination becoming relevant only after the agent reaches the vicinity of the goal. In this regard, the gap between OSR and SR provides a complementary view of termination behavior, since OSR reflects whether the trajectory reaches the success region whereas SR additionally requires successful termination within it. Overall, the distance-wise results further reveal the distinct roles of historical memory, receding-horizon diffusion planning, and LiteStop in the DreamFly framework.

Qualitative Analysis

Fig. 6 qualitatively illustrates how the three proposed components affect closed-loop navigation. In the first example, DreamFly steadily approaches the target and reduces the navigation error from 78.9 m to 10.9 m, whereas removing historical memory results in much slower progress and a final error of 43.5 m. This comparison shows that retaining previously observed visual evidence helps maintain a consistent reference to the target under changing viewpoints. In the second example, DreamFly follows a coherent sequence of actions and reaches the target with an error of 2.2 m, while removing action-chunk prediction leads to an early deviation and eventually fails with an error of 58.1 m, highlighting the benefit of incorporating short-horizon future-action structure into the current decision. In the third example, both variants approach the target region, but the model without LiteStop continues moving after entering the success region and drifts away, ending at 23.0 m. DreamFly instead terminates successfully at 12.5 m, demonstrating the importance of explicitly modeling task completion.

Overall, the examples reveal complementary failure modes addressed by the proposed components: historical memory preserves cross-step visual context, receding-horizon planning improves action consistency, and LiteStop prevents unnecessary motion after reaching the goal region.

Conclusion

In this work, we present DreamFly, a diffusion-based framework for aerial vision-language navigation. Building upon the Dream-VLA backbone, we revisit aerial navigation from the perspective of temporal decision making and identify three essential capabilities: retaining informative historical observations, anticipating future actions while preserving closed-loop feedback, and reliably determining when navigation should terminate. To this end, DreamFly integrates causally aligned historical memory, receding-horizon diffusion planning, and LiteStop for explicit termination into a unified closed-loop framework. The historical memory provides the policy with accumulated visual evidence under a strict causal prefix constraint, while diffusion-based action-chunk prediction enables joint modeling of the current action and short-horizon future actions. By following a plan-K, execute-one strategy, DreamFly uses future-action predictions as planning variables while replanning after every executed action, thereby preserving closed-loop feedback. LiteStop further decouples termination from action generation by estimating the stop probability from the initial all-mask action logits.

Extensive experiments on OpenFly validate the effectiveness of the proposed framework. DreamFly achieves the best NE, SR, and SPL on both the test-seen and test-unseen splits, reaching 32.04%/29.46% SR and 28.22%/23.54% SPL, respectively. The consistent performance improvements across both splits further demonstrate that DreamFly remains effective when navigating previously unseen environments.

Despite the consistent improvements observed in simulation, the current evaluation of DreamFly is still limited to simulated environments. Future work will focus on deploying the framework on physical UAV platforms to assess its real-world navigation capability and robustness under sensing noise, environmental disturbances, and sim-to-real domain shifts.

Improvements for AI systems

Based on this paper, here are the specific improvements you can make to AI systems and what the improved systems can do:

1. Implement causally aligned historical memory with read-before-write protocol

  • Construct memory exclusively from observations preceding the current decision step

  • Use a gated cross-attention module to selectively retrieve task-relevant historical evidence

  • Maintain fixed-capacity memory slots with anchor/prototype representations and validity masks

2. Adopt receding-horizon diffusion planning with plan-K, execute-one strategy

  • Jointly predict a K-step action chunk using bidirectional diffusion, but execute only the first action

  • Use valid-prefix supervision and horizon-aware loss weighting (γ=0.7) to emphasize the immediately executable action

  • Apply deterministic geometric kernel for context-aware reweighting during training

3. Integrate decoupled termination control (LiteStop)

  • Extract stop probability from initial all-mask action logits before diffusion generation completes

  • Train termination head separately while freezing the main policy

  • Use positive-class weighting (λ+=4.0) to handle class imbalance

  • Combine LiteStop decision with action-level Stop condition via logical OR

4. Apply staged optimization rather than joint training

  • Train memory-conditioned action policy first with horizon-aware objective

  • Freeze the complete policy, then train only the termination module

  • This prevents interference between objectives and enables modular improvements

Navigation and Control:

  • Maintain temporal consistency across long trajectories without losing track of previously observed landmarks

  • Make decisions with short-horizon lookahead while remaining responsive to new observations

  • Terminate reliably at goal locations, reducing both premature stops and overshooting

  • Handle partial observability by retrieving relevant historical context adaptively

Robustness and Generalization:

  • Operate effectively in unseen environments (29.46% SR on test-unseen vs. 32.04% on test-seen)

  • Maintain performance across varying initial distances from goals

  • Correct navigation errors through replanning rather than committing to open-loop action sequences

Efficiency:

  • Reduce computational overhead by using a lightweight termination module that reuses existing representations

  • Avoid redundant processing by discarding unpredicted future actions and replanning only when needed

  • Achieve lower navigation error (NE) compared to baselines while maintaining path efficiency (SPL)

Decision Quality:

  • Distinguish between motion errors (correctable) and termination errors (irreversible)

  • Exploit dependencies among future actions without sacrificing closed-loop feedback

  • Adapt memory retention based on instruction relevance rather than storing all past observations

Sources

Related papers