Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models

arXiv:2608.10525 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

Yuhang Song, Bor-Jiun Lin, Jiaxu Liu, Te-Chuan Chiu, Anh Nguyen, Chun-Yi Lee

University of Liverpool · National Tsinghua University · Imperial College London · National Taiwan University

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

Terminology

Summary

Affiliation: University of Liverpool, National Tsinghua University, Imperial College London, National Taiwan University

arXiv: 2608.10525v1 [cs.CV] 11 Aug 2026


The paper addresses the fundamental challenge of integrating historical context into Vision-Language Models (VLMs) for sequential decision-making tasks. As stated in the abstract: "Historical context integration presents a fundamental challenge for Vision-Language Models (VLMs) in sequential decision-making tasks. Current VLMs process visual inputs independently, which creates critical limitations for downstream applications that require temporal understanding."

The authors note that in Vision-and-Language Navigation (VLN), agents must synthesize information from multiple past observations to navigate complex multi-room environments where current visual input alone provides insufficient context for decision-making. They specifically highlight that "under partial observability, an agent's onboard camera captures only a limited field of view at each timestep, making historical context essential for inferring occluded landmarks, retracing steps, and maintaining spatial awareness."

The paper identifies three main categories of existing strategies for integrating historical context, each with significant drawbacks:

  1. Token concatenation approaches: Direct incorporation of historical frames into Transformer inputs produces quadratic attention complexity and excessive memory consumption. These methods disrupt the downstream token order and flood the model with redundant information.

  2. Recurrent compression methods: These employ RNNs or LSTMs to compress the entire frame history into a single state vector but lack the capacity to represent fine temporal structure, leading to information loss over extended sequences.

  3. External memory methods: These depend on manually constructed maps that may not generalize across different environments.

The paper introduces DCA, described as a novel context injection approach for pretrained VLMs. The method employs fixed-size, dynamically compressed memory to preserve historical semantics without frame concatenation and bridges static VLMs and recurrent policies and enables memory capabilities in pretrained models while maintaining computational efficiency.

The authors employ a compact pretrained VLM backbone using Prismatic-VLM with the phi-2+3b variant with only 3B parameters, which incorporates a ViT-based CLIP visual encoder, a lightweight Phi-2 language model, and multi-layer cross-modal projection.

The architecture consists of two parallel pathways:

  1. Standard VLM pathway: Processes current frame tokens and instruction tokens through the pretrained VLM to produce decoder embeddings for action prediction.

  2. Context adaptation pathway: Operates in parallel with two key modules:

  • Memory Compression Module: A fixed-size learnable compression vector Minit queries past embeddings through our Memory Compression Module, producing compressed memory.

  • Memory Integration Module: Attends over compressed memory with current decoder queries to extract context-enhanced outputs that can adapt into LLM layers without inflating input sequences.

The compression process works as follows: We initialize a learnable compression vector Minit = nn.Embedding(C, d).weight ∈ RC×d for each timestep t, where C denotes memory token count and d represents embedding dimension. Historical frames are encoded via the vision encoder and processed through grid pooling operator G: RP×d → Rp×d (with p ≪ P) to reduce spatial redundancy.

The compression computation is: M1:t−1 = Scps VF ∈ RC×d, where Scps = Softmax(QM KFT) ∈ RC×p achieving O(C · p) complexity.

For each Transformer layer k, the integration process is: zkcontext = Sintg VM, where Sintg = Softmax(Qk−1 KMT). The final layer output combines representations: zk+1 ← zk+1 + λzk+1context.

This achieves favorable O(S · C) for efficient context integration compared to the prohibitive O(S · t · p) for the naive concatenation approaches.

The paper claims three key advantages for DCA:

  1. Computational efficiency: DCA ensures computational efficiency by maintaining constant input token length regardless of past frame quantity, achieving linear complexity growth with extended context.

  2. Preservation of fine-grained details: "DCA preserves fine-grained contextual details through dynamic compression of critical information into fixed context vectors. This approach avoids temporal detail loss common in recurrent models while removing redundant features."

  3. Preservation of pretrained knowledge: DCA retains the original input and fully preserves the priors of the pretrained VLM, which enables maintaining its learned knowledge to the greatest extent possible.

From Table 1, DCA demonstrates substantial efficiency gains compared to the No-Adapt baseline:

  • Average inference time decreases from 3.21s to 2.71s per step

  • FLOPs reduce from 4.77T to 4.23T

  • Peak GPU memory usage drops from 37.84 GB to 34.31 GB

The paper reports over 25% reduction in attention FLOPs and 13% memory savings while improving performance on long-horizon tasks. At history length δ = 30, DCA achieves over 25% reduction in additional FLOPs relative to No-Adapt.

For memory efficiency, DCA consistently uses approximately 30% less memory than No-Adapt for δ ≥ 1.

On the VLN-CE R2R Val-Unseen split, DCA achieves:

  • Success Rate (SR): 13.7%

  • Success Rate weighted by Path Length (SPL): 12.9%

  • Oracle Success Rate (OS): 25.3%

  • Navigation Error (NE): 6.77

  • Trajectory Length (TL): 6.73

Key comparisons:

  • Compared to recurrent baselines RGB-Seq2Seq and RGB-CMA, DCA achieves relative success rate improvements of 13.7% and 8.7%, respectively.

  • Against Recurrent-Adapt, which shares our backbone and adaptation framework, DCA delivers 7.11% SR improvement, validating dynamic compression effectiveness over recurrent approaches.

  • "DCA outperforms concatenation-based approaches: it surpasses No-Adapt by 6.47% in SR while matching NaVid-IL performance despite using a smaller backbone (3B vs. 7B parameters) and standard training rather than auxiliary co-training."

The paper analyzes attention patterns within the Memory Compression Module and finds selective focus on semantically relevant observations. Specifically: "Early observations without visible targets (bedroom door) receive negligible attention weights, reflecting limited utility for decision-making at T = 68. Conversely, frames containing critical visual cues exhibit pronounced attention peaks: the target door at t = 53 and t = 61, and bedroom interior at t = 67 show substantially elevated weights."

This confirms "our method's ability to identify and prioritize critical contextual features while efficiently discarding temporally irrelevant information, achieving efficiency through intelligent temporal filtering rather than indiscriminate reduction."

Table 3 presents ablation results:

  1. Feature adaptation: FiLM-based fusion shows substantial SR and SPL drops because additional scaling parameters hinder stable context integration. Varying λ (0.5, 0.8) shows larger values consistently improve success and SPL.

  2. Context compression: Augmenting with instruction attention underperforms direct compression due to data quality issues in R2R where instructions and trajectories are misaligned.

  3. Memory capacity: Varying memory tokens C (24, 48, 64) shows Performance improvements correlate with increased C values, indicating greater capacity captures richer temporal patterns, though excessive increases risk overfitting.

The paper concludes: "We introduced DCA, a lightweight framework that efficiently integrates historical context into pretrained VLMs without inflating input token lengths. Our proposed approach employed a Memory Compression Module to distill past frame embeddings into fixed-size learnable memory vectors and a Memory Integration Module to adapt these compressed representations into each Transformer layer. This design preserved the pretrained VLM architecture while achieving linear scaling with extended context lengths. Our extensive evaluations on downstream VLN tasks demonstrated that DCA can achieve superior efficiency-performance trade-offs compared to existing approaches."

Improvements for AI systems

Improvements to AI Systems:

  1. Temporal Context Integration for VLMs: I can add a fixed-size memory compression module that dynamically distills historical visual frames into a small set of learnable context vectors (e.g., 48 tokens), then inject these vectors into every Transformer layer via cross-attention. This enables the VLM to maintain spatial and temporal awareness over long horizons (e.g., 30+ steps) without quadratic attention growth or token flooding.

  2. Efficient Memory Management: I can implement a grid-pooling operator that reduces spatial redundancy in each frame (e.g., from 576 to 64 patches) before compression, cutting per-step FLOPs by over 25% and peak GPU memory by 13% while preserving critical landmarks and occluded-object cues.

  3. Selective Attention for Relevance Filtering: I can train the compression module to assign near-zero weights to irrelevant past frames (e.g., empty corridors) and high weights to semantically critical ones (e.g., target doors, room interiors), enabling the system to prioritize actionable history rather than indiscriminately compressing all data.

  4. Preservation of Pretrained Knowledge: I can keep the original VLM input (current frame + instruction) completely unchanged and only add parallel context-adaptation pathways, so the system retains all pretrained priors (e.g., object recognition, language grounding) without fine-tuning or catastrophic forgetting.

  5. Linear Scaling with Extended Context: I can replace recurrent or concatenation-based memory with a fixed-size compression vector that queries past embeddings via softmax attention, achieving O(C·p) compression and O(S·C) integration complexity—linear in context length—enabling deployment on resource-constrained robots or edge devices.

  6. Adaptive Context Weighting per Layer: I can add a learnable scalar λ (e.g., 0.8) to blend context-enhanced outputs into each Transformer layer’s residual stream, allowing the model to control how strongly historical information influences current decisions, improving stability and success rate by 7% over recurrent baselines.

  7. Memory Capacity Tuning: I can expose memory token count (C) as a hyperparameter (e.g., 24, 48, 64) to trade off between richer temporal patterns and overfitting risk, enabling task-specific optimization for navigation, video QA, or robotic manipulation.

What the Improved AI System Can Do:

  • Navigate complex multi-room environments (e.g., VLN-CE) with partial observability, inferring occluded landmarks and retracing steps using compressed historical context, achieving 13.7% success rate on unseen splits—outperforming recurrent and concatenation baselines with a 3B-parameter backbone.

  • Operate in real-time on embodied agents: inference time drops from 3.21s to 2.71s per step, and memory usage falls by 30% compared to naive history concatenation, making it feasible for drones or mobile robots with limited GPU memory.

  • Maintain long-horizon coherence (e.g., 30+ timesteps) without forgetting early critical cues, such as a target door seen 15 steps ago, by dynamically re-weighting memory based on current query relevance.

  • Retain full pretrained VLM capabilities (e.g., instruction following, visual grounding) while adding memory, so it can be dropped into existing systems without retraining from scratch or losing prior knowledge.

  • Generalize across environments without manually constructed maps, since memory is learned end-to-end from visual embeddings, not hand-crafted spatial representations.

  • Scale to longer episodes (e.g., 50+ steps) with only linear cost growth, enabling tasks like search-and-rescue, warehouse navigation, or long-term activity recognition where current-frame-only models fail.

Abstract

Historical context integration presents a fundamental challenge for Vision-Language Models (VLMs) in sequential decision-making tasks. Current VLMs process visual inputs independently, which creates critical limitations for downstream applications that require temporal understanding. Direct incorporation of historical frames into Transformer inputs produces quadratic attention complexity and excessive memory consumption. Existing approaches suffer from significant drawbacks: computational inflation or substantial information loss through temporal compression. To address these challenges, we introduce Dynamic Context Adapter (DCA), a novel context injection approach for pretrained VLMs. Our method employs fixed-size, dynamically compressed memory to preserve historical semantics without frame concatenation. DCA bridges static VLMs and recurrent policies and enables memory capabilities in pretrained models while maintaining computational efficiency. DCA achieves over 25% reduction in attention FLOPs and 13% memory savings while improving performance on long-horizon tasks.

Sources

Related papers