StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

arXiv:2608.10949 · cs.CV, cs.CL · Submitted 2026-08-11 · Read on arXiv

Muxin Fu, Yifan Zhang, Wentao Zhang, Fangming Guo, Qian Chen, Guibin Zhang, Shuicheng Yan, Bo An

Tongji University · Nanyang Technological University · University of Michigan · The Hong Kong University of Science and Technology · National University of Singapore

cs.CV, cs.CL

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: StreamFlow introduces an efficient visual memory framework for streaming video understanding that enables dynamic, on-demand access to historical visual information.

Terminology

Summary

StreamFlow introduces an efficient visual memory framework for streaming video understanding that enables dynamic, on-demand access to historical visual information. It combines a lightweight, dynamics-aware mid-term memory that filters temporal redundancy before visual encoding with a latent long-term memory that consolidates historical video content into visual latents accessible to subsequent reasoning. During generation, an attention-guided retrieval mechanism injects relevant visual latents when the model's reliance on visual evidence weakens. StreamFlow achieves state-of-the-art streaming video understanding performance, reaching 67.73% overall accuracy on StreamingBench, while also delivering strong performance on offline long-video benchmarks. Relative to the vanilla setting, it improves the visual attention score (VAS) by 59.1% while reducing end-to-end latency and peak memory by 50.4% and 21.1%, respectively, enabling more visually grounded and efficient reasoning.

The framework comprises three main components: (1) Dynamics-aware mid-term memory identifies temporal changes directly from raw pixel differences and selectively encodes dynamic visual content, reducing redundant encoding while providing the MLLM with direct access to recent spatiotemporal information. (2) Latent long-term memory stores and consolidates earlier visual content as latent representations within a fixed capacity, keeping memory usage bounded as the video stream grows while making long-term visual information available for flexible access during generation. (3) Attention-guided memory injection monitors the MLLM's attention to visual information and injects visual latents retrieved from the long-term memory when needed. These components together form a continuous memory flow in which visual information is selectively encoded, consolidated over time, and dynamically reactivated when grounding weakens during generation.

The mid-term memory partitions the video stream into consecutive groups of pictures (GOPs), identifies temporal changes from raw pixel differences, and selectively encodes dynamic patches. It uses reference-anchored grouping where each GOP comprises a single I-frame and multiple P-frames, patch-level residual scoring to localize changes directly in raw pixel space before visual encoding, and sparse patch selection that allocates asymmetric patch budgets across each GOP to suppress temporally redundant content while preserving informative dynamics. All N I-frame patches are retained as the spatial reference, whereas each P-frame keeps the top ⌈ρN⌉ patches, with ρ ∈ (0, 1].

The latent long-term memory encodes earlier video segments as GOP-level visual latents that remain accessible to subsequent reasoning. When a sparse GOP leaves the mid-term memory, StreamFlow encodes its retained patches at their original spatiotemporal coordinates using the frozen visual encoder of the backbone MLLM. The memory maintains a temporally ordered latent memory with a fixed capacity of C GOPs. Once the memory reaches capacity, StreamFlow selects the most similar adjacent pair by averaging the cosine similarities between spatially corresponding tokens of their I-frames, and merges them by anchoring on the earlier I-frame, retaining the ⌈ρN⌉ least similar tokens from the later I-frame as a sparse P-frame, and uniformly subsampling the combined 2T frames to T frames.

The attention-guided memory injection uses the visual attention score (VAS) as an online grounding signal, defined as the visual attention mass averaged across all layers and heads. A low VAS indicates that linguistic context is dominating visual evidence, prompting StreamFlow to reactivate long-term memory for subsequent generation. Once the model's visual attention score falls below the threshold τ, StreamFlow retrieves the K GOPs most relevant to the current reasoning state using a neural compressor that jointly processes the chronologically concatenated GOPs, the query tokens, and learned queries, retaining only the outputs of the learned queries as a fixed visual bottleneck.

Extensive evaluations across diverse video understanding benchmarks consistently demonstrate the effectiveness of StreamFlow in both streaming and offline settings. On StreamingBench, StreamFlow surpasses the previous best result by 4.63%, achieving an overall score of 67.73%. Its benefits extend to long-video understanding, outperforming leading streaming methods by 2.24% and 8.22% on MLVU and VideoMME, respectively. On MLVU, StreamFlow obtains 75.34%, and on VideoMME, it achieves the best results across the short, medium, and long subsets with scores of 87.78%, 70.67%, and 62.11%, respectively, leading to an overall score of 73.52%. Mechanistic analysis shows that attention-guided injection improves the mean visual attention score (VAS) by 59.1% relative to the vanilla setting, indicating that StreamFlow restores attention to historical visual evidence as generation proceeds. Relative to the vanilla setting, StreamFlow further reduces end-to-end latency and peak memory by 50.4% and 21.1%, respectively.

The paper also includes sensitivity analysis and ablation studies. Increasing the retention ratio from 0.25 to 0.50 improves accuracy from 78.75% to 81.55%, while subsequent increases to 0.75 and 1.00 yield smaller gains of 0.85% and 0.55%. VAS-guided insertion achieves 81.55% accuracy, outperforming budget-matched delimiter and random insertion by 0.23% and 0.49%, respectively. Removing mid-term memory reduces accuracy from 81.55% to 76.86% on RTVU and from 62.11% to 60.33% on VideoMME-Long, while removing long-term memory lowers accuracy to 80.18% on RTVU and 51.67% on VideoMME-Long, establishing the complementarity of the two pathways: mid-term memory supplies detailed recent context, while long-term memory preserves evidence required for long-horizon reasoning.

Improvements for AI systems

Improvements to AI systems:

  1. Add dynamics-aware selective encoding to vision encoders. Instead of encoding every frame or patch uniformly, the system can compute raw pixel-level residual scores between reference and subsequent frames, then allocate asymmetric patch budgets (e.g., retain all I-frame patches, keep top ⌈ρN⌉ patches per P-frame). This reduces redundant computation by up to 50.4% latency and 21.1% peak memory while preserving informative dynamics.

  2. Implement a two-tier hierarchical memory with bounded capacity for streaming inputs. Use a short-term, high-resolution memory (recent GOPs with sparse dynamic patches) and a long-term, fixed-capacity latent memory (consolidated GOP-level visual latents). When capacity is exceeded, merge the most similar adjacent GOPs by anchoring on the earlier I-frame, retaining the least similar tokens from the later I-frame, and uniformly subsampling frames. This keeps memory usage constant regardless of video length.

  3. Add attention-guided retrieval during autoregressive generation. Monitor the mean visual attention score (VAS) across all layers and heads in real time. When VAS falls below a threshold τ, retrieve the K most relevant historical GOPs using a neural compressor that processes concatenated GOPs, query tokens, and learned queries, then inject only the learned-query outputs as a fixed visual bottleneck. This restores visual grounding mid-generation, improving VAS by 59.1%.

  4. Enable reference-anchored grouping for temporal change detection. Partition video into GOPs (one I-frame + multiple P-frames), compute patch-level residual scores in raw pixel space before any visual encoding, and use sparse patch selection to suppress temporally redundant content. This allows the system to skip encoding static regions entirely, reducing computational load while retaining all spatial reference information.

  5. Use a dual-pathway memory architecture for complementary reasoning. Maintain a mid-term memory for detailed recent spatiotemporal context (dynamic patches from recent GOPs) and a long-term memory for consolidated historical evidence (latent representations). This separation improves long-horizon reasoning—removing mid-term memory drops accuracy by 4.69% on RTVU and 1.78% on VideoMME-Long; removing long-term memory drops accuracy by 1.37% on RTVU and 10.44% on VideoMME-Long.

What the improved AI system can do:

  • Stream real-time video understanding with state-of-the-art accuracy (67.73% on StreamingBench, +4.63% over prior best) while using 50.4% less latency and 21.1% less peak memory than vanilla baselines.

  • Handle arbitrarily long videos without unbounded memory growth, by consolidating old content into fixed-capacity latent memories and merging similar adjacent segments.

  • Maintain visual grounding during long generations—when the model starts relying too much on language context, it automatically retrieves and injects relevant historical visual latents, improving visual attention by 59.1%.

  • Achieve strong offline long-video performance (75.34% on MLVU, 73.52% on VideoMME overall, including 87.78% on short, 70.67% on medium, 62.11% on long subsets) without sacrificing streaming capability.

  • Adaptively allocate compute based on temporal dynamics—static scenes require minimal encoding, while dynamic scenes get higher patch budgets, enabling efficient processing of both surveillance footage and action-heavy content.

Sources

Related papers