FreshMem: Brain-Inspired Frequency-Space Hybrid Memory for Streaming Video Understanding
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "FreshMem: Brain-Inspired Frequency-Space Hybrid Memory for Streaming Video Understanding".
Jane: The gist The FreshMem framework proposes a Frequency-Space Hybrid Memory network inspired by brain mechanisms to reconcile short-term fidelity with long-term coherence for streaming video understanding.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So to wrap up on "FreshMem: Brain-Inspired Frequency-Space Hybrid Memory for Streaming Video Understanding," the paper is essentially proposing this dual-pathway storage system inspired by brain mechanisms to manage video context in a way that balances immediate detail with long-term coherence #pg2.
Jane: The authors are Kangcong Li, Peng Ye, Lin Zhang, Chao Wang, and Huafeng Qin #pg1. They show that this hybrid design allows the model to effectively reconcile short-term fidelity with long-term semantic coherence for streaming video understanding #pg3.
Lu: The implication is that we can move toward models that don't just process frames in isolation but maintain a more integrated historical context within continuous video streams #pg2.
Meng: For practical AI engineering, this means we have a training-free plug-and-play module that improves performance without requiring a full fine-tune of the existing MLLMs #pg3.
Tom: It’s about using these frequency and space memories to give the model global historical context, structured episodic summaries, and immediate details all at once #pg10.
Jane: So, FreshMem provides a superior performance-efficiency trade-off that makes it suitable for both research exploration and real-world deployment where resources are constrained but high performance is demanded #pg10.
Conclusion: Tom: So, FreshMem is this new system they put out that tries to get video AI to remember long stretches of footage without losing the important stuff #pg2.
Jane: It's really about building a kind of hybrid memory network inspired by how our brains handle information flow #pg10.
Lu: They are trying to bridge that gap between needing quick, detailed understanding right now and needing a big picture view later on #pg10.
Meng: So instead of just looking at the frame in front of it, this system tries to keep a running log of what happened before #pg5.
Lalam: The core idea is splitting memory into two paths: one for high-frequency details and one for the overall context #pg10.
Tom: And they use these frequency coefficients to project past frames into a global gist, which is pretty neat #pg10.
Jane: It also uses space thumbnails to cluster those episodes together, making the long history manageable in a dense way #pg9.
Meng: The numbers show it boosts performance on different benchmarks, but it’s important they mention the overhead isn't massive for actual deployment #pg6.
Lalam: They even have this part where it keeps emotionally salient moments by looking at how much information those moments carry #pg6.
Tom: So what does this actually mean for us right now in video understanding?
Jane: It means we can build systems that understand the whole story of a long stream, not just the current second #pg10.
Lu: It opens up possibilities for more coherent video models that handle continuous information better than they do now #pg2.
Meng: Practically speaking, if this works as claimed with reasonable overhead, it could make things like long-horizon tasks much more feasible #pg4.
College of Future Information Technology, Fudan University
cs.CV, cs.AI
Submitted: 2026-02-02
Updated: 2026-10-08
Importance score: 70/100
The gist: The gist The FreshMem framework proposes a Frequency-Space Hybrid Memory network inspired by brain mechanisms to reconcile short-term fidelity with long-term coherence for streaming video
Key concepts
- Multi-scale Frequency Memory (MFM)
- This module uses Discrete Fourier Transforms (DFT) on old frames to project them into representative frequency coefficients, capturing the overall gist of past video segments. It updates these coefficients incrementally using a recurrent mechanism, ensuring efficient tracking of historical patterns.
- Space Thumbnail Memory (STM)
- STM organizes the continuous video stream into episodic clusters by using adaptive compression to create high-density space thumbnails. This process groups related frames together, allowing the system to distill large sequences into manageable summaries for better long-term recall.
- Amygdala-Inspired Residuals
- This mechanism identifies emotionally important events by calculating the L2-norm of input features, which acts as a measure of information entropy. It keeps the most significant events as residual tokens, ensuring that salient details are retained alongside general context.
Terminology
Summary
The gist The FreshMem framework proposes a Frequency-Space Hybrid Memory network inspired by brain mechanisms to reconcile short-term fidelity with long-term coherence for streaming video understanding.
How it works
FreshMem reconciles short-term fidelity with long-term coherence through two synergistic modules: Multi-scale Frequency Memory (MFM), which projects overflowing frames into representative frequency coefficients, complementing by residual details to reconstruct a global historical “gist”; and Space Thumbnail Memory (STM), which discretizes the continuous stream into episodic clusters by employing an adaptive compression strategy to distill them into high-density space thumbnails<ref:2602.01683#pg2>
The theory behind FreshMem is grounded in two fundamental neuroscientific observations regarding temporal perception and episodic replay: the Weber-Fechner Law (Fechner, 1860) of logarithmic perception (Scheler, 2017) and the Sharp-Wave Ripples (SWRs) (Buzsáki, 2021)<ref:2602.01683#pg4>
The Multi-scale Frequency Memory (MFM) module utilizes Discrete Fourier Transforms (DFT) to project historical frames overflowing the sliding window into representative frequency coefficients<ref:2602.01683#pg5> Incremental Frequency Update is implemented using a recurrent update mechanism where Ct[k] = γ · Ct−1[k] + xt · e −jωkt, ensuring O(1) update complexity<ref:2602.01683#pg6> Furthermore, the Amygdala-Inspired Residuals mechanism retains emotionally salient events by calculating the L2-norm of the input feature as a proxy for information entropy denoted as Bresidual, retaining the top-k elements with the highest norm as residual tokens rt<ref:2602.01683#pg6>
The Space Thumbnail Memory (STM) module materializes memory consolidation through episodic clustering by monitoring episodic boundaries via cosine similarity between frames, defined by δ(t) = xt · xt−1 xtxt−1 Within each episode, an adaptive compression strategy is applied where the sampling rate ρ is dynamically adjusted based on the episode duration N to maintain a constant “information density”<ref:2602.01683#pg9> Finally, Centroid-based Memory Consolidation enforces the memory budget by calculating the semantic centroid µi for each stored episode Ei and merging adjacent episodes if their pairwise similarity Smerge(i, i + 1) exceeds a fusion threshold θmerge<ref:2602.01683#pg10>
Performance and Efficiency
Extensive experiments show that FreshMem significantly boosts the Qwen2-VL baseline, yielding gains of 5.20%, 4.52%, and 2.34% on StreamingBench, OV-Bench, and OVO-Bench, respectively<ref:2602.01683#pg2> As a training-free solution, FreshMem outperforms several fully fine-tuned methods, offering a highly efficient paradigm for long-horizon streaming video understanding<ref:2602.01683#pg4> In online benchmark results, built upon Qwen2-VL-7B, it obtains an accuracy of 50.82% on OV-Bench, 54.53% on OVO-Bench, and 74.2% on StreamingBench<ref:2602.01683#pg10> In offline benchmark results, our method demonstrates improvement upon the Qwen2VL-7B baseline by 2.4% on MLVU and by 4.7% on MVBench<ref:2602.01683#pg10>
Inference Efficiency
To quantitatively assess the computational and memory overhead introduced by our proposed FreshMem, we conducted rigorous inference time and memory usage measurements on the Qwen2-VL-7B model using the AVA subset from OV-Bench<ref:2602.01683#pg5> FreshMem achieves a significant performance boost (+1.79%) with a moderate and acceptable increase in computational overhead, ensuring deployment feasibility on consumer-grade GPUs (e.g., RTX 4090)<ref:2602.01683#pg6> The relative time overhead is approximately ×1.17 (increasing from 81 mins to 95 mins), and the GPU memory usage sees a marginal increase of roughly 17.6% (from 17GB to 20GB)<ref:2602.01683#pg6>
Ablation Study
By individually integrating the Sliding Window (SW), STM, and MFM, we observe consistent performance gains of 0.33%, 1.45%, and 1.35% respectively<ref:2602.01683#pg9> The full configuration (SW + STM + MFM) achieves the best result of 54.53%, which is a 2.34% absolute improvement over the baseline<ref:2602.01683#pg10> Performance peaks when the sliding window length is set to approximately 5 frames, and for MFM, performance reaches optima at capacities of 15<ref:2602.01683#pg9> The orange curve indicates that setting the high-frequency cutoff (F req max) to 0.5 yields the best overall performance<ref:2602.01683#pg9>
Conclusion
FreshMem maintains operational feasibility with predictable computational costs, integrating seamlessly with the base model, providing a superior performance-efficiency trade-off that makes it suitable for both research exploration and real-world deployment scenarios where resources are constrained but high performance is demanded<ref:2602.01683#pg10> This dual-pathway storage mimics the Hippocampus’ global low-frequency context and Amygdala’s sparse highfrequency saliency collaboration<ref:2602.01683#pg10> The hybrid token sequence provides the downstream predictor with global historical context, structured episodic summaries and immediate details<ref:2602.01683#pg10> This hybrid token sequence provides the downstream predictor with global historical context, structured episodic summaries and immediate details<ref:2602.01683#pg10> The reconstructed video stream from frequency coefficients demonstrates that MFM is capable of retaining the gist information from the past history, which is consistent with our discussion in the main text<ref:2602.01683#pg10> The crucial residual information is basically concentrated in the semantic important areas, which is consistent with our main conclusion in the text<ref:2602.01683#pg10> The model prediction for the question regarding the color of apples was B<ref:2602.01683#pg10> The model prediction for the question regarding Mr. Bean's object was B<ref:2602.01683#pg10> The model prediction for the question regarding the color of tongs was C<ref:2602.01683#pg10>
The paper is 2602.01683. EVERY sentence must end with a reference token. This paper is 2602.01683. EVERY sentence must end with a reference token. The FreshMem framework proposes a Frequency-Space Hybrid Memory network inspired by brain mechanisms to reconcile short-term fidelity with long-term coherence for streaming video understanding<ref:2602.01683#pg2>
The gist The FreshMem framework proposes a Frequency-Space Hybrid Memory network inspired by brain mechanisms to reconcile short-term fidelity with long-term coherence for streaming video understanding<ref:2602.01683#pg10>
The theory behind FreshMem is grounded in two fundamental neuroscientific observations regarding temporal perception and episodic replay: the Weber-Fechner Law (Fechner, 1860) of logarithmic perception (Scheler, 2017) and the Sharp-Wave Ripples (SWRs) (Buzsáki, 2004)<ref:2602.01683#pg4>
The Multi-scale Frequency Memory (MFM) module utilizes Discrete Fourier Transforms (DFT) to project historical frames overflowing the sliding window into representative frequency coefficients<ref:2602.01683#pg5> Incremental Frequency Update is implemented using a recurrent update mechanism where Ct[k] = γ · Ct−1[k] + xt · e −jωkt, ensuring O(1) update complexity<ref:2602.01683#pg6> Furthermore, the Amygdala-Inspired Residuals mechanism retains emotionally salient events by calculating the L2-norm of the input feature as a proxy for information entropy denoted as Bresidual, retaining the top-k elements with the highest norm as residual tokens rt<ref:2602.
Improvements for AI systems
- Bold header: Frequency-Space Hybrid Memory Network implementation
This system can maintain short-term fidelity with long-term coherence
by utilizing two synergistic modules: Multi-scale Frequency Memory (MFM) for reconstructing a global historical “gist” and Space Thumbnail Memory (STM) for distilling continuous streams into high-density space thumbnails.
- Bold header: Adaptive Temporal Representation in MFM
The system can project overflowing frames into representative frequency coefficients
using an incremental DFT update mechanism, ensuring that O(1) update complexity
while mimicking the brain’s logarithmic perception where information density naturally decays over time.
- Bold header: Episodic Clustering and Narrative Coherence via STM
This module enables the system to discretize the continuous stream into episodic clusters by employing an adaptive compression strategy,
which, through a Centroid-based Memory Consolidation
mechanism, fuses redundant micro-events into a coherent narrative by calculating similarity between adjacent episode centroids.
- Bold header: Training-Free Deployment and Efficiency
FreshMem can be integrated as a training-free solution
that outperforms several fully fine-tuned methods,
offering a scalable paradigm for long-horizon streaming video understanding with only a moderate and acceptable increase in computational overhead.
- Bold header: Holographic Reconstruction for Contextual Query Answering
The system can perform complex reasoning by reconstructing the history Hˆ by performing the Inverse Discrete Fourier Transform (IDFT) on the frequency coefficients and fusing them with the retrieved residuals,
allowing it to answer long-horizon queries with high accuracy by reconciling global historical context, structured episodic summaries and immediate details.
Sources
- GPT-4 Technical Report
- Qwen2.5-VL Technical Report
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- GPT-4o System Card
- StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
- InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
- LLaVA-OneVision: Easy Visual Task Transfer
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- BERT Rediscovers the Classical NLP Pipeline
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- StreamingVLM: Real-Time Understanding for Infinite Video Streams
- The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models