Compressing History into Memory: Distilling Transformers into Recurrent Transformers
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Compressing History into Memory".
Tom: Recurrent Transformers address computational limitations in long-horizon sequential data processing by distilling the compression strategy from full-history transformers into fixed-size memory.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, to recap what Weinzaepfel et al. are proposing in "Compressing History into Memory: Distilling Transformers into Recurrent Transformers," they are focusing on a distillation approach. Their main thesis is that you can transfer the compression mechanism from a full-history transformer to a recurrent variant by creating a teacher model that explicitly compresses history into a fixed-size bottleneck representation, and then using that representation to supervise the student's memory.
Jane: That means they design this teacher model, called the Latent Bottleneck History Transformer or LBHT, which learns how to condense all past observations into these contextualized tokens that act like a latent map of the scene. Then, they use this as a direct supervisor for their student model, which is a recurrent transformer variant.
Lu: It’s fascinating that they frame the problem as transferring a compression strategy rather than just trying to make the recurrent model perform better on its own; it’s about leveraging prior knowledge of effective compression. This approach directly addresses the difficulty recurrent models face when deciding what to keep in their limited memory bank at each step without knowing what future information will be needed.
Meng: From a practical standpoint, it sounds like they are solving the explicit decision-making problem by giving the student a pre-learned, high-quality way to compress past observations into a manageable form. That makes the learning task for the student much less ambiguous than just letting it learn compression from scratch in its constrained recurrent loop.
Lalam: For me, this means we are moving toward AI that doesn't just passively store data but actively learns to summarize and distill what is truly essential from a long sequence of inputs into a compact, meaningful representation. That ability to maintain a stable core of memory tokens while capturing dynamic elements sounds like it could really enhance how our culture interacts with complex digital environments.
Conclusion: Tom: So, wrapping up this discussion on "Compressing History into Memory: Distilling Transformers into Recurrent Transformers," the authors are showing how we can bridge the performance gap between computationally expensive full-history transformers and memory-based recurrent models by using this specific distillation technique.
Jane: Essentially, they've shown that by explicitly training a teacher to compress history into a fixed memory bank, we can teach a recurrent model how to do the same efficiently, allowing it to keep its fast processing speed while achieving performance comparable to much larger transformer models on long sequences.
Lu: The implication here is that we might be able to deploy very capable sequential AI agents in environments where keeping a complete history is physically impossible due to memory constraints, which opens up new avenues for robotics and complex scene understanding.
Meng: For us at the engineering side, this means we can build systems that are both powerful and lightweight enough for real-time use, which is a huge step forward in making advanced AI practical outside of super-computing centers.
Lalam: I think the bigger picture here is about building more robust and context-aware AI systems that can handle long narratives or complex visual scenes without getting overwhelmed by the sheer volume of past data, leading to richer and more nuanced interactions with our digital world.
Naver Labs Europe
cs.CV, cs.LG
Submitted: 2026-06-19
Updated: 2026-10-01
Code: https://github.com/wgw101/habitat_sim2real
Importance score: 76/100
The gist: Recurrent Transformers address computational limitations in long-horizon sequential data processing by distilling the compression strategy from full-history transformers into fixed-size memory.
Key concepts
- Full-History Transformer (HT)
- A standard transformer architecture that processes all past observations in a sequence. While powerful, it suffers from quadratic computational complexity ($O(T^2)$) when handling long sequences, making it impractical for real-time applications requiring large histories.
- Recurrent Transformer (RT)
- A model designed for sequential data that maintains a fixed-size memory bank instead of processing the entire history at every step. This allows for linear time complexity ($O(1)$ per step) but often lags behind full-history models in performance due to difficulties in explicitly deciding what past information to retain.
- Latent Bottleneck History Transformer (LBHT)
- The teacher model that explicitly learns how to compress a long sequence of observations into a small, fixed-size latent representation. This learned representation acts as a 'map' of the scene, teaching the student model the optimal way to summarize history.
Terminology
Summary
Recurrent Transformers address computational limitations in long-horizon sequential data processing by distilling the compression strategy from full-history transformers into fixed-size memory. This work proposes a distillation approach that transfers the compression mechanism of a classical full-history transformer to a recurrent variant, enabling recurrent models to retain linear-time complexity while substantially narrowing the performance gap with full-history transformers.
The gist
We propose a distillation approach that transfers the compression strategy of a classical full-history transformer to a recurrent variant by designing a teacher model that explicitly compresses its observation history into a fixed-size bottleneck representation, and then directly supervising the student’s memory with this bottleneck representation.
Motivation and Problem Context
Transformers are powerful for sequential data but suffer from quadratic complexity, O(T2), when processing long sequences, making them impractical for applications like map-free pose estimation where storing and maintaining a history of observations is prohibitive. Recurrent Transformers (RTs) offer O(1) complexity at each step by maintaining a fixed-size memory bank, but their performance lags behind full-history Transformers (HTs). The authors argue this gap stems from differences in how they learn to compress past information: HTs implicitly compress history via attention and gradients, whereas RTs must explicitly decide at each time step what information to store in their fixed-size memory without having access to future requirements for prediction,
making the learning problem significantly harder.
Proposed Method: Distillation Architecture
The core of the proposal is to bridge this gap by turning memory learning into a supervised compression task. The method involves two main components:
-
A Teacher model, the Latent Bottleneck History Transformer (LBHT), which explicitly compresses its observation history into a fixed-size bottleneck representation, similar to the Perceiver Resampler architecture. This teacher learns how to take a full sequence of observations and compress them into a memory bank in the form of contextualized tokens, interpreted as a
latent representation, or latent 'map', of the scene.
-
A Student model, implemented as a Recurrent Transformer (RT), which maintains and updates a fixed-size memory state Mt using an update equation: Mt = fθ(Mt−1, ot). The student must learn a
smart lossy compression mechanism
to decide which information from the current observation ot needs to be retained in its limited memory.
Distillation Strategy
The distillation process aligns the two models by ensuring they share the same latent representation dimensions. Key steps include:
-
Aligning dimensions: The teacher’s latent representation B˜t and the student’s memory representation Mt are aligned using an identical number of embeddings and embedding dimensions, allowing direct distillation via an L1 loss between the flattened representations B˜t and Mt.
-
Segment-wise distillation: To best benefit from how the teacher learned to compress sequences of different lengths, a rollout of the student over a full sequence is decomposed into N segments. The final end memory bank Men for each segment is distilled with an L1 loss targeting the teacher representation B˜en obtained by a teacher having observed that specific sub-sequence.
Results and Findings
The approach, called Chimera for Compressing History Into Memory with Transformers, was validated on a long-horizon map-free pose estimation task. The experiments confirm two primary hypotheses:
-
"It is easier to learn to compress a full sub-sequence, i.e. the compression of a full sub-sequence, i.e., it is easier to learn to compress a full sub-sequence, i.e., that it is easier to learn to compress a full sub-sequence."
-
Once the LBHT teacher is trained and applied to sequences of increasing lengths T, thus obtaining sequences of bottleneck representations,
these sequences can be sufficiently approximated through a learned recurrence function.
The resulting Chimera model retains the favorable O(1) inference complexity of recurrent models while achieving performance levels previously reserved for computationally expensive transformers. Chimera significantly outperforms the state-of-the-art recurrent models, including Kinaema, and performs close to the LBHT teacher across various query and memory timesteps. Furthermore, analysis shows that the distilled model develops a more skewed distribution of memory tokens, suggesting it learns to maintain a stable core of memory tokens that are updated less frequently while capturing discriminative information in a few highly dynamic tokens.
Architectural Details
The student architecture is based on the Kinaema model, featuring a recurrent memory update step and a decoding step that queries memory via cross-attention. The teacher (LBHT) uses 7 transformer layers compared to the student's 3 layers, and both architectures were optimized. Observation encoding for both models involves a fine-tuned DINO-v2 ViT-s backbone, where visual embeddings are down-projected to 64 values, concatenated with odometry embeddings.
Improvements for AI systems
Based on the provided paper, here are the specific improvements to AI systems that can be achieved by implementing its proposed architecture, Chimera:
-
Improve performance in long-horizon streaming vision and robotics applications (e.g., map-free pose estimation).
-
Enable agents to accurately answer complex visual queries about their environment relative to a moving coordinate frame (
Where is this thing with respect to your current position?
). -
Achieve efficient, linear-time complexity inference for memory-based tasks, overcoming the quadratic complexity limitations of standard full-history Transformers in sequential processing.
-
Reduce the performance gap between recurrent models (Recurrent Transformers) and full-history Transformers by explicitly learning effective compression strategies rather than relying on implicit history access.
-
Develop agents that can maintain a fixed-size, relevant memory bank of past observations without needing to store or attend over the entire observation history, leading to more scalable and deployment-friendly models (O(1) update cost).
-
Improve the stability and generalization of learned memory representations by enforcing a compression mechanism distilled from a high-capacity teacher model (Latent Bottleneck History Transformer), resulting in memory states with better temporal stability.
-
Create models that learn to selectively retain critical information, allowing them to maintain a
stable core
of persistent memory tokens while capturing highly dynamic information in specialized, frequently updated tokens.
In summary, the improved AI system (Chimera) will be a more efficient and capable agent for real-world embodied AI tasks where long sequences of observations are common but full history processing is computationally prohibitive.
Sources
- S-MUSt3R: Sliding Multi-view 3D Reconstruction
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences
- On Predictive Information in RNNs
- Partially Observable Reinforcement Learning with Memory Traces
- DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
- Neural Turing Machines
- VGGT-SLAM 2.0: Real-time Dense Feed-forward Scene Reconstruction
- BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning
- Data Efficient Any Transformer-to-Mamba Distillation via Attention Bridge
- Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
- LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models