Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It
cs.CL
Submitted: 2026-06-09
Updated: 2026-08-30
Code: https://github.com/LARK-AI-Lab/QK-Restore
Terminology
Sources
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes
- Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
- Training Verifiers to Solve Math Word Problems
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- OpenThoughts: Data Recipes for Reasoning Models
- RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Repeat After Me: Transformers are Better than State Space Models at Copying
- Mistral 7B
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
- Distilling to Hybrid Attention Models via KL-Guided Layer Selection
- TL;DR: Too Long, Do Re-weighting for Efficient LLM Reasoning Compression
- Let's Verify Step by Step
- NVIDIA Nemotron 3: Efficient and Open Intelligence
- Training language models to follow instructions with human feedback
- Why think step by step? Reasoning emerges from the locality of experience
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering