PERK: Long-Context Reasoning as Test-Time Learning
cs.CL, cs.LG
Submitted: 2025-07-08
Updated: 2026-08-31
Comments: ICLR 2026, 10 pages, 7 figures, test-time learning, long context, meta-learning
Journal ref: The Fourteenth International Conference on Learning Representations, 2026
Code: https://github.com/facebookresearch/higher
License: http://creativecommons.org/licenses/by/4.0/
The gist: Long-context reasoning requires accurately identifying relevant information in extensive, noisy input contexts.
Terminology
Abstract
Long-context reasoning requires accurately identifying relevant information in extensive, noisy input contexts. In this work, we propose PERK (Parameter Efficient Reasoning over Knowledge), a scalable approach for learning to encode long contexts using gradient updates at test time. Specifically, PERK employs two nested optimization loops in a meta-training phase. The inner loop rapidly encodes contexts into a low-rank adapter (LoRA) that serves as a parameter-efficient memory module for the base model. Concurrently, the outer loop learns to use the updated adapter to accurately recall and reason over relevant information from the encoded long context. Our evaluations on several long-context reasoning tasks show that PERK significantly outperforms the standard long-context finetuning, achieving average absolute performance gains of up to 20% for Qwen-2.5 (0.5B & 7B) on synthetic and real-world long-context reasoning. PERK also maintains its advantages across model scales and families. Compared to specialized long-context LLMs, PERK matches or surpasses their performance. Finally, our analyses show PERK is more robust to reasoning complexity, length extrapolation, and the positions of relevant information in contexts. https://perk-long-context.web.app
Sources
- Simple linear attention language models balance the recall-throughput tradeoff
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
- xLSTM: Extended Long Short-Term Memory
- Titans: Learning to Memorize at Test Time
- ATLAS: Learning to Optimally Memorize the Context at Test Time
- In-Context Learning with Long-Context Models: An In-Depth Exploration
- Extending Context Window of Large Language Models via Positional Interpolation
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
- Data Engineering for Scaling Language Models to 128K Context
- How to Train Long-Context Language Models (Effectively)
- The Llama 3 Herd of Models
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- HyperAttention: Long-context Attention in Near-Linear Time
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Online Adaptation of Language Models with a Memory of Amortized Contexts
- Meta-Learning Online Adaptation of Language Models
- Repeat After Me: Transformers are Better than State Space Models at Copying
- LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
- A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering