Prefix Sliding for efficient test-time scaling
cs.CL, cs.AI, cs.LG
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 28 pages (9 main), 22 figures, 3 tables
Code: https://github.com/Muennighoff/prefix-sliding
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
- The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
- ETC: Encoding Long and Structured Inputs in Transformers
- TokenButler: Token Importance is Predictable
- Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Language Models are Few-Shot Learners
- Recurrent Memory Transformer
- R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
- Value-Aware Stochastic KV Cache Eviction for Reasoning Models
- TheoremQA: A Theorem-driven Question Answering dataset
- Reasoning Models Don't Always Say What They Think
- NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time
- Adapting Language Models to Compress Contexts
- Generating Long Sequences with Sparse Transformers
- Rethinking Attention with Performers
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Composer 2 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering