Test-Time Training with Next-Token Prediction
cs.CL
Submitted: 2026-06-19
Updated: 2026-08-30
Comments: 17 pages, 2 figures, 6 tables. Accepted to Findings of EMNLP 2026. Code: https://github.com/yancyou/TTT-NTP
Code: https://github.com/yancyou/TTT-NTP
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- Longformer: The Long-Document Transformer
- Extending Context Window of Large Language Models via Positional Interpolation
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
- In-Place Test-Time Training
- Data Engineering for Scaling Language Models to 128K Context
- The Llama 3 Herd of Models
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Measuring Massive Multitask Language Understanding
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Mistral 7B
- Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
- Hopfield Networks is All You Need
- Linear Transformers Are Secretly Fast Weight Programmers
- Learning to (Learn at Test Time): RNNs with Expressive Hidden States
- Retentive Network: A Successor to Transformer for Large Language Models
- End-to-End Test-Time Training for Long Context
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering