TTSR: Test-Time Self-Evolving via Reflection
cs.CL, cs.AI, cs.LG
Submitted: 2026-02-06
Updated: 2026-09-17
Comments: EMNLP 2026 Main Conference
License: http://creativecommons.org/licenses/by/4.0/
The gist: Test-time training (TTT) adapts large language models (LLMs) during inference using only unlabeled test inputs.
Terminology
Abstract
Test-time training (TTT) adapts large language models (LLMs) during inference using only unlabeled test inputs. Existing methods, however, face two major bottlenecks on hard reasoning tasks: (1) lack of learnable samples, as self-generated pseudo-labels on difficult questions are often noisy and yield unstable rewards; and (2) inefficient exploration, as performance gains depend on repeatedly sampling many rollouts without explicit diagnosis of why previous attempts fail. We propose TTSR (Test-Time Self-Reflection), a self-evolving framework based on a reflect-then-synthesize paradigm. A single pretrained model alternates between a Student role and a Teacher role: the Student solves test questions and updates, while the Teacher analyzes failed trajectories and synthesizes targeted variant questions closer to the Student's capability frontier. TTSR further maintains a cross-iteration weakness memory and compiles persistent weaknesses into a lightweight strategy note prepended to subsequent Student inputs, so diagnostic knowledge can guide exploration and gradually fade as weaknesses are resolved. Experiments on challenging mathematical reasoning benchmarks show consistent test-time improvements, strong cross-backbone generalization, and transfer to general-domain reasoning tasks.
Sources
- The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
- VisPlay: Self-Evolving Vision-Language Models from Images
- Large Language Models Cannot Self-Correct Reasoning Yet
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning
- Guided Self-Evolving LLMs with Minimal Human Supervision
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering