Unsupervised Post-Training of Foundation Models: A Survey
cs.CL, cs.AI, cs.CV, cs.LG, cs.MM
Submitted: 2026-08-25
Updated: 2026-08-27
Terminology
Sources
- Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
- Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
- Self-Improvement in Multimodal Large Language Models: A Survey
- Distribution-Aware Reward Estimation for Test-Time Reinforcement Learning
- Dual Consensus: Escaping from Spurious Majority in Unsupervised RLVR via Two-Stage Vote Mechanism
- Data Engineering for Scaling Language Models to 128K Context
- LongMagpie: A Self-synthesis Method for Generating Large-scale Long-context Instructions
- AERO: Autonomous Evolutionary Reasoning Optimization via Endogenous Dual-Loop Feedback
- One-shot Entropy Minimization
- G-Zero: Self-Play for Open-Ended Generation from Zero Data
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- Simple and Scalable Strategies to Continually Pre-train Large Language Models
- Model Whisper: Steering Vectors Unlock Large Language Models' Potential in Test-time
- TTSR: Test-Time Self-Evolving via Reflection
- Continual Pre-training of Language Models
- Self-Calibrating Language Models via Test-Time Discriminative Distillation
- LLM Post-Training: A Deep Dive into Reasoning Large Language Models
- SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
- SLOT: Sample-specific Language Model Optimization at Test-time
- Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering