Don't Repeat Yourself: Self-Supervised Fine-Tuning for Coverage
cs.CL
Submitted: 2026-09-16
Updated: 2026-09-30
Terminology
Sources
- Sampling More, Getting Less: Calibration is the Diversity Bottleneck in LLMs
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Evaluating Large Language Models Trained on Code
- SED-SFT: Selectively Encouraging Diversity in Supervised Fine-Tuning
- Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
- Reasoning with Exploration: An Entropy Perspective
- Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- Training Large Language Models To Reason In Parallel With Global Forking Tokens
- The Road Less Traveled: Enhancing Exploration in LLMs via Sequential Sampling
- Understanding the Effects of RLHF on LLM Generalisation and Diversity
- Diversity in Large Language Models under Supervised Fine-Tuning
- Diverse Preference Optimization
- EXAONE 3.5: Series of Large Language Models for Real-world Use Cases
- Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road
- Qwen2.5 Technical Report
- Qwen3 Technical Report
- Learning by Distilling Context
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- 2 OLMo 2 Furious
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering