Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRA
cs.AI
Submitted: 2026-09-22
Updated: 2026-09-22
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Federated Sketching LoRA: A Flexible Framework for Heterogeneous Collaborative Fine-Tuning of LLMs
- Measuring Massive Multitask Language Understanding
- Measuring Mathematical Problem Solving With the MATH Dataset
- Angles between subspaces and their tangents
- Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
- Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
- On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training
- SocialIQA: Commonsense Reasoning about Social Interactions
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- RL's Razor: Why Online Reinforcement Learning Forgets Less
- Self-Distillation Enables Continual Learning
- ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
- Denser not equal to Better: Limits of On-Policy Self-Distillation for Continual Post-Training
- Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
- HellaSwag: Can a Machine Really Finish Your Sentence?
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection