Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation".
Jane: Reasoning language models (RLMs) have shown impressive capabilities in domains like mathematics and coding,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Welcome back to the show folks! Today we’re diving into a paper that’s been making waves in the AI community. We have "Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation," and I'm really excited to break down what this means for how we train these powerful language models.
Jane: It sounds like a really interesting approach, Tom. This paper is looking at how we can get those impressive reasoning abilities out of models even when we don't have perfect verification traces during training. It seems to tackle a real challenge in the current AI landscape where relying only on input-output pairs can be tricky for complex tasks.
Lu: What caught my eye about this paper, "Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation," is how they address the difficulty of training models when there aren't reliable reasoning traces available, especially in areas like mathematics and coding where verification is crucial. It’s smart that they are looking at leveraging existing instruction fine-tuning data for this purpose.
Meng: From an engineering standpoint, the idea of a two-step pipeline sounds promising, but I'm curious about how robust this method is when applied to very different tasks. Can we really expect it to work well across both verifiable and harder-to-verify domains simultaneously?
Lalam: I think the core idea is powerful because it suggests we don't need perfect reasoning traces upfront to get a model performing better on specific tasks, which could significantly speed up development cycles for new applications.
Tom: Exactly, Lalam. The paper claims they show that this technique improves RLM performance in both verifiable and hard-to-verify domains, specifically mentioning coding and text summarization. Jane, can you explain what the authors are actually proposing they do?
Jane: Certainly. Essentially, the paper proposes a lightweight two-step pipeline for adapting Reasoning Language Models. First, they use standard instruction fine-tuning without reasoning traces on a training set to create a model that's better at the target task. Then, they merge this instruction-tuned model with the original reasoning model to recover the original reasoning behavior on that specific target task.
Lu: That two-step pipeline is clever because it separates the task adaptation from the core reasoning preservation step. It seems to be a practical way to handle this distributional mismatch they talked about, where training data doesn't explicitly show the reasoning steps taken.
Paper summary: Meng: So, if we look at the specifics of their method, how do they ensure that when they merge these models, the general reasoning capabilities aren't completely wiped out? What mechanism keeps that balance?
Lalam: The paper discusses selecting a merge ratio using a held-out target-task calibration set to maximize the preservation of reasoning behavior on that task. They define this preservation by looking at a reasoning rate, where it's the fraction of examples for which the model produces a non-empty reasoning trace.
Tom: That selection process sounds like the real trick here. So, how do they quantify what "acceptable" means when choosing that merge ratio, and what is the minimum threshold they use?
Jane: They select the optimal merge ratio, alpha*, by maximizing preservation while keeping the reasoning rate above a minimum acceptable level, which they set to zero point nine in their experiments. This allows them to find a point where the model is good at the new task but still retains enough of its original thinking ability.
Lu: It’s interesting that they focused on selecting alpha* based on reasoning behavior, rather than just accuracy on the target task. That shows they prioritized the core capability recovery over just chasing higher scores on the specific dataset.
Meng: I see the focus there. It makes sense because if we only focused on accuracy, we might fine-tune away the very reasoning skills we want to keep for general problems later. Does this technique hold up across different types of models they tested?
Lalam: They evaluated this method across four different Reasoning Language Models, including OpenThinker 7B and DeepSeek R1 Qwen 7B Distilled. The ablation studies also confirm that the method is robust to variations in merging algorithms and fine-tuning approaches.
Tom: That robustness is something I appreciate when we’re deploying these things, Meng. So, across those four models and two tasks—Rust coding and text summarization—what did the results actually show in terms of performance gain?
Jane: In Rust coding, their merging technique resulted in the highest average increase of seven point zero percentage points compared to just using instruction fine-tuning alone, which only gave them three point eight. For text summarization, it restored target-task reasoning and MATH500 almost completely while keeping ninety-five point five percent of the instruction fine-tuning gain.
Paper summary: Lu: Those figures are substantial, especially seeing that seven point zero percentage point lift in coding compared to just the IFT method alone. It suggests that this method isn't just incrementally better; it’s providing a significant boost when we need specialized reasoning skills.
Meng: From a practical deployment view, knowing that they achieved this with less than USD three on average is also quite relevant. The cost-effectiveness of the entire process seems pretty solid for real-world use cases.
Lalam: I think that cost aspect makes a big difference for widespread adoption, Meng. If we can achieve these kinds of improvements efficiently, it opens the door for more diverse and specialized AI applications across various industries.
Tom: It really does, Lalam. So, to wrap up this part of the discussion on "Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation," what’s the big picture takeaway here? What is this paper ultimately telling us about how we should approach training models without perfect reasoning traces?
Jane: The main point is that you can use widely available instruction fine-tuning data to get a model performing better on new tasks, and then use a smart merging process to recover its core reasoning abilities while retaining much of the task-specific gains.
Lu: It suggests that we don't have to wait for perfect verification traces during initial training to leverage supervised data effectively for improving model performance in complex domains like coding.
Meng: For practical implementation, it gives us a clear path: fine-tune first, then use calibration data to guide the merge ratio selection. It’s a structured way to handle this adaptation problem.
Lalam: I think this work opens up a lot of possibilities for developing more versatile and capable AI agents that can tackle specialized tasks efficiently, which is really exciting for the culture of our development team.
Tom: Fantastic stuff, everyone. We’ve seen how this paper uses instruction tuning and model merging to adapt RLMs efficiently while keeping their reasoning intact. That gives us a solid foundation for thinking about next steps in training these systems. We’ll be right back after the break with more AI research updates on arXiv.
Conclusion: Tom: So, we've seen how this paper uses instruction tuning and model merging to adapt RLMs efficiently while keeping their reasoning intact. Now, let's talk about the conclusion of this work and what it really means for us.
Jane: Exactly, Tom; the authors are basically saying that you don't need perfect reasoning traces during training to get a model doing better on specific tasks if you use instruction fine-tuning data. It's about using what we already have to make models more useful without needing those super detailed verification steps right away.
Lu: I think the core message is that this pipeline lets us separate the task adaptation from preserving the model's fundamental reasoning skills, which is a really neat structural idea. It opens up new ways to think about how we can guide models toward specific capabilities using less direct data.
Meng: From an engineering standpoint, it suggests a much more flexible way to deploy AI systems because we don't have to start from scratch with perfect reasoning logs for every single application. That flexibility is what makes this technique interesting for real-world startup development.
Lalam: For the future of AI culture, this means we can build applications that are much more specialized and robust because they are less dependent on perfect, traceable training data. It pushes us toward building models that work better with the messy reality of real-world instruction sets.
Tom: So, to sum it up, the authors have created a structured way—fine-tune then merge—to get specialized performance without sacrificing the core reasoning engine. It's about making models smarter faster using existing data structures.
Jane: That's right, Tom; it gives us a practical blueprint for how to adapt these complex reasoning models to new domains in a more manageable way than before.
Lu: And I think the authors' focus on preserving the reasoning rate above zero point nine provides a very concrete metric for us to track when we are performing this adaptation. That level of control over that balance is quite significant for theoretical exploration.
Meng: It gives us a clear path on how to reduce the computational overhead while still getting meaningful gains, which is exactly what we need when deploying these systems at scale.
Lalam: I really see this as a step toward building AI that can learn from instruction sets much more naturally, which will fundamentally improve how we interact with these powerful reasoning tools.
Tom: Alright, so what's next for this research? We need to know what the authors are looking at next to push this adaptation even further.
Jane: The paper mentions that they tested this across several models and tasks, so it's worth seeing if they can generalize these findings to even more diverse architectures in their future work.
Lu: I suspect the next steps involve exploring how this merging mechanism interacts with different types of model distillation methods, which could lead to even more nuanced control over behavior.
Meng: I'm curious if they plan to look at how these reasoning rates change when we introduce even more complex, non-standard instructions into the fine-tuning process. That's where the real engineering challenge lies.
Lalam: If they can show that this approach works reliably across a wider variety of model types, it really validates the idea that instruction data is a powerful tool for general AI development.
Department of Computer Science, ETH Zurich
cs.LG, cs.CL
Submitted: 2026-07-16
Updated: 2026-10-01
Code: https://github.com/eth-sri/rlm-training-merging
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: Reasoning language models (RLMs) have shown impressive capabilities in domains like mathematics and coding, but their performance gains are often limited because training them on tasks lacking
Key concepts
- Instruction Fine-Tuning (IFT)
- This is the first step where the base language model is trained on a specific task's training data. Crucially, the target answers are formatted with an empty reasoning trace followed by the desired output, teaching the model how to structure its thinking for that task.
- Model Merging
- This technique combines two sets of model weights using a linear interpolation formula. It is used to recover reasoning behavior lost during fine-tuning by blending the weights of two models, $ heta_1$ and $ heta_2$, controlled by a merge ratio $\alpha$. The goal is to find the best blend that keeps reasoning intact.
- Reasoning Rate ($ ho(M', D)$)
- This metric quantifies how well a model performs on a target task. It is defined as the fraction of examples where the model produces a non-empty reasoning trace. The optimal merge ratio $\alpha^*$ is chosen to ensure this rate stays above a minimum threshold, like 0.9, guaranteeing acceptable reasoning quality.
- Calibration Set ($D_{cal}$)
- This is a held-out set of target tasks used specifically to select the best model merging coefficient $\alpha$. By testing different merge ratios on this set, the method finds the point where the merged model balances high task performance with sufficient reasoning preservation.
Terminology
Summary
Reasoning language models (RLMs) have shown impressive capabilities in domains like mathematics and coding, but their performance gains are often limited because training them on tasks lacking reliable verifiers remains challenging. This work proposes a lightweight two-step pipeline—standard instruction fine-tuning followed by model merging using a calibration set—to adapt RLMs to new tasks efficiently while preserving their core reasoning abilities.
Core Methodology
The proposed method utilizes a two-step pipeline to mitigate the distributional mismatch that occurs when training RLMs on input-output pairs without explicit reasoning traces. First, the base Reasoning Language Model (RLM) M is fine-tuned using standard Instruction Fine-Tuning (IFT) on a task-specific training set Dtrain, where each target is serialized as an empty reasoning trace followed by the desired output: We set the target answer to y(o) = serial(ε, o), rendering the answer with model-native thinking trace and final answer formatting.
This results in a fine-tuned model MIFT. The second step involves linear merging, where a coefficient α is selected using a held-out target-task calibration set Dcal such that the merged model Mα remains close to MIFT while preserving the reasoning behavior of M on the target task, as measured by non-empty reasoning traces.
Model Merging and Ratio Selection
Model merging is employed to recover forgotten reasoning behavior, combining two sets of model weights using linear interpolation: θα = (1 − α)θ1 + αθ2, α ∈ [0, 1].
The merge ratio α is determined by maximizing the preservation of reasoning behavior on the target task. This is quantified by defining a reasoning rate ρ(M′, D) as the fraction of examples for which the model produces a non-empty reasoning trace.
The optimal merge ratio α⋆ is then selected as: "α⋆ = max α∈A
θ: ρmin ≤ ρ(Mα, Dcal)," where ρmin is the minimum acceptable reasoning rate (set to 0.9 in experiments). This procedure selects the largest merge ratio whose target-task calibration reasoning rate remains acceptable.
Evaluation and Performance
The technique is evaluated across four open RLMs—OpenThinker 7B, Apriel Nemotron 15B Thinker, Olmo3 7B Think, and DeepSeek R1 Qwen 7B Distilled—on two target tasks: Rust coding and text summarization. The evaluation focuses on the trade-off between task performance gain from IFT and the preservation of general reasoning capabilities, measured by MATH500. The results demonstrate that the method recovers most or all of the lost reasoning capability while retaining significant parts of the target-task gain from IFT.
For Rust coding, our merging technique resulted in the highest average increase of 7.0 percentage points
compared to only 3.8 due to IFT, while for text summarization, it restores target-task reasoning and MATH500 almost completely while retaining 95.5% of the IFT SummEval gain.
Optimizations and Cost-Effectiveness
To enhance efficiency, two optimizations are applied: first, running calibration only on the first few tokens by observing that when a fine-tuned model no longer reasons, it emits the endof-reasoning token immediately after the start-of-reasoning token.
Second, to reduce computational effort during search in Algorithm 1, we employ binary search instead of a grid search because the reasoning rate decreases monotonically as the merge ratio increases,
allowing us to abort the search as soon as we observe a reasoning rate that is less than 100% and greater than or equal to the minimum threshold ρmin.
This method is noted for being highly cost-effective, requiring less than USD 3
on average.
Ablation and Robustness
Ablations confirm the stability of the method across various settings. Ablating over merging techniques showed that our method reliably picks a point close to a good trade-off between reasoning and task performance.
Furthermore, ablating fine-tuning techniques confirmed that our method is robust to variations of merging algorithms, fine-tuning approaches, and hyperparameter choices,
indicating the general trend of the approach is stable. The study also highlights inconsistencies in reasoning loss across different models (e.g., OpenThinker 7B loses reasoning in all settings, while Olmo3 7B never loses it), suggesting that mechanisms causing loss of reasoning capabilities might depend on how models were originally trained via RLVR or distillation.
Conclusion
The work successfully demonstrates a method for training RLMs without requiring explicit reasoning traces by leveraging widely available IFT datasets.
Improvements for AI systems
Based on the provided scientific paper, Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation,
here are specific improvements that can be made to AI systems using this technique, along with what those improved systems can achieve:
The core improvement offered by this research is a lightweight, cost-effective method to adapt pre-trained Reasoning Language Models (RLMs) to new tasks (like coding or summarization) without requiring expensive verifiers or reasoning traces during fine-tuning.
Here are the specific improvements and their capabilities:
Adaptation of RLMs using Instruction Fine-Tuning (IFT) on Reasoning-Free Data:
The system can be fine-tuned on large amounts of readily available, high-quality task descriptions paired with solutions (input/output pairs) that lack explicit reasoning traces. This allows the model to acquire new task-specific knowledge efficiently.
Recovery of Lost Reasoning Capabilities via Targeted Model Merging:
After IFT, the system recovers its original reasoning behavior by linearly merging the instruction-tuned checkpoint with the original RLM using a carefully calibrated merge ratio (selected on a target-task calibration set).
Improved Performance in Verifier-Free Domains:
The improved RLMs can achieve significant performance gains on domains that traditionally lack reliable verifiers, specifically coding and text summarization, by retaining the model's general reasoning capabilities while acquiring task-specific skills.
Cost-Effective Model Adaptation:
The entire adaptation pipeline (IFT + Merging) is highly cost-effective, requiring less than USD 3 for adaptation on a single H200 GPU, making it accessible for rapid deployment across diverse tasks.
Preservation of General Reasoning Skills:
Unlike standard IFT, which often collapses reasoning behavior, the proposed method ensures that the model retains its general reasoning capabilities (measured by performance on held-out datasets like MATH500) while improving its target-task performance.
Enhanced Code Generation Accuracy:
For coding tasks, the improved system can produce code that correctly handles complex logic and type mismatches (as demonstrated in the case study), achieving higher success rates when compared to standard IFT methods that fail to recognize implicit type conflicts.
Targeted Summarization Quality Improvement:
In text summarization, the system can improve summary quality metrics (like SummEval scores) by retaining a high percentage of the performance gain achieved during IFT, while simultaneously restoring the model's ability to produce coherent and relevant summaries.
The improved AI system (the RLM after adaptation) can perform the following:
-
Generate accurate and functional code in Rust that correctly solves complex programming problems described in natural language, including handling subtle type mismatches.
-
Produce high-quality, contextually relevant, cohesive summaries of long-form text without losing the model's ability to maintain logical flow and consistency.
-
Perform mathematical reasoning tasks (like MATH500) with near-original performance levels despite undergoing task adaptation, indicating robust preservation of foundational logic skills.
-
Adapt to new, specialized domains rapidly and cheaply by leveraging existing instruction-following data rather than requiring expensive reinforcement learning with verifiers or massive amounts of reasoning trace data.
Abstract
Reasoning language models (RLMs) demonstrate impressive performance by leveraging test-time compute in the form of reasoning tokens. However, this behavior makes adapting RLMs to new domains challenging and expensive. The reason is that further training can disturb the learned behavior and degrade model performance. This makes it difficult to leverage supervised fine-tuning data with human-written solutions: although it contains high-quality annotations, it lacks reasoning tokens. In this work, we show how, despite this challenge, such data can be used efficiently for RLM adaptation. For this, we first use standard instruction tuning. Next, we leverage model merging to combine the instruction-tuned model with the original RLM, picking the merging ratio such that the resulting model's reasoning behavior on the target domain is recovered. We evaluate our method across four RLMs on coding and text summarization tasks, where it improves target-task performance by up to 11.0% while preserving reasoning behavior and limiting the out-of-distribution score degradation to on average 0.7%. Importantly, our adaptations are efficient and economical, costing less than USD 10 per model.
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Concrete Problems in AI Safety
- MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation
- Evaluating Large Language Models Trained on Code
- Scaling Instruction-Finetuned Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- ScaleRTL: Scaling LLMs with Reasoning Data and Test-Time Compute for Accurate RTL Code Generation
- Linear Mode Connectivity and the Lottery Ticket Hypothesis
- OpenThoughts: Data Recipes for Reasoning Models
- Let's Verify Step by Step
- On the Impact of Fine-Tuning on Chain-of-Thought Reasoning
- Natural Emergent Misalignment from Reward Hacking in Production RL
- Olmo 3
- OpenAI o1 System Card
- GPT-4 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Self-Distillation Enables Continual Learning
- Learning to summarize from human feedback
- Qwen2.5 Technical Report
- Qwen3 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks