Reasoning Pattern Alignment Merging for Adaptive Reasoning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "QA-Merging: Query-Adaptive Reasoning via Layer Selective Model Merging".
Jane: The paper was written by Zhaofeng Zhong, Wei Yuan, Tong Chen, Liang Qu, Xiangyu Zhao et al. from The University of Queensland and City University of Hong Kong and Griffith University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, listeners, welcome back. Today we're diving into a paper that's got a mouthful of a title: "Reasoning Pattern Alignment Merging for Adaptive Reasoning." Jane, I gotta say, when I first read that, I thought, okay, that's a lot of buzzwords.
Jane: It is, but the idea behind it is actually pretty simple. You know how some AI models are really good at solving hard math problems by thinking for a long time, and other models are quick and efficient but not as smart?
Tom: Right, like the tortoise and the hare, but for AI.
Jane: Exactly. This paper is about taking those two different kinds of models and merging them into one. So you get a single model that can be the tortoise when it needs to be, and the hare when it doesn't.
Tom: And that's the "adaptive" part. It's not just a dumb average of the two. The authors are from the University of Queensland, City University of Hong Kong, and Griffith University. They're trying to solve a real problem.
Jane: Which is that those smart, slow models, they "overthink." They'll write a five-page essay to solve a problem that only needs a single sentence. It costs a lot of money and takes forever.
Tom: So they're calling this new approach RPAM, which is a much better acronym. And the core idea is to merge a Long-CoT model, that's the slow thinker, with a Short-CoT model, the fast one.
Jane: And they do it in a clever way. They don't just mix the weights of the two models. They actually look at the internal representations of the models, layer by layer, and figure out which one is doing a better job for a given question.
Tom: So it's like a surgical merge, not just a blender. I'm really curious to see how they pull that off. This could be huge for making AI actually usable in the real world.
Jane: It could be. And the results they're showing are pretty impressive. But let's not get ahead of ourselves. We should get into the nitty-gritty of the method.
Tom: Absolutely. Let's get into it.
Summary: Jane: So, Tom, we're back with "Reasoning Pattern Alignment Merging for Adaptive Reasoning." Let's talk about the core problem they're tackling. The paper starts by pointing out that these big reasoning models are amazing, but they have a serious flaw.
Tom: The overthinking problem. They generate these massive chains of thought for everything, even simple arithmetic.
Jane: Right. And the paper says this doesn't just waste time and money, it can actually hurt accuracy. The model gets so caught up in its own rambling that it talks itself out of the right answer.
Tom: So they propose this framework, RPAM, to fix that. And the first step is figuring out which model is better for which question.
Jane: They build a small "pattern-labeled" dataset. They take a bunch of questions, and for each one, they ask both the long-thinking model and the short-thinking model to answer it multiple times. Whichever model gets more answers right is labeled the "positive" model for that question.
Tom: So it's like a teacher grading the two models on each specific problem.
Jane: Exactly. Then, they use this dataset to train the merging process. They don't merge the whole model at once. They go layer by layer, and for each layer, they adjust the merging coefficients.
Tom: And this is where the "alignment" part comes in. They want the merged model's internal thoughts, at each layer, to look like the thoughts of the positive model.
Jane: Right. They use a feature alignment loss to pull the merged model's representations closer to the chosen model. But they also add a "contrastive" loss.
Tom: Which pushes it away from the other model. So it's not just about getting closer to the good one, it's about actively avoiding the bad one.
Jane: Exactly. It's a one-two punch. And the results are pretty wild. On the MATH dataset, they reduced the number of generated tokens by forty-eight percent while actually improving accuracy.
Tom: That's the dream. Less work, better results. And on the harder OlympiadBench, they got a fifty percent reduction in cost with only a tiny drop in accuracy.
Jane: So the summary is they've found a way to make a model that's smart enough to know when to be lazy. That's a big deal.
Tom: It is. But I'm wondering, how does this stack up against other methods? We'll have to look at the experiments.
Improvements: Tom: We're back with "Reasoning Pattern Alignment Merging for Adaptive Reasoning." So, Jane, we've talked about the problem and the method. Now let's get into the details of the experiments. How does RPAM actually perform against the competition?
Jane: They ran it on seven benchmarks, from simple math like GSM8K to the brutal AIME and OlympiadBench. And they compared it to a bunch of other approaches.
Tom: And the big comparison is against other merging methods, right?
Jane: Yeah. There are the simple ones, like just averaging the weights of the two models. Those are fast but they hurt accuracy a lot. Then there are more sophisticated ones, like AIM and ACM, which use activation data to guide the merge.
Tom: And RPAM beats them all?
Jane: In the 4B model setting, RPAM gets an average accuracy of seventy-five point nine, which is only about four point four percent lower than the slow, expensive Long-CoT model. But it does it with forty-eight percent fewer tokens. The other merging methods either had lower accuracy or didn't save as many tokens.
Tom: So it's the best trade-off. But what about the training-based methods? I know there are some that use reinforcement learning to teach a model to be concise.
Jane: Right, they compared against those too. And this is where it gets interesting. The RL methods can get better accuracy, but they don't reduce the response length as much. And they take over twenty hours to train.
Tom: That's a huge cost.
Jane: RPAM, on the other hand, trains in under an hour. It's a lightweight calibration, not a full retraining. So you get a huge efficiency win with a much smaller training bill.
Tom: And they also did an ablation study, right? They showed that each part of their method matters.
Jane: Yes. If you take away the contrastive loss, or the feature alignment, performance drops. And if you replace their carefully labeled dataset with random labels, it gets even worse. So every piece of the puzzle is important.
Tom: So it's not just a clever idea, it's a well-engineered solution. I'm impressed.
Jane: Me too. And the fact that it works on a smaller 1 point 5B model too suggests it's pretty robust.
Tom: So where does this leave us? What's the bigger picture here?
Conclusion: Tom: So, we're wrapping up our discussion on "Reasoning Pattern Alignment Merging for Adaptive Reasoning." Jane, give it to me straight. What's the one thing you want our listeners to remember?
Jane: That we can have our cake and eat it too. We don't have to choose between a smart, slow AI and a fast, dumb one. This paper shows a practical way to get a single model that's both smart and efficient.
Tom: And that's a huge deal for the real world. Think about running these models on your phone, or in a data center where every token costs money. Being able to cut the cost by half while keeping the quality is massive.
Jane: Absolutely. And it's not just about cost. It's about making AI more responsive and usable. No one wants to wait five minutes for an AI to answer a simple question.
Tom: Right. The authors have really nailed the "adaptive" part. It's not a static model, it's one that can change its thinking style on the fly.
Jane: And they did it without needing a massive dataset or a huge training run. That's the part that gets me excited. It's a lightweight solution that could be applied to many different models.
Tom: So, a final thought. This feels like a step towards AI that doesn't just think, but thinks about how to think.
Jane: That's a nice way to put it. It's a step towards more efficient, more practical, and more human-like reasoning.
Tom: Well said. That's all for this paper. It was a great one.
Jane: It really was. Thanks for listening, everyone. We'll be back soon with another exciting paper to break down.
Tom: Until next time, keep thinking, but maybe not too much.
Zhaofeng Zhong, Wei Yuan, Tong Chen, Liang Qu, Xiangyu Zhao, Quoc Viet Hung Nguyen, Hongzhi Yin
The University of Queensland · City University of Hong Kong · Griffith University
cs.CL, cs.AI
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: 16 pages, 4 figures
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 70/100
The gist: The paper introduces Reasoning Pattern Alignment Merging (RPAM), a layer-wise model merging framework designed to achieve query-adaptive reasoning by combining a Long-CoT reasoning model with a
Key concepts
- Long-CoT Model
- This is a slow, highly intelligent AI model that excels at complex tasks but often 'overthinks' simple problems. It generates massive chains of thought, which is accurate but extremely costly and time-consuming.
- Short-CoT Model
- This model represents the fast, efficient side of the AI. It is quick to provide answers but lacks the deep reasoning capability of a Long-CoT model, making it less smart for complex challenges.
- Layer Selective Merging
- Instead of simply blending two models, this technique merges them layer by layer. The system analyzes internal representations to decide which model performs better for a specific question, allowing the adaptive merge to occur at that level.
- Adaptive Reasoning
- This is the core goal of the research. It means creating a single AI system that is not static but can dynamically choose its operational style—be slow and thorough when needed, or fast and efficient when appropriate.
Terminology
Summary
The paper introduces Reasoning Pattern Alignment Merging (RPAM), a layer-wise model merging framework designed to achieve query-adaptive reasoning by combining a Long-CoT reasoning model with a Short-CoT instruction model. The authors state: by combining a long chain-of-thought (Long-CoT) reasoning model with a Short-CoT instruction model, we obtain an adaptive reasoner without training from scratch or requiring large-scale additional data.
The paper addresses the issue that large reasoning models (LRMs) have recently achieved strong performance on complex reasoning tasks,
yet they often generate lengthy reasoning paths for every query, incurring unnecessary computation and latency.
The authors note that while Long-CoT is beneficial for difficult problems, it can be counterproductive on simple tasks that require few reasoning steps: models may 'overthink' by introducing unnecessary intermediate reasoning.
This overthinking not only increases inference cost, but can also hurt accuracy by amplifying the chance of spurious reasoning and obscuring the straightforward solution.
The paper identifies two existing paradigms for adaptive reasoning: "Training-based methods optimize LRMs to exhibit adaptive thought processes via supervised fine-tuning (SFT) or reinforcement learning (RL), but they typically require large-scale data and incur substantial training cost. Training-free methods, often implemented through prompt-guided strategies, can introduce adaptivity without additional optimization, yet they rely heavily on instruction following and can be sensitive to prompt constraints and query phrasing."
RPAM consists of two main components:
The authors construct a small pattern-labeled dataset that empirically compares the effectiveness of the Long-CoT and Short-CoT models on each query.
Given a seed set, they sample k solutions per query from both the Long-CoT model θL and the Short-CoT model θS
and define the empirical expected accuracy of a model θ ∈ θL, θS on input xi.
They then select the one with the higher expected accuracy as the positive model for the query
and designate the other model as the negative model θneg.
This yields a pattern-labeled dataset DPL = (x, θpos, θneg) x ∈ D.
The method learns layer-wise merging coefficients using the constructed pattern-labeled dataset DPL.
For each layer l, the merged model parameters are computed as θM(l) = λL(l)θL(l) + λS(l)θS(l). The authors minimize the squared l2 distance between the merged feature and the positive model feature
using the loss Lalign = ‖zM(l) − zpos(l)‖2.
The paper introduces a lightweight contrastive objective that pulls the merged model's representation closer to the positive model's while pushing it away from the negative model's.
The contrastive loss is defined as Lcl = −log[exp(zM(l)⊤zpos(l)/τ) / Σ N∈ pos,neg exp(zM(l)⊤zN(l)/τ)], where τ is a temperature hyperparameter.
The final joint objective is L(l) = Lalign + ω·Lcl, where ω controls the strength of the contrastive term.
The authors evaluate RPAM on seven widely used reasoning benchmarks
: GSM8K, MATH500, AIME24 (in-distribution), and AIME25, Minerva Math, OlympiadBench, GPQA (out-of-distribution). They use two model pairs: Qwen3-4B-Thinking (Long-CoT) and Qwen3-4B-Instruct (Short-CoT)
and DeepSeek-R1-Distill-Qwen-1.5B (Long-CoT) and Qwen2.5-Math-1.5B (Short-CoT).
For pattern-labeled dataset construction, they randomly sample a total of 128 questions from the GSM8K, MATH500, and AIME24 test sets, and draw k=12 responses per question from each of the two base models.
In the 4B setting, RPAM attains an average accuracy of 75.9 with only 5,976 tokens on average, cutting generation length by 48.33% relative to the Long-CoT base model while incurring merely a 4.37% reduction in accuracy.
In the 1.5B setting, "RPAM achieves an average accuracy of 44.8 with only 2,746 tokens on average, outperforming the two strong efficient-reasoning baselines: the data-dependent merging method AIM (41.9) and the training-based approach Ada-R1 (40.0)."
The ablation study shows that when we progressively remove contrastive learning (-CES) and then both contrastive learning and feature alignment (-CES-FA), we observe a consistent drop in accuracy accompanied by longer reasoning traces.
Replacing pattern-labeled data with random labels +Random
causes substantial degradation in both accuracy and efficiency.
Compared to RL-based methods (HAPO, Arora and Zanette, LC-R1), RPAM attains substantially greater compression, reducing response length by around 63% while maintaining competitive accuracy.
Additionally, the training costs of RPAM are under one hour with the 1.5B model setting, while the RL-based method costs more than 20 hours.
The paper demonstrates that RPAM dynamically selects Long-CoT pattern when needed, achieving a better balance between accuracy and inference efficiency.
Specifically, RPAM has lower ratios on Level-1 and 2 and preserves deeper reasoning on the harder levels
of MATH500 difficulty levels.
The authors acknowledge: "(1) Due to limited computational resources, we restrict our experiments to 1.5B and 4B models. (2) Moreover, our experiments are limited to dense models, and we do not assess performance on Mixture-of-Experts (MoE) models. We also assume the two base models, Long-CoT and Short-CoT, follow the same architecture during merging, and do not consider merging across heterogeneous models in different families."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
-
What it does: Enables the AI to automatically decide between long, detailed reasoning (Long-CoT) and short, concise reasoning (Short-CoT) on a per-query basis, rather than using a fixed reasoning style for all inputs.
-
Specific capability: For a simple arithmetic question like
2+3=?
, the AI will produce a direct answer in under 50 tokens. For a complex math olympiad problem, it will generate a multi-step, self-verifying reasoning chain exceeding 5,000 tokens—without any user prompting or external control. -
What it does: Combines a Long-CoT reasoning model and a Short-CoT instruction model into a single merged model by learning layer-specific merging coefficients, rather than using a naive global weight average.
-
Specific capability: The merged model preserves the deep reasoning strengths of the Long-CoT model on hard problems while inheriting the token efficiency of the Short-CoT model on easy ones. This is achieved by aligning intermediate representations at each layer with the
positive
model (the one that performs better for that specific query) and pushing away from thenegative
model via a contrastive loss. -
What it does: Automatically builds a small calibration dataset (e.g., 128 questions) that labels each query with the optimal reasoning pattern (Long-CoT or Short-CoT) based on empirical accuracy comparison.
-
Specific capability: The AI can be calibrated on a tiny dataset without needing large-scale human annotations or expensive reinforcement learning. For each query, it samples multiple responses from both base models, computes expected accuracy, and assigns the label to the model with higher accuracy (or shorter response if tied).
-
What it does: Adds a binary contrastive loss that explicitly separates the merged model's representations from the non-selected model, preventing the merged model from accidentally resembling the wrong reasoning style.
-
Specific capability: This ensures that when the AI decides to use Short-CoT, it does not accidentally slip into long, overthinking behavior, and vice versa. The contrastive term is controlled by a temperature parameter τ and a strength weight ω, both tunable for different model scales.
-
What it does: Reduces average token generation by 48–64% compared to the Long-CoT baseline, while maintaining or even improving accuracy on several benchmarks.
-
Specific capability: On the MATH500 benchmark with Qwen3-4B models, RPAM achieves 96.4% accuracy (vs. 95.2% for Long-CoT) while reducing response length from 1,521 tokens to 3,191 tokens (a 48% reduction). On GSM8K, it cuts tokens from 6,125 to 1,086 (an 82% reduction) while improving accuracy from 96.0% to 95.7%.
-
What it does: Avoids end-to-end retraining or reinforcement learning, requiring only layer-wise coefficient optimization on a small dataset, typically under one hour on a single GPU.
-
Specific capability: The system can be deployed and adapted to new model pairs (e.g., 1.5B or 4B scales) in under an hour, whereas RL-based methods (e.g., HAPO, LC-R1) require over 20 hours of training. This makes it practical for rapid iteration and deployment in resource-constrained environments.
-
What it does: Maintains adaptive reasoning behavior on benchmarks not seen during calibration, such as AIME25, OlympiadBench, and GPQA.
-
Specific capability: Even when the input format or subject differs from the calibration data (e.g., a biology multiple-choice question from GPQA), the AI correctly selects Short-CoT for simple questions and Long-CoT for complex ones, as demonstrated in the case study (Figure 7).
-
What it does: Dynamically adjusts the
thinking ratio
(frequency of reflective keywords likewait,
double-check
) based on problem difficulty. -
Specific capability: On MATH500, the AI uses thinking in only 92% of Level-1 problems (vs. 100% for Long-CoT) but increases to 98% for Level-5 problems, matching Long-CoT's depth on hard questions while saving tokens on easy ones.
The improved AI system can:
-
Automatically adapt its reasoning depth to each query, eliminating overthinking on simple tasks and underthinking on complex ones.
-
Reduce inference cost by 48–64% on average without sacrificing accuracy, and in some cases improving it.
-
Be calibrated in under an hour on a single GPU using only 128 labeled examples, without retraining or RL.
-
Generalize to unseen benchmarks and question types, maintaining adaptive behavior across math, biology, and other domains.
-
Provide a tunable trade-off between accuracy and efficiency via initialization coefficients and contrastive loss strength, allowing system designers to prioritize speed or correctness as needed.
Abstract
Recent large reasoning models (LRMs) have made substantial progress in complex reasoning tasks, yet they often generate lengthy reasoning paths for every query, incurring unnecessary computation and latency. Existing speed-up approaches typically rely on retraining the model or designing sophisticated prompting, which are either prohibitively expensive or highly sensitive to the input and prompt formulation. In this work, we study model merging as a lightweight alternative for efficient reasoning: by combining a long chain-of-thought (Long-CoT) reasoning model with a Short-CoT instruction model, we obtain an adaptive reasoner without training from scratch or requiring large-scale additional data. Building on this idea, we propose Reasoning Pattern Alignment Merging (RPAM), a layer-wise model merging framework based on feature alignment to facilitate query-adaptive reasoning. RPAM first constructs a small pattern-labeled calibration set that assigns each query an appropriate reasoning pattern. It then optimizes layer-wise merging coefficients by aligning the merged model's intermediate representations with those of the selected model, while a contrastive objective explicitly pushes them away from the non-selected model. Experiments on seven widely used reasoning benchmarks show that RPAM substantially reduces inference cost while maintaining strong performance. Upon article acceptance, we will provide open-source code to reproduce experiments for RPAM.
Sources
- Training Verifiers to Solve Math Word Problems
- The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks
- Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?
- Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
- AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
- Efficient Reasoning via Chain of Unconscious Thought
- ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
- Optimizing Length Compression in Large Reasoning Models
- Measuring Mathematical Problem Solving With the MATH Dataset
- HAPO: Training Language Models to Reason Concisely via History-Aware Policy Optimization
- From System 1 to System 2: A Survey of Reasoning Large Language Models
- Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning Optimization
- Towards Reasoning Ability of Small Language Models
- Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
- Activation-Informed Merging of Large Language Models
- Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging
- Distilling System 2 into System 1
- Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools
- ARM: Adaptive Reasoning Model
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering