Towards Efficient Reasoning in LLM-Based Recommender Systems via Model Merging

arXiv:2608.10447 · cs.IR, cs.AI · Submitted 2026-08-11 · Read on arXiv

Linh Dieu Le, Tong Chen, Shazia Sadiq, Hongzhi Yin, Ming Jin, Junliang Yu

The University of Queensland · Griffith University

cs.IR, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Code: https://github.com/linhledieu/REAM

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

The gist: This paper introduces REAM (Reasoning-HEad-Aware Merging), the first model merging framework for reasoning compression in LLM-based recommender systems.

Terminology

Summary

This paper introduces REAM (Reasoning-HEad-Aware Merging), the first model merging framework for reasoning compression in LLM-based recommender systems. The work addresses the challenge that slow-thinking recommender systems, which generate step-by-step reasoning before making predictions, achieve higher accuracy than fast-thinking models but produce unnecessarily verbose reasoning traces that increase inference costs without commensurate accuracy gains.

The paper identifies that slow-thinking recommenders, echoing dual-process cognition, incorporate Chain-of-Thought (CoT) reasoning to accumulate evidence before producing a rating. While this additional deliberation "can improve accuracy, particularly when preference signals are sparse or implicit, but often produces unnecessarily long traces regardless of input complexity, increasing inference latency and computational cost without commensurate benefits."

Existing approaches to reasoning compression face limitations: "Training-based methods, such as distillation and length-penalised optimisation, can shorten reasoning traces but require additional data construction and model adaptation, limiting their practical scalability. By contrast, inference-time methods avoid retraining by imposing token budgets or instructing the model to reason in shorter steps, yet their effectiveness can vary across inputs."

The paper proposes model merging as a training-free alternative that transfers behaviours by combining models in parameter space. The key insight is that "fast- and slow-thinking recommenders share the same preference-prediction objective but differ in the amount of reasoning generated before reaching a prediction. Merging them may therefore transfer the concise generation behaviour of the fast-thinking model while preserving the reasoning required for accurate recommendation."

REAM operates at the level of individual attention heads, assigning each attention head a separate coefficient based jointly on its reasoning importance and sensitivity to parameter change. The method addresses three key questions:

The paper defines two complementary signals for reasoning importance:

Retrieval Criticality: We posit that a head's importance may be reflected in its ability to retrieve relevant user-item evidence from the input context and previously generated trace to support the developing reasoning. This is measured by how frequently each head retrieves segment-relevant evidence throughout the reasoning trace, using a criterion that aligns the generated token with the source position receiving a head's strongest attention.

Decision Faithfulness: At the rate step, we posit that its importance may instead arise from connecting the final prediction to the preceding reasoning, attending to key conclusions established in the trace. This is measured as the share of attention directed to match among the three tracked regions (prompt, analyze, and match), since experiments show that replacing match yields a divergence 3.3–5.1× larger than replacing user or item, indicating that the rating is most sensitive to perturbations of the match segment.

The paper estimates the sensitivity of the slow-thinking model using the diagonal empirical Fisher of theta S and weights this sensitivity by the squared fast-thinking update to obtain a component-level risk measure. This captures how safely its parameters can be modified — whether the loss is flat along some parameter directions, tolerating larger changes with little effect, but sharp along others, where even small perturbations can substantially alter the model's behaviour.

REAM combines retrieval criticality, decision faithfulness, and update sensitivity into a perturbation weight that represents the risk of perturbing unit b along the fast-thinking update: sb = exp(gamma(kappab + faithb))db.

The allocation solves a constrained optimization: we choose merge coefficients that maximise the total applied fast-thinking update while limiting aggregate reasoning-grounded risk, subject to the constraint that the total perturbation stays within a budget. The KKT conditions yield a water-filling solution: alphab = clip(wb/(2musb), 0, alpha¯).

Experiments on Amazon Book, Amazon Music, and Yelp datasets show that REAM reduces reasoning length by up to 24.3% while outperforming competitive model merging baselines in maintaining recommendation accuracy. Specifically:

  • On Amazon Book: MAE reduced by 4.7%, RMSE by 2.3%, and tokens by 24.3% compared to the slow-thinking model

  • On Yelp: MAE reduced by 2.6%, RMSE by 3.7%, and tokens by 17.9%

  • On Amazon Music: MAE reduced by 1.6%, RMSE by 2.2%, and tokens by 21.1%

REAM is the only comparable merging method that maintains accuracy across MAE and RMSE while shortening reasoning on all three datasets. The paper notes that general-purpose merging can shorten reasoning, but does not reliably preserve accuracy — for example, on Book, Task Arithmetic and DARE+TA achieve substantial compression... but both worsen MAE relative to theta S.

Removing Fisher sensitivity has the largest effect on both datasets, increasing MAE to 0.5533 on Music (vs. 0.5348) and 0.7616 on Yelp (vs. 0.7564). The two reasoning signals (retrieval criticality and decision faithfulness) identify complementary and transferable head-level roles rather than redundant rankings, with near-zero correlation (Spearman rho ≤ 0.018) and nearly disjoint top-5% head sets (Jaccard ≤ 1.75%).

Ablating retrieval-critical heads increases NLL/JSD by 11.57×/3.29× on Music, 4.93×/3.96× on Book, and 7.35×/3.19× on Yelp compared to random ablation. Faithfulness-critical heads cause similarly large increases of 9.65×/5.64×, 11.16×/5.20×, and 6.63×/4.19×, respectively.

REAM transfers across model scales, base-model families, and training paradigms. On Qwen2.5-7B, it shortens theta S's trace from 336.02 to 288.46 tokens and outperforms Task Arithmetic in MAE and RMSE. On Llama-3.2-3B, it improves on theta S across all three metrics. On RecOne, REAM outperforms both theta S and Task Arithmetic on all three metrics.

REAM's full pipeline completes in approximately 41 minutes (0.69 GPU-hours) per dataset, which is substantially lower than training the two source models (combined 15.7 GPU-hours). The paper notes this is approximately 4% of this combined training budget and requires no gradient updates to either source model, and involves no changes to decoding.

The paper concludes that reasoning compression in slow-thinking recommender systems depends not only on how much of the fast-thinking update is merged, but also on which components receive it. REAM demonstrates that the two reasoning signals identify causally important yet largely disjoint head sets, supporting their complementary rather than redundant contributions. The remaining accuracy gap at larger model scale highlights an important direction for extending selective merging to larger recommenders.

Improvements for AI systems

Improvements to AI systems based on this paper:

  1. Adaptive reasoning-length control: AI systems can dynamically adjust the verbosity of their internal reasoning traces based on task complexity, producing concise reasoning for simple inputs while preserving detailed deliberation for ambiguous or sparse-signal cases—reducing inference latency and cost without sacrificing accuracy.

  2. Training-free reasoning compression via selective parameter merging: Systems can compress reasoning chains by merging a fast-thinking model's parameters into a slow-thinking model at the attention-head level, using head-specific coefficients derived from reasoning importance and parameter sensitivity—eliminating the need for distillation, fine-tuning, or decoding-time constraints.

  3. Head-level importance attribution for interpretable reasoning: AI systems can identify which attention heads are causally critical for retrieving evidence and for connecting reasoning to final decisions, enabling targeted interventions (e.g., pruning, masking, or boosting specific heads) to control reasoning behavior without retraining.

  4. Risk-aware parameter modification: Systems can estimate the safety of modifying each parameter component using empirical Fisher information, allowing them to apply larger updates along flat loss directions while avoiding sharp directions—reducing the risk of catastrophic behavior change during merging or adaptation.

  5. Complementary reasoning-signal fusion: AI systems can combine multiple orthogonal signals (e.g., retrieval criticality and decision faithfulness) to identify nearly disjoint sets of important components, enabling richer, more transferable reasoning structures than any single signal alone.

  6. Water-filling constrained optimization for resource allocation: Systems can allocate computational or representational resources (e.g., merge coefficients, attention budgets) across components by solving a constrained optimization that maximizes utility (e.g., fast-thinking update applied) while respecting a total perturbation budget—ensuring principled trade-offs between compression and fidelity.

  7. Cross-scale and cross-family reasoning transfer: The selective merging approach works across model sizes (3B–7B), base-model families (Qwen, Llama), and training paradigms (including non-recommender tasks), enabling lightweight compression of reasoning-heavy models without requiring access to original training data or gradients.

  8. Causal validation of reasoning components: Systems can use perturbation-based causal analysis (e.g., NLL/JSD divergence) to validate which components truly matter for reasoning quality, providing a rigorous method for auditing and debugging chain-of-thought behavior in production AI systems.

What the improved AI system can do:

  • Generate shorter, cost-efficient reasoning traces for recommendation, question answering, or decision-making tasks while maintaining or improving prediction accuracy (e.g., up to 24.3% token reduction with better MAE/RMSE).

  • Compress reasoning in a training-free manner within 40 minutes per dataset, using only a few forward passes—suitable for rapid deployment in resource-constrained environments.

  • Provide interpretable, head-level explanations of why certain reasoning steps are kept or removed, enabling human oversight and targeted fixes for reasoning failures.

  • Adapt its reasoning depth on-the-fly to input complexity, avoiding overthinking on easy queries and underthinking on hard ones.

  • Transfer compressed reasoning behavior across model architectures and scales without retraining, making it practical for evolving model families.

Abstract

Large language model-based recommender systems are increasingly adopting slow-thinking models that generate step-by-step reasoning before making predictions, often achieving higher accuracy than fast-thinking models that predict directly. However, their reasoning traces are often unnecessarily verbose, increasing inference costs without commensurate accuracy gains. Existing training-based approaches to reasoning compression often incur substantial adaptation costs, while inference-time methods are brittle and difficult to scale. These limitations motivate model merging as a promising training-free direction for transferring specialised behaviours between models in a shared parameter space. In particular, merging a slow-thinking model with a fast-thinking counterpart provides a natural mechanism for balancing recommendation accuracy and reasoning conciseness. To this end, we propose, to our knowledge, the first model merging framework for reasoning compression in recommender systems. Unlike conventional merging methods that apply uniform merge coefficients across model components, our method performs fine-grained merging at the level of individual attention heads, capturing heterogeneous patterns in recommendation reasoning. Each attention head is assigned a distinct merge coefficient according to its contribution to critical reasoning evidence and its sensitivity to parameter change, enabling selective injection of the concise behaviour of the fast-thinking model into the slow-thinking model and reducing reasoning verbosity without compromising recommendation quality. Experiments on three benchmark datasets show that our method reduces reasoning length by up to 24.3% while outperforming competitive model merging baselines in maintaining recommendation accuracy. The code is available at https://github.com/linhledieu/REAM.

Sources

Related papers