InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models

arXiv:2505.13878 · cs.LG, cs.CL · Submitted 2025-05-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models".

Jane: The gist The current few fusion methods on PA phase, like WRPO,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: The paper’s title, "InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models," really tells you exactly what it does—it's about fusing models implicitly through preference optimization. They list Yanggan Gu, Yuanyi Wang, Zhaoyi Yan, Yiming Zhang, Qi Zhou, Fei Wu, and Hongxia Yang as the authors who developed this framework.

Jane: It’s interesting how they frame it as implicit fusion. This isn't about explicitly merging the model weights in a traditional sense; it’s about creating a fused source model to serve as the reference during preference optimization steps, which is what they focus on.

Lu: The implication here is that we can achieve this integration without having to do complicated vocabulary matching between models, which has been a real hurdle for prior work like WRPO six <ref:2505.13878#pg2>.

Meng: So, if they solve the vocabulary alignment problem by looking at sequence probabilities instead of just token outputs, that suggests a more robust way to combine models with different underlying architectures.

Lalam: It opens up a path where the pivot model can truly benefit from the diverse knowledge each source model brings, not just one narrow perspective.

The paper's summary: Tom: The core idea of this paper is that existing methods on preference alignment often simplify things by only looking at response outputs and throwing away the probability information from the source models. InfiFPO goes a step further by replacing the reference model in DPO with a fused source model that synthesizes multi-source probabilities at the sequence level, which maintains those crucial probability signals.

Jane: They are moving from just using response-level supervision to using sequence probabilities from all source models, including for both preferred and dispreferred responses. That makes the fusion process much more comprehensive.

Lu: The framework is built on a KL-constrained optimization problem where you keep the pivot model within a certain distance, an epsilon ball, of every single source model simultaneously across the dataset D <ref:2505.13878#pg3>.

Meng: So, they’re using this constraint to force the pivot model to stay close in behavior to all sources at once while maximizing preference rewards. That sounds like a very structured way to learn from multiple models at the same time.

Lalam: This means the pivot model learns not only what humans prefer but also it learns from the probabilistic tendencies of every source model in that constrained space, which is a significant improvement over just one reference point.

The paper's improvements: Tom: They introduce three main enhancements to make this approach more stable and effective. First, they have length normalization to reduce bias caused by different token sequence lengths across the models. Second, there's probability clipping to limit the influence of underperforming source models by suppressing noisy gradients.

Jane: And third, they have max-margin fusion, which adaptively prioritizes the source models that provide the most distinctive and informative differences when comparing them to the pivot model. That sounds like a smart way to handle sources that might be less reliable.

Lu: These specific enhancements help improve robustness, especially when you're dealing with heterogeneous models where some might be performing better on certain tasks than others <ref:2505.13878#pg2>.

Meng: From an engineering standpoint, the probability clipping and length normalization are practical because they directly address common noise issues we see in training across different model sizes and vocabularies.

Lalam: Max-margin fusion is particularly interesting because it doesn't treat all sources equally; it actively seeks out the most informative models to guide the pivot model's learning.

Conclusion: Tom: So, to wrap up, InfiFPO enables model fusion during preference alignment by replacing the reference model in DPO with this fused sequence-level distribution over multiple source models, effectively bypassing token-level vocabulary issues while keeping the probability signals intact. They show it performs well on benchmarks like Phi-four improving its average performance from seventy-nine point nine five to eighty-three point three three on eleven tasks.

Jane: The combination of length normalization and probability clipping gave them the strongest performance boost across all those benchmarks, with gains as high as plus five point six in coding tasks and plus three point one in math ones.

Lu: It’s also worth mentioning that their max-margin fusion strategy outperformed both average-based and confidence-based strategies, showing it prioritizes the most useful source models.

Meng: While the results are strong on those specific benchmarks, the paper points out that they only selected five mainstream open-source LLMs for their experiments, so we don't yet know if this scales to even larger or more diverse model sets.

Lalam: They also note that more rigorous theoretical analysis is still needed to fully understand the fusion mechanisms behind InfiFPO, which is something they are looking into next.

Tom: Exactly. So, "InfiFPO: Implicit Model Fusion via Preference Optimization in Large Language Models" gives us a principled framework for integrating diverse LLMs into a single model during preference alignment by leveraging full sequence probabilities instead of just token-level outputs. We'll be talking about other papers next on how AI systems are detecting free-riders in federated learning.

Hong Kong Polytechnic University (PolyU) · PolyU-Daya Bay Technology and Innovation Research Institute

cs.LG, cs.CL

Submitted: 2025-05-20

Updated: 2026-10-08

Journal ref: NeurIPS 2025

Code: https://github.com/InfiXAI/InfiFPO

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 84/100

The gist: The gist The current few fusion methods on PA phase, like WRPO, simplify the process by utilizing only response outputs from source models while discarding their probability information InfiFPO

Key concepts

Implicit Model Fusion
This technique aims to combine knowledge from several different language models into one unified representation without explicitly merging their entire weights. InfiFPO achieves this by fusing the probabilistic outputs of multiple source models at the sequence level, allowing the system to benefit from diverse model strengths.
Direct Preference Optimization (DPO)
DPO is a method used to align language models with human preferences by directly optimizing a reward function. InfiFPO adapts DPO by changing its reference model to a fused source model, ensuring that the optimization process incorporates the combined probabilistic knowledge from several underlying models.
Sequence-Level Probability Synthesis
Instead of using token-level outputs, InfiFPO synthesizes probabilities for an entire sequence by combining information from all source models. This approach bypasses difficult vocabulary alignment problems between different models while keeping the rich probability data necessary for effective preference learning.

Terminology

Summary

The gist The current few fusion methods on PA phase, like WRPO, simplify the process by utilizing only response outputs from source models while discarding their probability information InfiFPO replaces the reference model in Direct Preference Optimization (DPO) with a fused source model that synthesizes multi-source probabilities at the sequence level, circumventing complex vocabulary alignment challenges in previous works and meanwhile maintaining the probability information

How it works

InfiFPO is proposed as a preference optimization method for implicit model fusion by replacing the reference model in DPO with a fused source model that synthesizes multi-source probabilities at the sequence level This approach circumvents complex vocabulary alignment challenges in previous works and maintains probability information The framework is derived from an RLHF-style constrained optimization framework called FuseRLHF, which encourages the pivot model to maximize preference rewards while remaining close—in sequence-level KL divergence—to each source model This allows the pivot model to learn not only from preference data but also from the probabilistic behaviors of multiple source models

Key Enhancements

To improve the robustness and effectiveness of InfiFPO, three key enhancements are introduced

1 Length Normalization, which reduces bias arising from varying token sequence lengths across models

2 Probability Clipping, which limits the influence of underperforming source models by suppressing noisy gradients

3 Max-Margin Fusion, which adaptively prioritizes source models that offer the most distinctive and informative deviations from the pivot model

Optimization Objective

The unconstrained objective derived from relaxing the constraints in Eq. (5) is arg max Mp Ex∼D,y∼Mp(yx) [r(x, y)] − β X N i=1 γ i DSKL [Mp (yx)Ms i (yx)], where γ i ≥ 0 and PN i=1 γ i = 1 The final InfiFPO objective is LInfiFPO(Mp; Ms i N i=1) = −E(x,yw,yl)∼Dp "log σ βlog Mp (ywx) Msclip fu (ywx) − βlog Mp (ylx) Msclip fu (ylx)

Performance and Analysis

Comprehensive experiments on 11 widely-used benchmarks demonstrate that InfiFPO consistently outperforms existing model fusion and preference optimization methods When using Phi-4 as the pivot model, InfiFPO improve its average performance from 79.95 to 83.33 on 11 benchmarks, significantly improving its capabilities in mathematics, coding, and reasoning tasks The method effectively integrates the capabilities of the source models InfiFPO consistently outperforms preference optimization baselines after training on half of the data’s preferred responses yw, SFT improved by 1.62 on average compared to the original model The combination of Length Normalization and Probability Clipping delivers the strongest performance boost, with significant improvements across all benchmarks (+3.1 in Math, +5.6 in Code, +3.3 in All) The Max-margin fusion strategy consistently outperforms both the Average-based and Confidence-based strategies InfiFPO demonstrates recursive computation of Stirling numbers, while SFT simply recalls a wrong value, lacking procedural reasoning

Conclusion

InfiFPO enables model fusion during the preference alignment phase by replacing the reference model in DPO with a fused sequence-level distribution over multiple source models It leverages full-sequence probabilities rather than tokenlevel outputs, avoiding vocabulary alignment issues across heterogeneous models while preserving rich preference signals This framework offers a robust and scalable path to integrating diverse LLMs into a single model It is noted that more rigorous theoretical analysis is needed to better understand InfiFPO’s fusion mechanisms and further strengthen its theoretical foundation The paper also notes that due to computational resource constraints, only five mainstream open-source LLMs were selected as source models for experiments

Limitation

Despite InfiFPO demonstrating significant empirical performance, it still relies on existing preference optimization methods such as DPO More rigorous theoretical analysis is needed to better understand InfiFPO’s fusion mechanisms and further strengthen its theoretical foundation Additionally, due to computational resource constraints, we only selected five mainstream open-source LLMs as source models for our experiments, which cannot represent the SOTA performance of current advanced LLMs The paper also notes that experiments with larger-scale models and datasets remain unexplored

References

[1] Fanqi Wan, Xinting Huang, Deng Cai, Xiaojun Quan, Wei Bi, and Shuming Shi. Knowledge fusion of large language models. In ICLR, 2024

[2] Fanqi Wan, Longguang Zhong, Ziyi Yang, Ruijun Chen, and Xiaojun Quan. Fusechat: Knowledge fusion of chat models, 2025

[3] Yuanyi Wang, Zhaoyi Yan, Yiming Zhang, Qi Zhou, Yanggan Gu, Fei Wu, and Hongxia Yang. Infigfusion: Graph-on-logits distillation via efficient gromov-wasserstein for model fusion, 2025

[4] Zhaoyi Yan, Yiming Zhang, Baoyi He, Yuhao Fu, Qi Zhou, Zhijie Sang, Chunlin Ji, Shengyu Zhang, Fei Wu, and Hongxia Yang. Infifusion: A unified framework for enhanced cross-model reasoning via llm fusion, 2025

[5] Qi Zhou, Yiming Zhang, Yanggan Gu, Yuanyi Wang, Zhijie Sang, Zhaoyi Yan, Zhen Li, Shengyu Zhang, Fei Wu, and Hongxia Yang. Democratizing AI through model fusion: a comprehensive review and future directions. Nexus

[6] Ziyi Yang, Fanqi Wan, Longguang Zhong, Tianyuan Shi, and Xiaojun Quan. Weighted-reward preference optimization for implicit model fusion. In The Thirteenth International Conference on Learning Representations, 2025

[7] Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J. Hewett, Mojan Javaheripi, Piero Kauffmann, James R. Lee, Yin Tat Lee

[8] Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model. In NeurIPS, 2023

[9] Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey

[10] Qwen2.5 technical report

[11] Gemma 3 technical report

[12] Qwen2.5-coder technical report

[13] Qwen2.5-math technical report: Toward mathematical expert model via self-improvement. arXiv preprint arXiv:2409.

Improvements for AI systems

  1. Model fusion during preference alignment can be performed by replacing the reference model in Direct Preference Optimization (DPO) with a fused source model that synthesizes multi-source probabilities at the sequence level, circumventing complex vocabulary alignment challenges in previous works and meanwhile maintaining the probability information. This allows the pivot model to learn not only from preference data but also from the probabilistic behaviors of multiple source models.

  2. The InfiFPO framework can be implemented efficiently by deriving it from an offline relaxation of a sequence-KL constrained RLHF objective, which avoids expensive online sampling and reward model training. This converts the constrained RL problem into a fully offline optimization objective that is computationally efficient for training.

  3. Stability and robustness against source model degradation are enhanced through three key strategies: (1) Length Normalization, which reduces bias arising from varying token sequence lengths across models; (2) Probability Clipping, which limits the influence of underperforming source models by suppressing noisy gradients; and (3) Max-Margin Fusion, which adaptively prioritizes source models that offer the most distinctive and informative deviations from the pivot model.

  4. The improved system can achieve superior performance across diverse tasks by inheriting specialized strengths while avoiding weaknesses; for instance, it can maintain balanced high performance across these diverse task categories by selectively fusing domain-specific models based on their relevance to the input data.

  5. The system demonstrates superior reasoning capabilities in complex mathematical and coding tasks; for example, InfiFPO can construct solutions via complete, logically consistent derivation chains in symbolic manipulation and actively refining the reasoning chain to bridge the gap between instruction semantics and executable code.

Sources

Related papers