A Distributional Optimisation Perspective on Combining Models in Deep Learning
cs.LG, stat.ML
Submitted: 2026-09-21
Updated: 2026-09-21
Code: https://github.com/AGI-Edgerunners/LLM-Adapters
License: http://creativecommons.org/licenses/by/4.0/
The gist: Combining predictions from different models can improve performance at machine learning tasks, but the training of the individual models and the rule used to combine them are typically chosen
Terminology
Abstract
Combining predictions from different models can improve performance at machine learning tasks, but the training of the individual models and the rule used to combine them are typically chosen separately, and by ad hoc means. Recent advances in distributional optimisation (i.e. where the optimisation occurs over the set of probability distributions) offer an opportunity for principled joint training, viewing the collection of models as a discrete distribution whose support points are to be optimised, but the potential of these methods is not well-understood. In this paper we (1) cast two standard combination strategies - ensembles and low-rank adapter averaging - as entropy-regularised distributional optimisation, observing that the resulting objective is convex in the ensemble case but not in the adapter-averaging case, so that existing convergence guarantees for mean field Langevin dynamics transfer only to the former; (2) assess existing and novel algorithms for this task, including a functional variant of variational gradient descent; and (3) report an empirical study spanning synthetic classification tasks and fine-tuning of large language models on a commonsense reasoning benchmark.
Sources
- Secret mixtures of experts inside your LLM
- A Computable Measure of Suboptimality for Entropy-Regularised Variational Objectives
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Learning Factored Representations in a Deep Mixture of Experts
- A Review of Sparse Expert Models in Deep Learning
- Large sample analysis of the median heuristic
- The Llama 3 Herd of Models
- Adam: A Method for Stochastic Optimization
- Efficient Ensemble for Fine-tuning Language Models on Multiple Datasets
- Model Merging on Loss Landscape: A Geometry Perspective
- Stochastic Particle Gradient Descent for Infinite Ensembles
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks