HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging
Hyo Seo Kim, Ren Wang
Illinois Institute of Technology
cs.LG, cs.AI
Submitted: 2026-08-11
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging introduces a framework for task vector merging that addresses the limitations of existing methods, which typically require
Terminology
Summary
HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging introduces a framework for task vector merging that addresses the limitations of existing methods, which typically require per-subset scalar tuning and are restricted to linear rescaling. The paper formulates merging across varying task subsets as a combinatorial correction problem and proposes HyperFix, a lightweight hypernetwork that predicts subset-conditioned nonlinear corrections in weight space. Trained once on singleton, pair, and triple subsets from a task bank, HyperFix generalizes to larger subsets without per-subset optimization. The theoretical analysis bounds the residual correction beyond linear merging and motivates learning it from small task updates. Experiments across diverse benchmarks show that HyperFix outperforms existing task vector merging methods while reducing tuning cost.
The method operates as follows: For a subset of tasks S, linear merging produces a merged task vector τS = (1/S) Σ i∈S τi, where τi = θi − θ0 are task vectors. HyperFix augments this with a subset-conditioned correction term, yielding θS = θ0 + τS + ∆S, where ∆S captures structured interaction effects. The subset embedding zS is constructed as the average of task-level embeddings zi, which are rows of a task-level Gram matrix Gij = ⟨τi, τj⟩ computed over all encoder parameters. This embedding is permutation-invariant and encodes interaction statistics. The hypernetwork hϕ, a two-layer MLP with hidden dimension 512 and GELU activation, predicts low-rank factors Ul and Vl for each encoder weight matrix, forming a LoRA-style update ∆Wl = UlVl⊤ with rank r = 4. The training objective is KL-based knowledge distillation, aligning the predictive distributions of the merged model with those of the corresponding single-task models, averaged over tasks in the subset.
The theoretical analysis provides two main theorems. Theorem 1 bounds the residual correction magnitude: ∥∆⋆S∥ ≤ (1/µ)(∥∇LS(θ0)∥ + Hρ), and shows the nonlinear remainder scales quadratically with task update size, ∥RS∥ ≤ (M/2µ)ρ2, under assumptions of local smoothness, local conditioning, and small task updates. Theorem 2 establishes low-order interaction generalization: for any subset size m ≥ 2, E∥hϕ(zS) − ∆⋆S∥ ≤ ε + Lgσz/√m, where the √(1/m) term arises from concentration of the empirical subset embedding around its expectation. This explains why training on subsets up to size three is sufficient for generalization to larger subsets.
Experiments are conducted on eight image classification benchmarks (Cars, DTD, EuroSAT, GTSRB, MNIST, RESISC45, SUN397, SVHN) using CLIP models with ViT-B/32, ViT-B/16, and ViT-L/14 backbones. HyperFix is trained only on subsets with S ≤ 3 and evaluated on subset sizes from 2 to 8, averaging over all possible subsets for each size. Results show that under standard fine-tuning, Mean + HyperFix improves the S = 8 performance from 72.9% (Mean) to 92.0%, surpassing Sum + Scalar (77.0%) and TIES + Scalar (80.9%). TIES + HyperFix further reaches 92.9%. The performance gap between linear merging baselines and HyperFix widens as S increases, indicating the importance of modeling non-additive task interactions. Similar trends hold under tangent-space fine-tuning.
Additional analyses show that KL distillation consistently outperforms cross-entropy training, with gains increasing for larger subsets (e.g., +4.2 percentage points at S = 8). An ablation on maximum training subset size shows that training only on singletons yields 82.0%, including pairs improves to 90.4%, and triples further improve to 94.4%, with marginal gains beyond Smax ≥ 4 (up to 95.9%). The magnitude of corrections remains stable across subset sizes, and perturbing the subset embedding (shuffling or sign-flipping) degrades performance, confirming that corrections meaningfully depend on the subset representation. In terms of computational efficiency, for the full task subset with S = 8, scalar tuning requires 1634.26 seconds, while HyperFix requires only 284.16 seconds, reducing total execution time by 82.6%.
Improvements for AI systems
Improvements to AI Systems:
-
Zero-shot multi-task model merging without per-task tuning: HyperFix enables a single trained hypernetwork to merge any unseen combination of fine-tuned task vectors (e.g., 2 to 8 tasks) into one multi-task model, eliminating the need for expensive per-subset scalar optimization or retraining. The improved system can instantly combine pre-trained specialist models (e.g., image classifiers for cars, satellite images, traffic signs) into a unified model that performs all tasks simultaneously with near-single-task accuracy.
-
Scalable task interaction modeling via low-rank corrections: The system learns to predict nonlinear, subset-conditioned corrections (as LoRA-style low-rank updates) that capture non-additive interactions between tasks—something linear averaging or scalar rescaling cannot. This allows the AI to handle task sets where tasks conflict or interfere (e.g., combining fine-grained object recognition with scene classification) without degrading performance, and the correction magnitude remains stable as the number of tasks grows.
-
Data-efficient generalization from small subsets to large ones: By training only on subsets of size 1–3, the system leverages the theoretical guarantee (Theorem 2) that empirical subset embeddings concentrate, enabling accurate predictions for subsets up to size 8 or beyond. This reduces training data requirements by orders of magnitude—no need to enumerate all possible task combinations during development.
-
Permutation-invariant task embedding for robust merging: The system constructs a subset embedding as the average of task-level Gram-matrix rows (capturing pairwise task-vector similarities). This makes the merging process invariant to the order of tasks and robust to noisy or shuffled task representations, as shown by perturbation experiments. The improved AI can reliably merge tasks even when task identities are partially corrupted or reordered.
-
Distillation-based alignment for better multi-task behavior: Using KL-divergence distillation from single-task models (rather than cross-entropy on merged logits) yields consistent gains, especially for larger subsets (+4.2 points at 8 tasks). The improved system produces a merged model whose output distributions closely match those of the individual experts, improving calibration and reducing overconfidence in ambiguous multi-task scenarios.
-
82.6% faster task merging pipeline: The hypernetwork replaces iterative scalar tuning (which took 1634 seconds for 8 tasks) with a single forward pass (284 seconds). This enables real-time or near-real-time model merging in production environments—e.g., dynamically adding a new task to a deployed multi-task system without downtime or manual hyperparameter search.
-
Theoretical guarantees for safe deployment: The bounds on residual correction (Theorem 1) and low-order interaction generalization (Theorem 2) provide formal assurance that the system will not diverge or produce catastrophic interference when merging many tasks, given small task-update norms. This makes the system suitable for safety-critical applications like medical image analysis or autonomous driving, where predictable multi-task behavior is required.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks