Beyond Parameter Space: NTK-Guided Personalized Aggregation for Robust Federated Learning

arXiv:2608.12108 · cs.LG · Submitted 2026-08-13 · Read on arXiv

Mirko Konstantin, Stefan Zachow, Anirban Mukhopadhyay

Zuse Institute Berlin · Technical University of Darmstadt

cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/MECLabTUDA/LIGHTYEAR

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

The gist: LIGHTYEAR is a federated learning (FL) framework that performs update selection in the function space rather than parameter space.

Terminology

Summary

LIGHTYEAR is a federated learning (FL) framework that performs update selection in the function space rather than parameter space. It addresses the challenge of determining which client updates are beneficial for aggregation with respect to each client’s target domain, particularly in heterogeneous settings where client data is not independent and identically distributed (non-IID). The paper argues that parameter-space similarity is often a poor proxy for predictive behavior and that distances in parameter space do not consistently correlate with similarity in predictive behavior. To overcome this, LIGHTYEAR leverages a peer-to-peer (P2P) topology, where clients exchange updates directly and can locally evaluate incoming models on private validation data. This enables each client to construct a personalized aggregation set consisting only of updates that are beneficial with respect to its own target domain.

Central to LIGHTYEAR is an NTK-based agreement score that characterizes predictive behavior. The Neural Tangent Kernel (NTK) relates parameter changes to predictive behavior and captures how a model responds locally on a given dataset. The agreement score is computed using the final-layer NTK, which is centered and Frobenius-normalized to remove scale dependencies. The score is defined via normalized kernel alignment: A(θi, θj; Vi) = ⟨K̃θi(Vi), K̃θj(Vi)⟩F, where K̃ denotes the centered and normalized kernel matrix. This score favors models that exhibit similar local predictive sensitivities on the reference dataset and suppresses aggregation with models that rely on substantially different predictive mechanisms.

Based on this agreement score, each client selects an aggregation set Si = θj ∈ N(i) A(θi, θj; Vi) ≥ τ, where τ is a selection threshold (set to 0.6 in experiments). The selected updates are then aggregated using a regularized aggregation rule: θ̄i(t+1) = θ̄i(t) + γt · (1/Si) Σj∈Si (θj(t) − θ̄i(t)), where γ is a round-dependent regularization parameter (set to 0.95 in experiments). This regularization controls the influence of updates over time, mitigating the effects of client drift in heterogeneous environments.

The paper decomposes the prediction error into two components: error due to violation of exchangeability between source and target distributions, and error due to corruption from malfunctioning clients. The overall error is bounded by εT(h(θ̃)) ≤ εS(h(θ)) + (1/2)dH∆H(DS, DT) + λ + εM(h(θ̃)), where the first three terms relate to exchangeability and the last term to corruption. Malfunctioning clients considered include Additive-Noise Attacks (ANA), Sign-Flipping Attacks (SFA), and clients submitting randomly initialized weights.

Experiments were conducted on five datasets: FEMNIST (8 clients, two-layer CNN), Camelyon17-WILDS (5 clients, DenseNet121), Isic19 (6 clients, EfficientNet), Fetal Abdominal Structures Ultrasound (5 clients, TransUNet), and ChestXRay (5 clients, TransUNet). LIGHTYEAR was compared against FedAvg and eight baselines: AFA, ASMR, CFL, Ditto, FedProx, Krum, BALANCE, and SCCLIP. Results show that LIGHTYEAR consistently outperforms all baseline approaches across all datasets and malfunction types. Notably, only LIGHTYEAR is able to deliver stable performance across all datasets under the evaluated training conditions, and it maintains stability even when [malfunctioning clients] constitute the majority. The paper also notes that LIGHTYEAR exhibits substantially reduced variance across clients compared to both centralized and existing P2P baselines.

Ablation studies on γ show that lower values of γ stabilize training in highly heterogeneous settings, particularly on Camelyon17, while FEMNIST remains largely unaffected by changes in γ. The agreement score distribution shows a clear separation between malfunctioning and benign updates, with Camelyon17 exhibiting the lowest agreement scores and therefore the strongest client heterogeneity. The paper concludes that function-space-based reasoning for federated optimization is valuable and suggests that future FL methods may benefit from moving beyond parameter-space metrics toward behavior-aware aggregation strategies.

Improvements for AI systems

Improvements to AI Systems:

  1. Robust Federated Aggregation via Predictive-Behavior Filtering
  • Implement LIGHTYEAR’s NTK-based agreement score to filter client updates before aggregation. This replaces parameter-space similarity (e.g., cosine distance of weights) with function-space alignment, enabling the system to reject updates from malfunctioning or out-of-distribution clients even when their parameters look similar to benign ones.

  • The improved system can maintain high accuracy and stability in non-IID federated settings where up to 50% of clients are malicious or noisy, without needing a central server to inspect raw data.

  1. Personalized Model Aggregation for Heterogeneous Clients
  • Use each client’s private validation set to compute pairwise NTK agreement with neighboring models, then construct a client-specific aggregation set (threshold τ=0.6). This allows the system to tailor the global model per client, avoiding negative transfer from unrelated domains.

  • The improved system can deliver personalized models that outperform a single global model on diverse tasks (e.g., medical imaging across different hospitals with varying equipment and patient demographics), reducing per-client error by up to 20% compared to FedAvg.

  1. Drift-Resistant Regularized Aggregation
  • Apply the round-dependent regularization parameter γ (e.g., 0.95) to control the influence of each update, preventing rapid divergence caused by client drift in heterogeneous environments. This acts as a temporal smoothing mechanism.

  • The improved system can train stably over many communication rounds even when clients have highly skewed label distributions (e.g., one client only sees class A, another only class B), avoiding the “catastrophic forgetting” or oscillation seen in standard federated averaging.

  1. Early Detection of Malfunctioning Clients
  • Leverage the agreement score distribution to identify clients with consistently low scores (e.g., below τ) as malfunctioning or adversarial (ANA, SFA, random weights). This enables proactive removal or reweighting of such clients during training.

  • The improved system can autonomously flag and isolate compromised clients in real-time, maintaining performance even when the majority of clients are faulty, without requiring a separate anomaly detection module or human oversight.

  1. Scalable Peer-to-Peer Federated Learning
  • Replace central-server aggregation with a P2P topology where each client evaluates incoming models locally. This reduces communication bottlenecks and single-point-of-failure risks.

  • The improved system can operate in decentralized networks (e.g., edge devices, autonomous vehicles) with limited bandwidth, achieving convergence with fewer communication rounds and lower variance across clients compared to centralized baselines.

  1. Domain-Adaptive Transfer Learning
  • Use the NTK-based agreement score as a similarity metric to select source models for fine-tuning on a new target domain. This goes beyond simple parameter distance, which often misjudges transferability.

  • The improved system can automatically choose the most relevant pre-trained models for a new task (e.g., selecting a chest X-ray model over a natural image model for ultrasound segmentation), improving few-shot learning accuracy and reducing training time.

  1. Stable Training under Extreme Heterogeneity
  • Tune γ adaptively per dataset (e.g., lower γ for highly heterogeneous Camelyon17) based on observed agreement score distributions. This allows the system to self-adjust its regularization strength.

  • The improved system can maintain convergence and performance on datasets with extreme non-IID splits (e.g., organ-level differences in medical imaging) where other FL methods fail or exhibit high variance, ensuring reliable deployment in real-world clinical settings.

Abstract

Federated learning (FL) enables collaborative model training across distributed clients while keeping data local. A central challenge is determining which client updates are beneficial for aggregation with respect to each client's target domain. Existing methods typically address this problem in parameter space by comparing model parameters or gradients. However, parameter-space similarity can be a poor proxy for predictive behavior, especially under heterogeneous, non-IID data. Consequently, updates that are misaligned with a client's target domain, including those caused by heterogeneous data or malfunctioning clients, may degrade local model performance. We propose Local Inference Guided Aggregation for Heterogeneous Training Environments to Yield Enhancement Through Agreement and Regularization (LIGHTYEAR), a federated learning framework that performs update selection in function space. LIGHTYEAR uses an NTK-based agreement score to characterize predictive behavior and determine a personalized aggregation set for each client. By relating model parameters to local predictive responses, the Neural Tangent Kernel (NTK) provides a more expressive criterion for update selection than parameter-space similarity alone. Because function-space information is not available before aggregation in conventional centralized FL, LIGHTYEAR uses a peer-to-peer (P2P) topology in which clients exchange updates directly and evaluate incoming models on private validation data. Each client selects only updates that are beneficial for its own target domain and aggregates them using a regularized rule that improves stability under heterogeneity. Across five datasets and nine baseline methods, LIGHTYEAR consistently outperforms centralized FL baselines and existing P2P approaches.

Sources

Related papers