A Compositional Kernel Model for Feature Learning
cs.LG, math.OC
Submitted: 2025-09-17
Updated: 2026-09-01
Comments: Fix Typos
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study a compositional variant of kernel ridge regression in which the predictor is applied to a coordinate-wise reweighting of the inputs.
Terminology
Abstract
We study a compositional variant of kernel ridge regression in which the predictor is applied to a coordinate-wise reweighting of the inputs. Formulated as a variational problem, this model provides a tractable setting for studying feature learning in compositional architectures. From the perspective of variable selection, we show how relevant variables are recovered while noise variables are eliminated. We prove that both global minimizers and stationary points discard noise coordinates when the noise variables are Gaussian distributed. A central finding is that 1-type kernels, such as the Laplace kernel, succeed in recovering features contributing to nonlinear effects at stationary points, whereas Gaussian kernels recover only linear ones.
Sources
- On the Self-Penalization Phenomenon in Feature Selection
- A Theory of Feature Learning in Kernel Models
- Enhanced Feature Learning via Regularisation: Integrating Neural Networks and Kernel Methods
- A Variational Analysis of Kernel Learning with Learnable Linear Transformations
- Gradient flow in the kernel learning problem
- Iteratively reweighted kernel machines efficiently learn sparse functions
- xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
- Position: A Theory of Deep Learning Must Include Compositional Sparsity
- Taming Nonconvexity in Kernel Feature Selection -- Favorable Properties of the Laplace Kernel
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks