Hilbert Operator for Progressive Encoding (HOPE): A Mathematical Framework for Deconstructing Learned Representations in Deep Networks
Hossein Mobahi, Peter L. Bartlett
cs.LG, cs.AI, stat.ML
Submitted: 2026-07-23
License: http://creativecommons.org/licenses/by/4.0/
The gist: Deep neural networks encode complex representations, but deconstructing this internal knowledge remains a challenge.
Terminology
Abstract
Deep neural networks encode complex representations, but deconstructing this internal knowledge remains a challenge. Given the link between learning and compression, network compression offers a promising lens to analyze this knowledge. However, standard compression heuristics often suffer from scale symmetries and architectural biases. To resolve these, we introduce Hilbert Operator for Progressive Encoding (HOPE), a mathematical framework to gradually deconstruct the representations in trained network weights. HOPE shifts network compression from the discrete domain into a Hilbert space of continuous functions. By modeling individual neurons as rank-1 Hilbert-Schmidt operators, HOPE unifies pruning and neuron merging as low-rank subspace projection. Extending this formulation, HOPE introduces macro block eviction to encompass multi-layer structures like entire residual pathways under the same unified metric. This unified approach enables unbiased architectural decisions across layers with different types and sizes. HOPE is a data-free and hyperparameter-free framework. We present proof-of-concept experiments in model compression and fine-tuning to highlight the practical potential of our theory.
Sources
- Understanding symmetries in deep networks
- Not All Language Model Features Are One-Dimensionally Linear
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- What Do Compressed Deep Neural Networks Forget?
- Emergence of cyclic flux eruptions in kinetic simulations of magnetized spherical accretion onto a Schwarzschild black hole
- The Platonic Representation Hypothesis
- Equivariant Masked Position Prediction for Efficient Molecular Representation
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- Opening the Black Box of Deep Neural Networks via Information
- Pruning neural networks without any data by iteratively conserving synaptic flow
- Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
- TACO: Temporal Consensus Optimization for Continual Neural Mapping
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks