Optimal Pruning for Neural Architectures using Fisher Information Distances
cs.AI, cs.IT, math.DG, math.IT
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 21 pages, 4 figures, 4 tables
Code: https://github.com/edhirst/Fdist
License: http://creativecommons.org/licenses/by/4.0/
The gist: A new scheme for parameter pruning is introduced, derived from the differential-geometric distance in model space.
Terminology
Abstract
A new scheme for parameter pruning is introduced, derived from the differential-geometric distance in model space. Pruning a parameter sets its value to zero, representing a displacement of the model to the hypersurface on which that parameter vanishes. The minimal distance from the unpruned model to this hypersurface is naturally computed via the geodesic distance in the model space as determined by the Fisher information metric. This distance determines the true change in the model, and its performance, under pruning. By analysing progressively more faithful approximations of this geodesic distance a natural hierarchy of optimality for pruning methods is determined. This starts with the traditional magnitude pruning, then develops into new more sophisticated and effective pruning schemes. The method is demonstrated for both fully-connected networks and vision transformers, on MNIST and CIFAR-10, over the complete 0 - 100% pruning range and across five random seeds. It outperforms pruning by parameter magnitude and by the local Fisher information alone in every architecture and dataset combination considered, on both accuracy and the Matthews correlation coefficient. Additionally, analysis of different levels of geodesic approximation produces intermediate pruning schemes that are computationally efficient and maintain near-optimal performance. This geometric picture supplies not only a state-of-the-art pruning methodology for AI models, but also a verified and mathematically-motivated justification for pruning schemes.
Sources
- Learning both Weights and Connections for Efficient Neural Networks
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- What is the State of Neural Network Pruning?
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- Faster gaze prediction with dense networks and Fisher pruning
- Pruning Convolutional Neural Networks for Resource Efficient Inference
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Escaping the Big Data Paradigm with Compact Transformers
- Attention Is All You Need
- Limitations of the Empirical Fisher Approximation for Natural Gradient Descent
- Overcoming catastrophic forgetting in neural networks
- Optimizing Neural Networks with Kronecker-factored Approximate Curvature
- On the Dynamics of Inference and Learning
- The Inverse of Exact Renormalization Group Flows as Statistical Inference
- Bayesian Renormalization
- NCoder -- A Quantum Field Theory approach to encoding data
- Bayesian RG Flow in Neural Network Field Theories
- Grokking vs. Learning: Same Features, Different Encodings
- SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection