Emergence of Fibrations, Compression, and Symmetry Breaking in Artificial Neural Networks
cs.LG
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: 11 Figures
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Artificial neural networks are often regarded as powerful yet opaque black boxes.
Terminology
Abstract
Artificial neural networks are often regarded as powerful yet opaque black boxes. Here, we demonstrate that learning in deep neural networks generates local symmetries known in graph theory as fibrations and coverings. We prove that covering symmetries are stable attractors of stochastic gradient descent. Consistent with this theory, we report the emergence of covering symmetries across major network architectures, including multilayer, convolutional, recurrent, and transformer networks. Exploiting these symmetries enables drastic model compression - reducing networks to 17% of their original size without sacrificing performance. Furthermore, controlled breaking of covering symmetry overcomes the loss of plasticity, achieving state-of-the-art performance in continual learning. The theoretical results provide a new foundation for AI systems based on symmetries that convert black boxes into interpretable colored graphs and enable more efficient inference and lifelong learning.
Sources
- Scaling Laws for Neural Language Models
- Training Compute-Optimal Large Language Models
- Disentangling the Causes of Plasticity Loss in Neural Networks
- Symmetries of Living Systems: Symmetry Fibrations and Synchronization in Biological Networks
- Fibration symmetry-breaking supports functional transitions in a brain network engaged in language
- Semi-Supervised Classification with Graph Convolutional Networks
- Shrink-Perturb Improves Architecture Mixing during Population Based Training for Neural Architecture Search
- The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction
- CompactifAI: Extreme Compression of Large Language Models using Quantum-Inspired Tensor Networks
- Proximal Policy Optimization Algorithms
- Speeding up Convolutional Neural Networks with Low Rank Expansions
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks