Riemannian Deep Learning: Modules, Networks, and Geometries
cs.LG, cs.AI, math.DG
Submitted: 2026-07-21
Updated: 2026-09-15
Comments: PhD thesis manuscript, University of Trento; defense pending
Code: https://github.com/GitZH-Chen/LieBN
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Euclidean approximations, or require
Terminology
Abstract
Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific manifolds, rely on Euclidean approximations, or require costly and numerically fragile geometric operations. This thesis develops a unified framework for Riemannian deep learning from three complementary perspectives: reusable neural modules, manifold-specific network architectures, and the design of underlying geometries. It generalizes batch normalization from Euclidean spaces and individual manifolds to broad classes of Lie groups and gyrogroups, and extends multinomial logistic regression from Euclidean space to SPD manifolds and then to general Riemannian manifolds. It further develops neural networks for several important geometric representations, including an unconstrained model of hyperbolic space, Busemann-based hyperbolic learning, and full-rank correlation matrices. Finally, it introduces adaptive and computationally efficient Riemannian metrics on SPD manifolds, including learnable Log-Euclidean geometries and fast, stable Cholesky-based geometries. The proposed methods are supported by theoretical analysis and validated through numerical experiments and applications in vision, signal processing, graph learning, and genomics.
Sources
- Layer Normalization
- Stochastic gradient descent on Riemannian manifolds
- ManifoldNorm: Extending normalizations on Riemannian Manifolds
- Riemannian Batch Normalization: A Gyro Approach
- Power Euclidean metrics for covariance matrices with application to diffusion tensor imaging
- Training Deep Networks with Structured Layers by Matrix Backpropagation
- Adam: A Method for Stochastic Optimization
- Geoopt: Riemannian Optimization in PyTorch
- Klein Model for Hyperbolic Neural Networks
- Differentiation of the Cholesky decomposition
- Accelerating 3D Deep Learning with PyTorch3D
- Instance Normalization: The Missing Ingredient for Fast Stylization
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks