Stochastic Optimization of Tree Tensor Networks

arXiv:2609.00870 · math.OC, cs.CV, physics.comp-ph · Submitted 2026-09-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Stochastic Optimization of Tree Tensor Networks".

Tom: Tensor networks, originally developed for quantum many-body physics, are promising models for machine learning.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, to recap, the core of "Stochastic Optimization of Tree Tensor Networks" is the derivation of stochastic Riemannian optimizers for tree tensor networks (TTNs) on both their parameter and quotient manifolds. The authors claim they can handle adaptive and learning-rate-free schemes that are well-suited for minibatch training.

Jane: Essentially, the thesis is to optimize these models by respecting their underlying geometric properties instead of treating them as generic functions in Euclidean space. They focus on optimizing the orthogonal TTN submanifold for numerical stability and to leverage differential geometry results.

Lu: The paper formally introduces the quotient manifold arising from gauge freedom in TTNs and develops key geometric tools, including identifying horizontal space as a Cartesian product of individual horizontal spaces and defining a Riemannian metric on the quotient space through horizontal lifts.

Meng: This geometric framework allows them to relate the quotient gradients to the total space gradient by projecting onto the horizontal space, which is a crucial connection for their optimization logic.

Lalam: The paper also demonstrates that even though the TTN quotient doesn't exhibit a product structure, it shares its geodesic spray with a Cartesian product of Grassmann manifolds, which they use theoretically to back up considerations about the exponential map and Riemannian distances.

Tom: And they test all this out using a hybrid CNN–TTN architecture on Fashion-MNIST, CIFAR10, and Imagenette. The claim is that their proposed optimizers achieve predictive performance comparable to unconstrained optimization while simultaneously enabling numerically stable downstream compression.

Jane: So it matters because the results show that Riemannian optimizers keep the total norm of the TTN model stable during training, which is vital for tasks like model compression where standard ADAM can explode.

Lu: That stability is a direct consequence of working on these specific manifolds, and it confirms that optimizing within this geometric context leads to better adherence to the inherent hierarchical structure of the TTN.

Meng: So from an engineering standpoint, this means we can deploy these complex models more reliably because they won't suffer from numerical instability issues during inference or further optimization steps.

Lalam: I see this as a major step toward making AI systems that are not only accurate but also robust and dependable for deployment in complex environments.

Conclusion: Tom: We’re wrapping up this discussion on "Stochastic Optimization of Tree Tensor Networks" by Marius Willner, Maximilian Scharf, André Uschmajew, Timo Felser, and Marco Trenti. The title itself highlights the core topic: using stochastic optimization techniques specifically for tree tensor networks.

Jane: It’s a really interesting piece because it establishes a theoretical framework for stochastic optimization on orthogonal TTN manifolds and proves that these Riemannian optimizers work effectively for minibatch training.

Lu: The main implication I see is that this work provides a rigorous foundation showing how to apply differential geometry to the optimization of structured models, which isn't just applying existing methods in a new context.

Meng: Practically speaking, the impact is that we gain a way to optimize these large tensor networks in a way that guarantees numerical stability for subsequent model compression steps. That’s something engineers can actually build into their pipelines.

Lalam: I think this advances the field because it shows how deep structural understanding of a model, like its TTN structure, can be leveraged to create optimization algorithms that are inherently more reliable for complex AI systems.

Tom: The paper proves that while these Riemannian optimizers have some overhead in setup, the need for orthogonality in downstream tasks usually justifies it when dealing with tensor networks.

Jane: So we’re looking at a future where optimization methods are intrinsically tied to the model's geometry, which should lead to more efficient and stable AI development overall.

Lu: It opens doors for combining optimization and compression into one process, especially with the suggestions about stochastic optimization schemes with adaptive bond dimensions that could merge these two ideas.

Meng: I'm looking forward to seeing how researchers translate this theoretical work into practical frameworks that can handle the complexity of modern deep learning architectures reliably.

Lalam: Ultimately, this paper suggests a path toward building AI systems where the inherent structure is used not just as a representation, but as an active part of the optimization process itself.

Marius Willner, Maximilian Scharf, André Uschmajew, Timo Felser, Marco Trenti

Institute of Mathematics, University of Augsburg · Tensor AI Solutions GmbH

math.OC, cs.CV, physics.comp-ph

Submitted: 2026-09-01

Updated: 2026-10-02

Code: https://github.com/fastai/imagenette

Importance score: 69/100

The gist: Tensor networks, originally developed for quantum many-body physics, are promising models for machine learning.

Key concepts

Tree Tensor Networks (TTNs)
TTNs are mathematical models used in quantum physics that represent many-body systems. They are structured like trees, making them suitable for machine learning applications because they can efficiently model complex interactions.
Quotient Manifold
This manifold arises from the gauge freedom within TTNs. By optimizing on this quotient space, researchers can refine the optimization process by accounting for these internal symmetries, leading to a more structured and stable learning environment.
Riemannian Optimizers
These are advanced optimization algorithms designed to work on curved spaces (manifolds) instead of flat Euclidean space. They use geometric concepts like Riemannian metrics and geodesics to ensure that the steps taken during training respect the underlying geometry of the TTN structure.

Terminology

Summary

Tensor networks, originally developed for quantum many-body physics, are promising models for machine learning. The proposed methods derive stochastic Riemannian optimizers for tree tensor networks (TTNs) on both their parameter and quotient manifolds, including adaptive and learning-rate-free schemes suitable for minibatch training.

The gist

The paper derives stochastic Riemannian optimizers for tree tensor networks (TTNs) on both their parameter and quotient manifolds, including adaptive and learning-rate-free schemes suitable for minibatch training.

Optimization on TTN Manifolds

The work focuses on optimizing the TTN structure by respecting its underlying geometric properties rather than treating it as a generic parameterized function in Euclidean space. The core tensors of a TTN parameterize the Euclidean space, but working with the orthogonal TTN submanifold is preferred for numerical stability and to exploit differential geometry results. The parameter space of an orthogonal TTN is given by a specific expression involving Stiefel manifolds, and its tangent space is the Cartesian product of the individual tangent spaces of the core tensors.

Quotient Structure and Riemannian Metrics

The paper introduces the quotient manifold, which arises from gauge freedom in TTNs. This quotient structure allows for a more refined optimization setting. The key geometric tools developed include:

  1. The identification of horizontal space as the Cartesian product of individual horizontal spaces, given by Equation (3).

  2. A Riemannian metric on the quotient space, defined through horizontal lifts: the metric is given by Equation (5).

  3. The relation between quotient gradients and the total space gradient via projection onto the horizontal space: the quotient gradient lifted to T is actually equal to the projection of grad f to the horizontal space Equation (12).

Adaptive Optimization Schemes

The paper extends adaptive optimization mechanisms from Euclidean spaces to Riemannian manifolds. For optimization on the TTN manifold T, adaptive schemes iterate according to block-coordinate formulas where the individual components serve as block coordinates in adaptive Riemannian schemes. The work presents several algorithms:

  1. RADAM (Riemannian ADAM) Algorithm 1, which uses momentum updates and a retraction R.

  2. RDOG (Distance over gradients) Algorithm 2, which provides a tuning-free dynamic step-size formula based on geodesic distances, defined by Equation (14).

  3. RMUON and its Riemannian versions (Algorithm 3 and 4), which combine MUON updates with momentum or DOG-like schemes.

Numerical Experiments and Results

The developed algorithms are evaluated on three datasets: Fashion-MNIST, CIFAR10, and Imagenette, using a hybrid CNN–TTN architecture. The results show that the Riemannian optimizers achieve predictive performance comparable to unconstrained optimization while enabling numerically stable downstream compression. Specifically:

(Fashion MNIST)

The lowest training loss was achieved by RMUONDOG, followed closely by RADAM, ADAM & RMUON.

(CIFAR10 & Imagenette)

Riemannian optimizers perform on par with ADAM in terms of test accuracy, but they produce final iterates that conform significantly better to the hierarchical structure of the TTN. Crucially, Riemannian optimizers keep the total norm of the TTN model stable during training, whereas ADAM's norm explodes when optimizing with ADAM, exceeding the numerical range of float32 after just six epochs. This stability is vital for downstream tasks like model compression.

Conclusion and Future Directions

The work establishes a theoretical framework for stochastic optimization on orthogonal TTN manifolds. The main finding is that Riemannian optimizers learn a more numerically stable version for downstream tasks, which is critical because most tensor network algorithms, including the computation of explainability measures and model compression, are based on orthogonalization. Future research directions suggested include developing stochastic optimization schemes with adaptive bond dimensions to combine optimization and compression into one process. The paper concludes that while Riemannian optimizers have an overhead, it is usually justified by the need for orthogonality in downstream tasks.

A Geodesics of the TTN Quotient

The paper proves a deep connection between the quotient manifold TG and a Cartesian product of Grassmannians Z, showing that Any geodesic γˆ: r0, 1s Ñ T G can be mapped to a geodesic γ˜: r0, 1s Ñ Z of the same length. This equivalence confirms that optimization on the quotient space is well-founded. The exponential map on TG is related to the Grassmann exponential map via ExpGr X pδXq “ X V cospΣqV T ´ U sinpΣqV T.

Convergence of RADAGRAD

The convergence analysis for RADAGRAD on the TTN quotient manifold demonstrates that the regret is bounded by a term involving step-size and gradient norms, showing convergence of objective values in expectation for K ≥ 8.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that can be made to existing or future AI systems by implementing these stochastic Riemannian optimization techniques:


  1. Improve numerical stability and generalization of deep Tensor Network (TN) models during training and compression.

  2. Enable more effective model compression for deep TNs by ensuring the final learned iterates conform to the hierarchical structure of the TTN, preventing numerical instabilities that plague unconstrained optimization methods like standard ADAM when orthogonalization is required.

  3. Develop highly efficient, learning-rate-free stochastic optimizers (e.g., RADAM, RMUON) for large-scale TNs that are numerically stable and converge effectively on complex parameter and quotient manifolds, potentially requiring fewer hyperparameters than traditional adaptive methods.

  4. Enhance the interpretability of AI models by ensuring that the learned representations remain orthogonal during training, which is crucial for downstream tasks like feature entropy computation and model compression based on singular value decompositions of the orthogonality center.

  5. Create a hybrid CNN-TTN architecture that allows for the efficient processing of large image datasets (CIFAR10, Imagenette) by leveraging the structured sparsity and efficiency of TTNs as classification heads while using CNNs for robust feature extraction, ensuring that both components are optimized coherently on their respective manifolds.

  6. Implement Distance over Gradients (DOG) schemes to provide a tuning-free dynamic step-size schedule for stochastic optimization on TNs, leading to more conservative and potentially faster convergence by dynamically adjusting the learning rate based on geodesic distances across iterates.

  7. Develop specialized stochastic optimization schemes (like RMUONDOG) that combine momentum with DOG-like step-size control, offering superior convergence properties for complex TN architectures.

This improved AI system can specifically perform the following:

  1. Perform high-fidelity supervised classification and generative modeling using deep Tensor Network models trained on complex, high-dimensional data (e.g., images).

  2. Achieve robust and accurate model compression of deep TNs with significantly higher retention ratios compared to standard ADAM training, ensuring that compressed models retain high test accuracy for downstream tasks like image recognition.

  3. Generate numerically stable and mathematically sound learned representations where the tensors are naturally orthogonal, which is critical for reliable feature extraction (e.g., calculating meaningful feature entropies).

  4. Operate efficiently on large-scale datasets by leveraging a hybrid CNN-TTN architecture optimized end-to-end on the structured manifold of the TNs.

  5. Execute complex optimization tasks where learning rates are not known beforehand, relying instead on geometric properties (geodesic distances) to dynamically determine optimal step sizes, leading to faster and more reliable convergence in high-dimensional parameter spaces.

Sources

Related papers