Certified Inference and Training for Deep Equilibrium Networks: A Continuation Framework with Polynomial Complexity Guarantees
stat.ML, cs.LG
Submitted: 2026-09-15
Updated: 2026-09-20
Comments: Python scripts and Lean formalization are included as ancillary files
License: http://creativecommons.org/licenses/by/4.0/
The gist: We develop a certified continuation framework for equilibrium computation and for training deep equilibrium networks (DEQs), with training formulated as interpolation to accuracy 2-b.
Terminology
Abstract
We develop a certified continuation framework for equilibrium computation and for training deep equilibrium networks (DEQs), with training formulated as interpolation to accuracy 2-b. For inference, compact input homotopy selects a unique branch from a supplied start root, and a rounded Newton tracker follows it under certified boundary, conditioning, derivative, and tube-radius bounds. For training, we augment local-plus-low-rank recurrence with programmable dormant bilinear rank-one channels. Loaded Tikhonov solves diagnose a failed interpolation pass without spectral decomposition; an output-preserving repair aligned with the pass residual supplies the required direction. Training requires certified gate realization and column stability on each pass region, well-posed inference, and finite-update error budgets. With polynomial geometric, encoding, precision, and complete backend budgets, both certified inference and training have bit cost O(poly(L+b)), where L is the encoded instance length. The trainer uses O(b+) passes and reserve channels from an initial residual bounded by 2. These guarantees concern a certified promise class. Lean 4 verifies the quantitative core and concrete inference backend; numerical comparisons illustrate the loaded mechanism.
Sources
- Robust certified numerical homotopy tracking
- Multiscale Deep Equilibrium Models
- Robust Implicit Networks via Non-Euclidean Contractions
- JFB: Jacobian-Free Backpropagation for Implicit Networks
- SHINE: SHaring the INverse Estimate from the forward pass for bi-level optimization and implicit models
- On the Theory of Implicit Deep Learning: Global Convergence with Implicit Layers
- A global convergence theory for deep ReLU implicit networks via over-parameterization
- Net2Net: Accelerating Learning via Knowledge Transfer
- Lipschitz Multiscale Deep Equilibrium Models: A Theoretically Guaranteed and Accelerated Approach
- Response Renormalization for Critical Deep Equilibrium Models
- Adding One Neuron Can Eliminate All Bad Local Minima
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey