Stability and Generalization of Straight-Through Estimators for Training Two-Layer Quantized Neural Networks
cs.LG, stat.ML
Submitted: 2026-09-06
Updated: 2026-09-06
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study the identity straight-through estimator (STE) for training a two-layer binary-activation network with hinge loss from the perspective of Statistical Learning Theory (SLT).
Terminology
Abstract
We study the identity straight-through estimator (STE) for training a two-layer binary-activation network with hinge loss from the perspective of Statistical Learning Theory (SLT). Our central question is whether algorithmic stability can explain the statistical generalization of the estimator produced by the discontinuous STE training rule. In the saturated-output regime, the zero-initialized samplewise STE recursion is exactly the stochastic subgradient descent on the convex latent loss (-yu x)+. This representation makes a stability analysis possible. We derive an exact distance identity for two coupled updates and prove approximate non-expansiveness of the common-example map, with a quadratic defect only when the two latent margins straddle zero. We then obtain explicit 2 on-average model-stability and generalization bounds, transferring stability isometrically from the latent vector to the full first-layer matrix. Combining stability with a standard optimization bound yields an explicit excess induced-risk guarantee and the rate O(n-1/2) when T=n squared. Under margin separability, a complementary argument gives the optimal-order O(R 2/(γ 2n)) expected excess misclassification error for a randomized one-pass STE iterate and a corresponding majority-vote bound.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks