Multiscale Training of Convolutional Neural Networks
Shadab Ahamed, Niloufar Zakariaei, Eldad Haber, Moshe Eliasof
cs.LG
Submitted: 2026-02-28
Updated: 2026-08-19
Comments: 25 pages, 10 figures, 8 tables
Journal ref: Transactions on Machine Learning Research (TMLR), March 2026
Code: https://github.com/ahxmeds/multiscale-gradient-estimation
License: http://creativecommons.org/licenses/by/4.0/
The gist: Training convolutional neural networks (CNNs) on high-resolution images is often bottlenecked by the cost of evaluating gradients of the loss on the finest spatial mesh.
Terminology
Abstract
Training convolutional neural networks (CNNs) on high-resolution images is often bottlenecked by the cost of evaluating gradients of the loss on the finest spatial mesh. To address this, we propose Multiscale Gradient Estimation (MGE), a Multilevel Monte Carlo-inspired estimator that expresses the expected gradient on the finest mesh as a telescopic sum of gradients computed on progressively coarser meshes. By assigning larger batches to the cheaper coarse levels, MGE achieves the same variance as single-scale stochastic gradient estimation while reducing the number of fine mesh convolutions by a factor of 4 with each downsampling. We further embed MGE within a Full-Multiscale training algorithm that solves the learning problem on coarse meshes first and "hot-starts" the next finer level, cutting the required fine mesh iterations by an additional order of magnitude. Extensive experiments on image denoising, deblurring, inpainting and super-resolution tasks using UNet, ResNet and ESPCN backbones confirm the practical benefits: Full-Multiscale reduces the computation costs by 4-16x with no significant loss in performance. Together, MGE and Full-Multiscale offer a principled, architecture-agnostic route to accelerate CNN training on high-resolution data without sacrificing accuracy, and they can be combined with other variance-reduction or learning-rate schedules to further enhance scalability.
Sources
- Variance Reduction in SGD by Distributed Importance Sampling
- Bayesian Deep Learning with Multilevel Trace-class Neural Networks
- Stochastic Training of Graph Convolutional Networks with Variance Reduction
- Wavelet Convolutional Neural Networks
- Learning across scales - A multiscale method for Convolution Neural Networks
- Adam: A Method for Stochastic Optimization
- Neural Operator: Graph Kernel Network for Partial Differential Equations
- U-Net: Convolutional Networks for Biomedical Image Segmentation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks