Convergence, design and training of continuous-time dropout as a random batch method
cs.LG, math.OC
Submitted: 2025-10-15
Updated: 2026-09-12
Comments: 37 pages, 11 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study continuous-time dropout in controlled differential equations.
Terminology
Abstract
We study continuous-time dropout in controlled differential equations. We introduce a random-batch approximation of additive vector fields. On each time interval of length h, a random subset of components is activated and rescaled by its inclusion probabilities, yielding an unbiased approximation of the full field. We prove a uniform-in-time mean-square trajectory error of order O(h). At the distribution level, we derive Wasserstein and pointwise density estimates, together with global L 1 bounds under moment assumptions. For supervised training, we prove uniform O(sqrt h) root-mean-square fluctuations of the randomized objective and an O(h) weak error for its expectation, leading to consistency of optimal values and near-minimizers and, under quadratic growth, convergence in control space. We further characterize how the sampling law affects the error constant through inclusion probabilities and field variance, compare fixed-size sampling, fixed partitions, and Bernoulli dropout, and derive a work--accuracy model for selecting the switching scale. Numerical experiments with neural ODEs illustrate trajectory and transport errors, sampling effects, and objective consistency at a fixed control.
Sources
- Comprehensive Review of Neural Differential Equations for Time Series Analysis
- Stochastic Training of Residual Networks: a Differential Equation Viewpoint
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks