Convergence, design and training of continuous-time dropout as a random batch method

arXiv:2510.13134 · cs.LG, math.OC · Submitted 2025-10-15 · Read on arXiv

cs.LG, math.OC

Submitted: 2025-10-15

Updated: 2026-09-12

Comments: 37 pages, 11 figures

License: http://creativecommons.org/licenses/by/4.0/

The gist: We study continuous-time dropout in controlled differential equations.

Terminology

Abstract

We study continuous-time dropout in controlled differential equations. We introduce a random-batch approximation of additive vector fields. On each time interval of length h, a random subset of components is activated and rescaled by its inclusion probabilities, yielding an unbiased approximation of the full field. We prove a uniform-in-time mean-square trajectory error of order O(h). At the distribution level, we derive Wasserstein and pointwise density estimates, together with global L 1 bounds under moment assumptions. For supervised training, we prove uniform O(sqrt h) root-mean-square fluctuations of the randomized objective and an O(h) weak error for its expectation, leading to consistency of optimal values and near-minimizers and, under quadratic growth, convergence in control space. We further characterize how the sampling law affects the error constant through inclusion probabilities and field variance, compare fixed-size sampling, fixed partitions, and Bernoulli dropout, and derive a work--accuracy model for selecting the switching scale. Numerical experiments with neural ODEs illustrate trajectory and transport errors, sampling effects, and objective consistency at a fixed control.

Sources

Related papers