The hidden advantage of mask resampling: a theory of masked autoencoders
stat.ML, cond-mat.dis-nn, cs.LG
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/SPOC-group/mask-resampling
Terminology
Sources
- BEiT: BERT Pre-Training of Image Transformers
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- A solvable high-dimensional model where nonlinear autoencoders learn structure invisible to PCA while test loss misaligns with generalization
- Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing
- Drawing Multiple Augmentation Samples Per Image During Training Efficiently Decreases Test Error
- Decoupled Weight Decay Regularization
- Large Batch Training of Convolutional Networks
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey