The hidden advantage of mask resampling: a theory of masked autoencoders
Jorge Medina Moreira, Lorenzo Bardone, Lenka Zdeborová
Abstract
Why can masked prediction learn useful representations that unmasked reconstruction misses? We study this question in a high-dimensional model of a masked autoencoder (MAE) trained on data with shared latent structure and heterogeneous noise. We prove that masked linear reconstruction can recover the latent feature at linear sample complexity in regimes where unmasked linear reconstruction, equivalent to PCA, fails. The analysis also quantifies the statistical advantage of mask resampling, an established ingredient of masked pretraining. By introducing a fixed collection of K masks per sample, we characterize its effect on feature recovery and downstream performance, identifying regimes where greater mask diversity lowers sample complexity. Guided by this prediction, we find that random cropping and flipping in standard image-training pipelines can obscure the advantage of mask resampling by renewing the prediction task even when the patch mask is fixed. Removing these transformations reveals a downstream advantage for dynamic over static masking in CNN autoencoders and vision transformers. A complementary BERT pilot finds benefits from greater mask diversity on downstream language tasks. Our results separate the benefit of the masked prediction objective from that of mask diversity, and show how a tractable theory can guide experiments that uncover advantages hidden by standard training practices.
Create a lesson
Related papers
Wasserstein Gradient Flows and Forward-Only Diffusion Are Not Enough for Multimodal Sampling
Daniel McBride, Pratik Khandagale, Cristina Garcia-Cardona et al.
Error-Corrected Inference-Time Scaling for Imperfect Diffusion Models
Zuokai Wen, Louis Grenioux, Weinan E et al.
Zero Flux: Flow-Based Comparison of High-Dimensional Discrete Distributions
Leyang Wang, Yakun Wang, Song Liu et al.
Optimal Transport Meets Reinforcement Learning: A Survey
Yujie Zhu, Charles A. Hepburn, Matthew Thorpe et al.
Posterior sampling by source-space MCMC via prior-based few-step transport maps
Hoang Phuc Hau Luu, Marcelo Hartmann, Zhongjian Wang
Tolerance-Based Fairness Auditing: Violation Certification and Sensitivity Screening
Jie Tang, Chuanlong Xie, Lixing Zhu