On the Reliability of Generative Augmentation: A Wasserstein-Based Theoretical and Empirical Study
Chathurika S Abeykoon, Mathias Nthiani Muia, Mallory Goldstein
Abstract
Generative data augmentation is widely used to mitigate class imbalance, yet its theoretical effect on downstream generalization remains poorly understood. In this work, we develop a statistical framework for conditional generative augmentation and analyze its impact on classification risk. We formalize augmentation as a distribution-mixing process and show that the resulting risk distortion is controlled by both the augmentation strength and the class-conditional Wasserstein discrepancy between real and generated distributions. We further derive a capacity-dependent generalization bound based on Rademacher complexity, revealing an explicit trade-off between hypothesis complexity, augmentation intensity, and generative fidelity. Empirically, we evaluate the framework on binary and multiclass imbalanced classification tasks using Conditional GAN and Conditional WGAN-GP augmentation. Across datasets, CWGAN-GP consistently achieves lower Wasserstein discrepancies than CGAN, indicating improved distributional fidelity. However, improved fidelity does not necessarily translate into superior classification performance, with classical oversampling methods often remaining competitive. These findings support the central theoretical prediction that augmentation reliability is governed by distributional approximation error rather than predictive performance alone. Overall, this work establishes generative augmentation as a distributional perturbation process whose reliability can be quantified through Wasserstein-based measures and supported by finite-sample generalization guarantees. The proposed framework provides a principled foundation for evaluating synthetic data quality beyond classification accuracy alone.
Create a lesson
Related papers
Copula Transformations for Data-Consistent Inversion
Troy Butler, Tianyi Jiang, João Silva et al.
Full-Model Optimality for Tunable Linear Generative Priors in Compressed Sensing
Zhaoming Li, Paul Hand
Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency
Jia-Nan Wang, Zixun Huang, Kairui Li et al.
A computational approach to maximum likelihood thresholds for colored Gaussian graphical models
Roser Homs, Olga Kuznetsova, Bernadette J. Stolz
From topology learning to graph generation: A unifying perspective
Xiaowen Dong, Hoi-To Wai, Siheng Chen et al.
Schrödinger Bridges on Lie Group Manifolds for Probabilistic Intrinsic Generation
Shizhe Zhang, Mingyang Zhao, Lei Ma