LG MLOct 10, 2022

The good, the bad and the ugly sides of data augmentation: An implicit spectral regularization perspective

Chi-Heng Lin, Chiraag Kaushik, Eva L. Dyer, Vidya Muthukumar

arXiv:2210.05021v316.150 citationsh-index: 19Has Code

Originality Incremental advance

AI Analysis

This provides a theoretical foundation for data augmentation design, addressing a core issue in ML generalization, though it is incremental as it builds on existing regularization perspectives.

The authors tackled the problem of understanding how data augmentation affects generalization in machine learning by developing a theoretical framework for linear models, revealing that it induces implicit spectral regularization through two effects: manipulating eigenvalues and boosting the spectrum, which explains phenomena like discrepancies between over- and under-parameterized regimes.

Data augmentation (DA) is a powerful workhorse for bolstering performance in modern machine learning. Specific augmentations like translations and scaling in computer vision are traditionally believed to improve generalization by generating new (artificial) data from the same distribution. However, this traditional viewpoint does not explain the success of prevalent augmentations in modern machine learning (e.g. randomized masking, cutout, mixup), that greatly alter the training data distribution. In this work, we develop a new theoretical framework to characterize the impact of a general class of DA on underparameterized and overparameterized linear model generalization. Our framework reveals that DA induces implicit spectral regularization through a combination of two distinct effects: a) manipulating the relative proportion of eigenvalues of the data covariance matrix in a training-data-dependent manner, and b) uniformly boosting the entire spectrum of the data covariance matrix through ridge regression. These effects, when applied to popular augmentations, give rise to a wide variety of phenomena, including discrepancies in generalization between over-parameterized and under-parameterized regimes and differences between regression and classification tasks. Our framework highlights the nuanced and sometimes surprising impacts of DA on generalization, and serves as a testbed for novel augmentation design.

View on arXiv PDF Code

Similar