LGMar 29, 2023

Importance Sampling for Stochastic Gradient Descent in Deep Neural Networks

arXiv:2303.16529v12.03 citationsh-index: 3Has Code

Originality Synthesis-oriented

AI Analysis

This work addresses the challenge of optimizing training efficiency in deep learning, but it appears incremental as it reviews existing importance sampling methods without introducing a new paradigm.

The paper tackles the problem of inefficient uniform sampling in stochastic gradient descent for deep neural networks by proposing a metric to evaluate sampling schemes and studying their interaction with optimizers, but does not report concrete performance numbers.

Stochastic gradient descent samples uniformly the training set to build an unbiased gradient estimate with a limited number of samples. However, at a given step of the training process, some data are more helpful than others to continue learning. Importance sampling for training deep neural networks has been widely studied to propose sampling schemes yielding better performance than the uniform sampling scheme. After recalling the theory of importance sampling for deep learning, this paper reviews the challenges inherent to this research area. In particular, we propose a metric allowing the assessment of the quality of a given sampling scheme; and we study the interplay between the sampling scheme and the optimizer used.

View on arXiv PDF Code

Similar