LG CV MLMar 25, 2020

VaB-AL: Incorporating Class Imbalance and Difficulty with Variational Bayes for Active Learning

Jongwon Choi, Kwang Moo Yi, Jihoon Kim, Jinho Choo, Byoungjip Kim, Jin-Yeop Chang, Youngjune Gwon, Hyung Jin Chang

arXiv:2003.11249v214.752 citationsh-index: 41

Originality Incremental advance

AI Analysis

This addresses the challenge of efficient data labeling in machine learning, particularly for imbalanced datasets, though it is an incremental improvement over existing active learning methods.

The paper tackles the problem of active learning for discriminative models by incorporating class imbalance and difficulty, showing that ignoring these factors is harmful. The proposed method, which uses variational Bayes and a VAE, significantly outperforms state-of-the-art methods on multiple datasets, including a real-world imbalanced one.

Active Learning for discriminative models has largely been studied with the focus on individual samples, with less emphasis on how classes are distributed or which classes are hard to deal with. In this work, we show that this is harmful. We propose a method based on the Bayes' rule, that can naturally incorporate class imbalance into the Active Learning framework. We derive that three terms should be considered together when estimating the probability of a classifier making a mistake for a given sample; i) probability of mislabelling a class, ii) likelihood of the data given a predicted class, and iii) the prior probability on the abundance of a predicted class. Implementing these terms requires a generative model and an intractable likelihood estimation. Therefore, we train a Variational Auto Encoder (VAE) for this purpose. To further tie the VAE with the classifier and facilitate VAE training, we use the classifiers' deep feature representations as input to the VAE. By considering all three probabilities, among them especially the data imbalance, we can substantially improve the potential of existing methods under limited data budget. We show that our method can be applied to classification tasks on multiple different datasets -- including one that is a real-world dataset with heavy data imbalance -- significantly outperforming the state of the art.

View on arXiv PDF

Similar