ML CV LGFeb 23, 2018

A DIRT-T Approach to Unsupervised Domain Adaptation

Rui Shu, Hung H. Bui, Hirokazu Narui, Stefano Ermon

arXiv:1802.08735v241.1663 citationsHas Code

Originality Highly original

AI Analysis

This work addresses a critical problem in machine learning for adapting models across domains with scarce labels, offering a novel solution that advances state-of-the-art performance in specific applications.

The paper tackles limitations in domain adversarial training for unsupervised domain adaptation by proposing two models, VADA and DIRT-T, which incorporate the cluster assumption to avoid decision boundaries crossing high-density regions, resulting in significant improvements on benchmarks like digit, traffic sign, and Wi-Fi recognition.

Domain adaptation refers to the problem of leveraging labeled data in a source domain to learn an accurate model in a target domain where labels are scarce or unavailable. A recent approach for finding a common representation of the two domains is via domain adversarial training (Ganin & Lempitsky, 2015), which attempts to induce a feature extractor that matches the source and target feature distributions in some feature space. However, domain adversarial training faces two critical limitations: 1) if the feature extraction function has high-capacity, then feature distribution matching is a weak constraint, 2) in non-conservative domain adaptation (where no single classifier can perform well in both the source and target domains), training the model to do well on the source domain hurts performance on the target domain. In this paper, we address these issues through the lens of the cluster assumption, i.e., decision boundaries should not cross high-density data regions. We propose two novel and related models: 1) the Virtual Adversarial Domain Adaptation (VADA) model, which combines domain adversarial training with a penalty term that punishes the violation the cluster assumption; 2) the Decision-boundary Iterative Refinement Training with a Teacher (DIRT-T) model, which takes the VADA model as initialization and employs natural gradient steps to further minimize the cluster assumption violation. Extensive empirical results demonstrate that the combination of these two models significantly improve the state-of-the-art performance on the digit, traffic sign, and Wi-Fi recognition domain adaptation benchmarks.

View on arXiv PDF Code

Similar