LG MLApr 26, 2019

Classification from Pairwise Similarities/Dissimilarities and Unlabeled Data via Empirical Risk Minimization

Takuya Shimada, Han Bao, Issei Sato, Masashi Sugiyama

arXiv:1904.11717v112.849 citations

Originality Incremental advance

AI Analysis

This work addresses classification in scenarios where only pairwise information is available, such as privacy-aware settings, offering a more flexible approach than previous methods.

The paper tackles the problem of classification using pairwise similarities and dissimilarities along with unlabeled data, deriving an unbiased risk estimator that handles both types of pairwise information without strong geometric assumptions. The result includes theoretical error bounds and experimental validation of the method's practical usefulness.

Pairwise similarities and dissimilarities between data points might be easier to obtain than fully labeled data in real-world classification problems, e.g., in privacy-aware situations. To handle such pairwise information, an empirical risk minimization approach has been proposed, giving an unbiased estimator of the classification risk that can be computed only from pairwise similarities and unlabeled data. However, this direction cannot handle pairwise dissimilarities so far. On the other hand, semi-supervised clustering is one of the methods which can use both similarities and dissimilarities. Nevertheless, they typically require strong geometrical assumptions on the data distribution such as the manifold assumption, which may deteriorate the performance. In this paper, we derive an unbiased risk estimator which can handle all of similarities/dissimilarities and unlabeled data. We theoretically establish estimation error bounds and experimentally demonstrate the practical usefulness of our empirical risk minimization method.

View on arXiv PDF

Similar