IT LGDec 21, 2020

Neural Joint Entropy Estimation

Yuval Shalev, Amichai Painsky, Irad Ben-Gal

arXiv:2012.11197v18.615 citations

Originality Incremental advance

AI Analysis

This work provides a practical solution for more accurate entropy and related information-theoretic measure estimation, particularly beneficial for researchers and practitioners in machine learning, statistics, and data compression when dealing with limited data.

This paper addresses the challenge of estimating the entropy of discrete random variables, especially with small sample sizes relative to the alphabet size. It introduces a neural network-based approach that leverages the generalization capabilities of DNNs for cross-entropy estimation, resulting in improved accuracy across various information-theoretic measures.

Estimating the entropy of a discrete random variable is a fundamental problem in information theory and related fields. This problem has many applications in various domains, including machine learning, statistics and data compression. Over the years, a variety of estimation schemes have been suggested. However, despite significant progress, most methods still struggle when the sample is small, compared to the variable's alphabet size. In this work, we introduce a practical solution to this problem, which extends the work of McAllester and Statos (2020). The proposed scheme uses the generalization abilities of cross-entropy estimation in deep neural networks (DNNs) to introduce improved entropy estimation accuracy. Furthermore, we introduce a family of estimators for related information-theoretic measures, such as conditional entropy and mutual information. We show that these estimators are strongly consistent and demonstrate their performance in a variety of use-cases. First, we consider large alphabet entropy estimation. Then, we extend the scope to mutual information estimation. Next, we apply the proposed scheme to conditional mutual information estimation, as we focus on independence testing tasks. Finally, we study a transfer entropy estimation problem. The proposed estimators demonstrate improved performance compared to existing methods in all tested setups.

View on arXiv PDF

Similar