LGDec 20, 2024

Function Space Diversity for Uncertainty Prediction via Repulsive Last-Layer Ensembles

Sophie Steger, Christian Knoll, Bernhard Klein, Holger Fröning, Franz Pernkopf

arXiv:2412.15758v111.58 citationsh-index: 35

Originality Incremental advance

AI Analysis

This work addresses uncertainty prediction for neural networks, particularly in scenarios like active learning and distribution shifts, but it is incremental as it builds on existing function space inference with practical modifications.

The paper tackles the challenge of approximating Bayesian inference in function space for neural networks by proposing a repulsive last-layer ensemble method that uses a single multi-headed network to improve uncertainty estimation with minimal computational overhead. The method achieves competitive results in active learning, out-of-domain detection, and calibrated uncertainty under distribution shifts.

Bayesian inference in function space has gained attention due to its robustness against overparameterization in neural networks. However, approximating the infinite-dimensional function space introduces several challenges. In this work, we discuss function space inference via particle optimization and present practical modifications that improve uncertainty estimation and, most importantly, make it applicable for large and pretrained networks. First, we demonstrate that the input samples, where particle predictions are enforced to be diverse, are detrimental to the model performance. While diversity on training data itself can lead to underfitting, the use of label-destroying data augmentation, or unlabeled out-of-distribution data can improve prediction diversity and uncertainty estimates. Furthermore, we take advantage of the function space formulation, which imposes no restrictions on network parameterization other than sufficient flexibility. Instead of using full deep ensembles to represent particles, we propose a single multi-headed network that introduces a minimal increase in parameters and computation. This allows seamless integration to pretrained networks, where this repulsive last-layer ensemble can be used for uncertainty aware fine-tuning at minimal additional cost. We achieve competitive results in disentangling aleatoric and epistemic uncertainty for active learning, detecting out-of-domain data, and providing calibrated uncertainty estimates under distribution shifts with minimal compute and memory.

View on arXiv PDF

Similar