PAC-Bayes Analysis Beyond the Usual Bounds
This work addresses generalization guarantees in machine learning, offering incremental improvements by relaxing prior assumptions in PAC-Bayes theory.
The paper tackles the problem of guaranteeing generalization for stochastic learning models using PAC-Bayes analysis, presenting a basic inequality that extends known bounds and allows for data-dependent priors and unbounded losses, with specific examples like one for the unbounded square loss.
We focus on a stochastic learning model where the learner observes a finite set of training examples and the output of the learning process is a data-dependent distribution over a space of hypotheses. The learned data-dependent distribution is then used to make randomized predictions, and the high-level theme addressed here is guaranteeing the quality of predictions on examples that were not seen during training, i.e. generalization. In this setting the unknown quantity of interest is the expected risk of the data-dependent randomized predictor, for which upper bounds can be derived via a PAC-Bayes analysis, leading to PAC-Bayes bounds. Specifically, we present a basic PAC-Bayes inequality for stochastic kernels, from which one may derive extensions of various known PAC-Bayes bounds as well as novel bounds. We clarify the role of the requirements of fixed 'data-free' priors, bounded losses, and i.i.d. data. We highlight that those requirements were used to upper-bound an exponential moment term, while the basic PAC-Bayes theorem remains valid without those restrictions. We present three bounds that illustrate the use of data-dependent priors, including one for the unbounded square loss.