Gaussian Process Latent Factor Regression for Low-Data, High-Dimensional Output Problems

arXiv:2606.065766.9
Predicted impact top 69% in LG · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the problem of high-dimensional output regression in low-data regimes, which is common in scientific applications, by providing a principled method that outperforms existing compress-then-predict pipelines.

The authors propose Gaussian process latent factor regression (GPLFR), a model that jointly optimizes compression and prediction for high-dimensional outputs from few training examples, and demonstrate it by building the first spatially resolved emulator of global climate models for rocky exoplanets.

In the sciences, regression tasks often require predicting high-dimensional outputs from few training examples. Multi-output Gaussian processes excel in low-data regimes but typically struggle with high-dimensional outputs. Compress-then-predict pipelines such as PCA-GP (principal component analysis plus Gaussian process regression) handle high dimensionality, but rely on bases optimized for reconstruction rather than prediction. To address this gap, we propose a model that represents each output as a linear-Gaussian decoding of a low-dimensional latent state drawn from a Gaussian process prior. By analytically marginalizing the decoder weights, we couple compression and prediction in a single objective that scales to high-dimensional outputs. We refer to this model as Gaussian process latent factor regression (GPLFR). We demonstrate GPLFR by building the first spatially resolved emulator of global climate models for rocky exoplanets.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes