QMLGJul 2

Structured Gaussian Processes for Uncertainty-Aware Classification of High-Dimensional, Small-Sampled Omics Data

arXiv:2607.021031.2
Predicted impact top 97% in QM · last 90 daysOriginality Incremental advance
AI Analysis

For computational biologists analyzing high-dimensional, small-sample omics data with class imbalance, this work offers a method that leverages known biological interactions to improve classification and uncertainty quantification.

The paper introduces a structured Gaussian process classifier that integrates biological pathway graphs into kernel construction for high-dimensional, small-sample omics data. On three microbiome datasets, the hybrid approach outperforms unstructured baselines and matches established benchmarks, while providing calibrated uncertainty estimates.

Classifying heterogeneous omics data remains a fundamental challenge in computational biology, particularly in high-dimensional, small-sample settings where nonlinear interactions dominate and class imbalance further complicates reliable prediction of minority phenotypes. While traditional kernel methods rely on feature abundance, they fail to leverage the known interaction landscapes of biological systems. In this work, we propose a structured Gaussian process classification framework that integrates graph-encoded biological pathways directly into the kernel construction. By propagating information along known interaction networks and combining this with abundance-derived features, the resulting classifier captures both quantitative measurements and topological context. We benchmark our proposed methodology on three publicly available gut and fecal microbiome datasets. To address severe class imbalance, we evaluate complementary strategies, including data-level resampling, threshold calibration, and confusion-matrix-based adjustments, and report minority-class performance alongside accuracy. The hybrid approach yields a performance gain over unstructured baselines and matches the performance of established benchmarks for similar datasets. Furthermore, the probabilistic nature of the framework naturally provides calibrated predictive uncertainty, enabling robust differentiation between confident predictions and ambiguous samples.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes