Conformal calibration and look-elsewhere effect in anomaly detection for new-physics searches

arXiv:2606.137808.4h-index: 3
Predicted impact top 54% in HEP-PH · last 90 daysOriginality Highly original
AI Analysis

For experimental high-energy physics searches, this provides a statistically rigorous, detector-agnostic method to obtain trial-factor-aware significances from anomaly detectors, correcting miscalibration and false positives that plague current approaches.

The paper addresses the lack of calibrated significance in machine-learned anomaly detection for new-physics searches, proposing a conformal prediction-based calibration layer that converts any anomaly score into a valid significance with distribution-free guarantees. On LHC Olympics data, it removes a fabricated 46σ excess from background sculpting and eliminates false alarms in signal-free windows, while standard methods produce ≥10σ false excesses.

Machine-learned anomaly detection is reshaping searches for new physics, but it has outrun the statistics used to interpret it. A raw anomaly score has no calibrated meaning, a model that scans many regions inflates the look-elsewhere effect, and the asymptotic significances the field relies on are blind to the background mismodelling that anomaly detectors are especially prone to. We propose a calibration layer, built on conformal prediction, that turns any anomaly score into a defensible significance with distribution-free, finite-sample guarantees. Conformal prediction converts scores into valid local p-values, weighted and Mondrian variants repair the sideband-to-signal-region exchangeability failures that resonant searches suffer, and a Gross-Vitells step carries the result through to a look-elsewhere-aware global significance. The layer does two things at once. It exposes miscalibration that the standard pipeline cannot see, and it corrects it without retraining the detector. On public LHC Olympics data, a classifier develops a substructure-mass correlation that makes sideband-calibrated background p-values anti-conservative. Taken at face value, this manufactures a $\sim 46σ$ excess from background sculpting alone, which the label-free weighted correction removes, restoring an honest null. When run as a blind wide-mass bump hunt, the standard asymptotic and unweighted procedures fabricate $\gtrsim10σ$ excesses and $\approx5σ$ excesses even in signal-free windows, while the conformal layer raises no false alarms and its global false-positive rate is verified on background-only pseudoexperiments. The result is an auditable, detector-agnostic path from an uncalibrated score to a trials-factor-aware significance, ready to be folded into experimental anomaly searches.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes