ML LG DSJun 2, 2025

Signature Maximum Mean Discrepancy Two-Sample Statistical Tests

Andrew Alden, Blanka Horvath, Zacharia Issa

arXiv:2506.01718v115.54 citationsh-index: 4

Originality Synthesis-oriented

AI Analysis

This work addresses statistical testing for stochastic processes, which is incremental as it extends MMD to path spaces with practical error mitigation.

The paper tackles the problem of using signature Maximum Mean Discrepancy (sig-MMD) for two-sample testing on path space distributions, identifying and mitigating Type 2 errors that occur in limited data settings.

Maximum Mean Discrepancy (MMD) is a widely used concept in machine learning research which has gained popularity in recent years as a highly effective tool for comparing (finite-dimensional) distributions. Since it is designed as a kernel-based method, the MMD can be extended to path space valued distributions using the signature kernel. The resulting signature MMD (sig-MMD) can be used to define a metric between distributions on path space. Similarly to the original use case of the MMD as a test statistic within a two-sample testing framework, the sig-MMD can be applied to determine if two sets of paths are drawn from the same stochastic process. This work is dedicated to understanding the possibilities and challenges associated with applying the sig-MMD as a statistical tool in practice. We introduce and explain the sig-MMD, and provide easily accessible and verifiable examples for its practical use. We present examples that can lead to Type 2 errors in the hypothesis test, falsely indicating that samples have been drawn from the same underlying process (which generally occurs in a limited data setting). We then present techniques to mitigate the occurrence of this type of error.

View on arXiv PDF

Similar