2.3NAJan 9, 2015
Parameter choice strategies for least-squares approximation of noisy smooth functions on the sphereSergei. V. Pereverzyev, Ian. H. Sloan, Pavlo Tkachenko
We consider a polynomial reconstruction of smooth functions from their noisy values at discrete nodes on the unit sphere by a variant of the regularized least-squares method of An et al., SIAM J. Numer. Anal. 50 (2012), 1513--1534. As nodes we use the points of a positive-weight cubature formula that is exact for all spherical polynomials of degree up to $2M$, where $M$ is the degree of the reconstructing polynomial. We first obtain a reconstruction error bound in terms of the regularization parameter and the penalization parameters in the regularization operator. Then we discuss a priori and a posteriori strategies for choosing these parameters. Finally, we give numerical examples illustrating the theoretical results.
1.2NAJan 2, 2015
Two-parameter regularization of ill-posed spherical pseudo-differential equations in the space of continuous functionsHui Cao, Sergei V. Pereverzyev, Ian H. Sloan et al.
In this paper, a two-step regularization method is used to solve an ill-posed spherical pseudo-differential equation in the presence of noisy data. For the first step of regularization we approximate the data by means of a spherical polynomial that minimizes a functional with a penalty term consisting of the squared norm in a Sobolev space. The second step is a regularized collocation method. An error bound is obtained in the uniform norm, which is potentially smaller than that for either the noise reduction alone or the regularized collocation alone. We discuss an a posteriori parameter choice, and present some numerical experiments, which support the claimed superiority of the two-step method.
9.2STAug 15, 2023
On regularized Radon-Nikodym differentiationDuc Hoan Nguyen, Werner Zellinger, Sergei V. Pereverzyev
We discuss the problem of estimating Radon-Nikodym derivatives. This problem appears in various applications, such as covariate shift adaptation, likelihood-ratio testing, mutual information estimation, and conditional probability estimation. To address the above problem, we employ the general regularization scheme in reproducing kernel Hilbert spaces. The convergence rate of the corresponding regularized algorithm is established by taking into account both the smoothness of the derivative and the capacity of the space in which it is estimated. This is done in terms of general source conditions and the regularized Christoffel functions. We also find that the reconstruction of Radon-Nikodym derivatives at any particular point can be done with high order of accuracy. Our theoretical results are illustrated by numerical simulations.
10.7LGJul 30, 2023
Adaptive learning of density ratios in RKHSWerner Zellinger, Stefan Kindermann, Sergei V. Pereverzyev
Estimating the ratio of two probability densities from finitely many observations of the densities is a central problem in machine learning and statistics with applications in two-sample testing, divergence estimation, generative modeling, covariate shift adaptation, conditional density estimation, and novelty detection. In this work, we analyze a large class of density ratio estimation methods that minimize a regularized Bregman divergence between the true density ratio and a model in a reproducing kernel Hilbert space (RKHS). We derive new finite-sample error bounds, and we propose a Lepskii type parameter choice principle that minimizes the bounds without knowledge of the regularity of the density ratio. In the special case of quadratic loss, our method adaptively achieves a minimax optimal error rate. A numerical illustration is provided.
Domain Generalization by Functional RegressionMarkus Holzleitner, Sergei V. Pereverzyev, Werner Zellinger
The problem of domain generalization is to learn, given data from different source distributions, a model that can be expected to generalize well on new target distributions which are only seen through unlabeled samples. In this paper, we study domain generalization as a problem of functional regression. Our concept leads to a new algorithm for learning a linear operator from marginal distributions of inputs to the corresponding conditional distributions of outputs given inputs. Our algorithm allows a source distribution-dependent construction of reproducing kernel Hilbert spaces for prediction, and, satisfies finite sample error bounds for the idealized risk. Numerical implementations and source code are available.
3.8LGJul 21, 2023
General regularization in covariate shift adaptationDuc Hoan Nguyen, Sergei V. Pereverzyev, Werner Zellinger
Sample reweighting is one of the most widely used methods for correcting the error of least squares learning algorithms in reproducing kernel Hilbert spaces (RKHS), that is caused by future data distributions that are different from the training data distribution. In practical situations, the sample weights are determined by values of the estimated Radon-Nikodým derivative, of the future data distribution w.r.t.~the training data distribution. In this work, we review known error bounds for reweighted kernel regression in RKHS and obtain, by combination, novel results. We show under weak smoothness conditions, that the amount of samples, needed to achieve the same order of accuracy as in the standard supervised learning without differences in data distributions, is smaller than proven by state-of-the-art analyses.
1.2STJan 28
Towards regularized learning from functional data with covariate shiftMarkus Holzleitner, Sergiy Pereverzyev, Sergei V. Pereverzyev et al.
This paper investigates a general regularization framework for unsupervised domain adaptation in vector-valued regression under the covariate shift assumption, utilizing vector-valued reproducing kernel Hilbert spaces (vRKHS). Covariate shift occurs when the input distributions of the training and test data differ, introducing significant challenges for reliable learning. By restricting the hypothesis space, we develop a practical operator learning algorithm capable of handling functional outputs. We establish optimal convergence rates for the proposed framework under a general source condition, providing a theoretical foundation for regularized learning in this setting. We also propose an aggregation-based approach that forms a linear combination of estimators corresponding to different regularization parameters and different kernels. The proposed approach addresses the challenge of selecting appropriate tuning parameters, which is crucial for constructing a good estimator, and we provide a theoretical justification for its effectiveness. Furthermore, we illustrate the proposed method on a real-world face image dataset, demonstrating robustness and effectiveness in mitigating distributional discrepancies under covariate shift.
3.1MLMay 7, 2024
Multiparameter regularization and aggregation in the context of polynomial functional regressionElke R. Gizewski, Markus Holzleitner, Lukas Mayer-Suess et al.
Most of the recent results in polynomial functional regression have been focused on an in-depth exploration of single-parameter regularization schemes. In contrast, in this study we go beyond that framework by introducing an algorithm for multiple parameter regularization and presenting a theoretically grounded method for dealing with the associated parameters. This method facilitates the aggregation of models with varying regularization parameters. The efficacy of the proposed approach is assessed through evaluations on both synthetic and some real-world medical data, revealing promising results.
Addressing Parameter Choice Issues in Unsupervised Domain Adaptation by AggregationMarius-Constantin Dinu, Markus Holzleitner, Maximilian Beck et al.
We study the problem of choosing algorithm hyper-parameters in unsupervised domain adaptation, i.e., with labeled data in a source domain and unlabeled data in a target domain, drawn from a different input distribution. We follow the strategy to compute several models using different hyper-parameters, and, to subsequently compute a linear aggregation of the models. While several heuristics exist that follow this strategy, methods are still missing that rely on thorough theories for bounding the target error. In this turn, we propose a method that extends weighted least squares to vector-valued functions, e.g., deep neural networks. We show that the target error of the proposed algorithm is asymptotically not worse than twice the error of the unknown optimal aggregation. We also perform a large scale empirical comparative study on several datasets, including text, images, electroencephalogram, body sensor signals and signals from mobile phones. Our method outperforms deep embedded validation (DEV) and importance weighted validation (IWV) on all datasets, setting a new state-of-the-art performance for solving parameter choice issues in unsupervised domain adaptation with theoretical error guarantees. We further study several competitive heuristics, all outperforming IWV and DEV on at least five datasets. However, our method outperforms each heuristic on at least five of seven datasets.
3.1LGMay 12, 2021
A function approximation approach to the prediction of blood glucose levelsH. N. Mhaskar, S. V. Pereverzyev, M. D. van der Walt
The problem of real time prediction of blood glucose (BG) levels based on the readings from a continuous glucose monitoring (CGM) device is a problem of great importance in diabetes care, and therefore, has attracted a lot of research in recent years, especially based on machine learning. An accurate prediction with a 30, 60, or 90 minute prediction horizon has the potential of saving millions of dollars in emergency care costs. In this paper, we treat the problem as one of function approximation, where the value of the BG level at time $t+h$ (where $h$ the prediction horizon) is considered to be an unknown function of $d$ readings prior to the time $t$. This unknown function may be supported in particular on some unknown submanifold of the $d$-dimensional Euclidean space. While manifold learning is classically done in a semi-supervised setting, where the entire data has to be known in advance, we use recent ideas to achieve an accurate function approximation in a supervised setting; i.e., construct a model for the target function. We use the state-of-the-art clinically relevant PRED-EGA grid to evaluate our results, and demonstrate that for a real life dataset, our method performs better than a standard deep network, especially in hypoglycemic and hyperglycemic regimes. One noteworthy aspect of this work is that the training data and test data may come from different distributions.
1.2NAJul 8, 2015
A Parameter Choice Strategy for the Inversion of Multiple ObservationsC. Gerhards, S. Pereverzyev, P. Tkachenko
In many geoscientific applications, multiple noisy observations of different origin need to be combined to improve the reconstruction of a common underlying quantity. This naturally leads to multi-parameter models for which adequate strategies are required to choose a set of 'good' parameters. In this study, we present a fairly general method for choosing such a set of parameters, provided that discrete direct, but maybe noisy, measurements of the underlying quantity are included in the observation data, and the inner product of the reconstruction space can be accurately estimated by the inner product of the discretization space. Then the proposed parameter choice method gives an accuracy that only by an absolute constant multiplier differs from the noise level and the accuracy of the best approximant in the reconstruction and in the discretization spaces.