Alexander Selvikvåg Lundervold

LG
h-index13
5papers
122citations
Novelty48%
AI Score25

5 Papers

7.7LGFeb 21, 2023
Does the evaluation stand up to evaluation? A first-principle approach to the evaluation of classifiers

K. Dyrland, A. S. Lundervold, P. G. L. Porta Mana

How can one meaningfully make a measurement, if the meter does not conform to any standard and its scale expands or shrinks depending on what is measured? In the present work it is argued that current evaluation practices for machine-learning classifiers are affected by this kind of problem, leading to negative consequences when classifiers are put to real use; consequences that could have been avoided. It is proposed that evaluation be grounded on Decision Theory, and the implications of such foundation are explored. The main result is that every evaluation metric must be a linear combination of confusion-matrix elements, with coefficients - "utilities" - that depend on the specific classification problem. For binary classification, the space of such possible metrics is effectively two-dimensional. It is shown that popular metrics such as precision, balanced accuracy, Matthews Correlation Coefficient, Fowlkes-Mallows index, F1-measure, and Area Under the Curve are never optimal: they always give rise to an in-principle avoidable fraction of incorrect evaluations. This fraction is even larger than would be caused by the use of a decision-theoretic metric with moderately wrong coefficients.

2.0LGFeb 21, 2023
Don't guess what's true: choose what's optimal. A probability transducer for machine-learning classifiers

K. Dyrland, A. S. Lundervold, P. G. L. Porta Mana

In fields such as medicine and drug discovery, the ultimate goal of a classification is not to guess a class, but to choose the optimal course of action among a set of possible ones, usually not in one-one correspondence with the set of classes. This decision-theoretic problem requires sensible probabilities for the classes. Probabilities conditional on the features are computationally almost impossible to find in many important cases. The main idea of the present work is to calculate probabilities conditional not on the features, but on the trained classifier's output. This calculation is cheap, needs to be made only once, and provides an output-to-probability "transducer" that can be applied to all future outputs of the classifier. In conjunction with problem-dependent utilities, the probabilities of the transducer allow us to find the optimal choice among the classes or among a set of more general decisions, by means of expected-utility maximization. This idea is demonstrated in a simplified drug-discovery problem with a highly imbalanced dataset. The transducer and utility maximization together always lead to improved results, sometimes close to theoretical maximum, for all sets of problem-dependent utilities. The one-time-only calculation of the transducer also provides, automatically: (i) a quantification of the uncertainty about the transducer itself; (ii) the expected utility of the augmented algorithm (including its uncertainty), which can be used for algorithm selection; (iii) the possibility of using the algorithm in a "generative mode", useful if the training dataset is biased.

1.2GNMar 31, 2022
rfPhen2Gen: A machine learning based association study of brain imaging phenotypes to genotypes

Muhammad Ammar Malik, Alexander S. Lundervold, Tom Michoel

Imaging genetic studies aim to find associations between genetic variants and imaging quantitative traits. Traditional genome-wide association studies (GWAS) are based on univariate statistical tests, but when multiple traits are analyzed together they suffer from a multiple-testing problem and from not taking into account correlations among the traits. An alternative approach to multi-trait GWAS is to reverse the functional relation between genotypes and traits, by fitting a multivariate regression model to predict genotypes from multiple traits simultaneously. However, current reverse genotype prediction approaches are mostly based on linear models. Here, we evaluated random forest regression (RFR) as a method to predict SNPs from imaging QTs and identify biologically relevant associations. We learned machine learning models to predict 518,484 SNPs using 56 brain imaging QTs. We observed that genotype regression error is a better indicator of permutation p-value significance than genotype classification accuracy. SNPs within the known Alzheimer disease (AD) risk gene APOE had lowest RMSE for lasso and random forest, but not ridge regression. Moreover, random forests identified additional SNPs that were not prioritized by the linear models but are known to be associated with brain-related disorders. Feature selection identified well-known brain regions associated with AD,like the hippocampus and amygdala, as important predictors of the most significant SNPs. In summary, our results indicate that non-linear methods like random forests may offer additional insights into phenotype-genotype associations compared to traditional linear multi-variate GWAS methods.

1.2NAMay 4, 2015
On the Lie enveloping algebra of a post-Lie algebra

Kurusch Ebrahimi-Fard, Alexander Lundervold, Hans Munthe-Kaas

We consider pairs of Lie algebras $g$ and $\bar{g}$, defined over a common vector space, where the Lie brackets of $g$ and $\bar{g}$ are related via a post-Lie algebra structure. The latter can be extended to the Lie enveloping algebra $U(g)$. This permits us to define another associative product on $U(g)$, which gives rise to a Hopf algebra isomorphism between $U(\bar{g})$ and a new Hopf algebra assembled from $U(g)$ with the new product. For the free post-Lie algebra these constructions provide a refined understanding of a fundamental Hopf algebra appearing in the theory of numerical integration methods for differential equations on manifolds. In the pre-Lie setting, the algebraic point of view developed here also provides a concise way to develop Butcher's order theory for Runge--Kutta methods.

8.0NASep 23, 2010
Hopf algebras of formal diffeomorphisms and numerical integration on manifolds

Alexander Lundervold, Hans Munthe-Kaas

B-series originated from the work of John Butcher in the 1960s as a tool to analyze numerical integration of differential equations, in particular Runge-Kutta methods. Connections to renormalization theory in perturbative quantum field theory have been established in recent years. The algebraic structure of classical Runge-Kutta methods is described by the Connes-Kreimer Hopf algebra. Lie-Butcher theory is a generalization of B-series aimed at studying Lie-group integrators for differential equations evolving on manifolds. Lie-group integrators are based on general Lie group actions on a manifold, and classical Runge-Kutta integrators appear in this setting as the special case of R^n acting upon itself by translations. Lie--Butcher theory combines classical B-series on R^n with Lie-series on manifolds. The underlying Hopf algebra combines the Connes-Kreimer Hopf algebra with the shuffle Hopf algebra of free Lie algebras. We give an introduction to Hopf algebraic structures and their relationship to structures appearing in numerical analysis, aimed at a general mathematical audience. In particular we explore the close connection between Lie series, time-dependent Lie series and Lie--Butcher series for diffeomorphisms on manifolds. The role of the Euler and Dynkin idempotents in numerical analysis is discussed. A non-commutative version of a Faa di Bruno bialgebra is introduced, and the relation to non-commutative Bell polynomials is explored.