ME LGAug 6, 2025

Accept-Reject Lasso

arXiv:2508.04646v1h-index: 1

Originality Incremental advance

AI Analysis

This addresses a specific problem in statistical learning for researchers and practitioners using Lasso, but it is incremental as it builds on existing ensemble and clustering methods to enhance feature selection stability.

The paper tackles the instability of Lasso with highly correlated features, which leads to errors in omitting falsely redundant features and including truly redundant ones, by introducing the Accept-Reject Lasso (ARL) that uses clustering and subset analysis to differentiate between true and spurious correlations, resulting in improved feature selection as shown in simulations and real-data experiments.

The Lasso method is known to exhibit instability in the presence of highly correlated features, often leading to an arbitrary selection of predictors. This issue manifests itself in two primary error types: the erroneous omission of features that lack a true substitutable relationship (falsely redundant features) and the inclusion of features with a true substitutable relationship (truly redundant features). Although most existing methods address only one of these challenges, we introduce the Accept-Reject Lasso (ARL), a novel approach that resolves this dilemma. ARL operationalizes an Accept-Reject framework through a fine-grained analysis of feature selection across data subsets. This framework is designed to partition the output of an ensemble method into beneficial and detrimental components through fine-grained analysis. The fundamental challenge for Lasso is that inter-variable correlation obscures the true sources of information. ARL tackles this by first using clustering to identify distinct subset structures within the data. It then analyzes Lasso's behavior across these subsets to differentiate between true and spurious correlations. For truly correlated features, which induce multicollinearity, ARL tends to select a single representative feature and reject the rest to ensure model stability. Conversely, for features linked by spurious correlations, which may vanish in certain subsets, ARL accepts those that Lasso might have incorrectly omitted. The distinct patterns arising from true versus spurious correlations create a divisible separation. By setting an appropriate threshold, our framework can effectively distinguish between these two phenomena, thereby maximizing the inclusion of informative variables while minimizing the introduction of detrimental ones. We illustrate the efficacy of the proposed method through extensive simulation and real-data experiments.

View on arXiv PDF

Similar