LGMay 9, 2025

Deep-ICE: The first globally optimal algorithm for empirical risk minimization of two-layer maxout and ReLU networks

arXiv:2505.05740v111.43 citationsh-index: 2

Originality Highly original

AI Analysis

It solves the problem of finding exact solutions for neural network training, which is incremental as it builds on existing optimization methods but with provable guarantees.

The paper introduces the first globally optimal algorithm for empirical risk minimization of two-layer maxout and ReLU networks, achieving a worst-case time complexity of O(N^{DK+1}) and demonstrating a 20-30% reduction in misclassifications compared to state-of-the-art methods.

This paper introduces the first globally optimal algorithm for the empirical risk minimization problem of two-layer maxout and ReLU networks, i.e., minimizing the number of misclassifications. The algorithm has a worst-case time complexity of $O\left(N^{DK+1}\right)$, where $K$ denotes the number of hidden neurons and $D$ represents the number of features. It can be can be generalized to accommodate arbitrary computable loss functions without affecting its computational complexity. Our experiments demonstrate that the proposed algorithm provides provably exact solutions for small-scale datasets. To handle larger datasets, we introduce a novel coreset selection method that reduces the data size to a manageable scale, making it feasible for our algorithm. This extension enables efficient processing of large-scale datasets and achieves significantly improved performance, with a 20-30\% reduction in misclassifications for both training and prediction, compared to state-of-the-art approaches (neural networks trained using gradient descent and support vector machines), when applied to the same models (two-layer networks with fixed hidden nodes and linear models).

View on arXiv PDF

Similar