LG MLDec 4, 2019

Extreme Learning Machine design for dealing with unrepresentative features

Nicolás Nieto, Francisco Ibarrola, Victoria Peterson, Hugo Rufiner, Ruben Spies

arXiv:1912.02154v11.03 citationsHas Code

Originality Incremental advance

AI Analysis

This work addresses a specific issue in ELM design for handling unrepresentative features, particularly in domains like EEG signal processing, and is incremental as it builds on existing ELM methods.

The authors tackled the problem of unrepresentative features in Extreme Learning Machines (ELMs) by proposing to use a much larger number of hidden nodes than traditionally recommended, along with a pruning algorithm to reduce computational burden. Experimental results on EEG signals showed improved performance over traditional ELM approaches while diminishing extra computing time.

Extreme Learning Machines (ELMs) have become a popular tool in the field of Artificial Intelligence due to their very high training speed and generalization capabilities. Another advantage is that they have a single hyper-parameter that must be tuned up: the number of hidden nodes. Most traditional approaches dictate that this parameter should be chosen smaller than the number of available training samples in order to avoid over-fitting. In fact, it has been proved that choosing the number of hidden nodes equal to the number of training samples yields a perfect training classification with probability 1 (w.r.t. the random parameter initialization). In this article we argue that in spite of this, in some cases it may be beneficial to choose a much larger number of hidden nodes, depending on certain properties of the data. We explain why this happens and show some examples to illustrate how the model behaves. In addition, we present a pruning algorithm to cope with the additional computational burden associated to the enlarged ELM. Experimental results using electroencephalography (EEG) signals show an improvement in performance with respect to traditional ELM approaches, while diminishing the extra computing time associated to the use of large architectures.

View on arXiv PDF Code

Similar