LG AIOct 11, 2021

Homogeneous Learning: Self-Attention Decentralized Deep Learning

arXiv:2110.05290v15.511 citationsHas Code

Originality Incremental advance

AI Analysis

This addresses efficiency and robustness issues in decentralized learning for applications like medical imaging, though it appears incremental over existing decentralized approaches.

The paper tackles the problem of slow convergence in decentralized deep learning with non-IID data by proposing Homogeneous Learning (HL), which uses a self-attention mechanism to select nodes for training, resulting in a 50.8% reduction in training rounds and 74.6% reduction in communication cost compared to random policy-based methods.

Federated learning (FL) has been facilitating privacy-preserving deep learning in many walks of life such as medical image classification, network intrusion detection, and so forth. Whereas it necessitates a central parameter server for model aggregation, which brings about delayed model communication and vulnerability to adversarial attacks. A fully decentralized architecture like Swarm Learning allows peer-to-peer communication among distributed nodes, without the central server. One of the most challenging issues in decentralized deep learning is that data owned by each node are usually non-independent and identically distributed (non-IID), causing time-consuming convergence of model training. To this end, we propose a decentralized learning model called Homogeneous Learning (HL) for tackling non-IID data with a self-attention mechanism. In HL, training performs on each round's selected node, and the trained model of a node is sent to the next selected node at the end of each round. Notably, for the selection, the self-attention mechanism leverages reinforcement learning to observe a node's inner state and its surrounding environment's state, and find out which node should be selected to optimize the training. We evaluate our method with various scenarios for an image classification task. The result suggests that HL can produce a better performance compared with standalone learning and greatly reduce both the total training rounds by 50.8% and the communication cost by 74.6% compared with random policy-based decentralized learning for training on non-IID data.

View on arXiv PDF Code

Similar