LG AIJun 13, 2022

Rank Diminishing in Deep Neural Networks

Ruili Feng, Kecheng Zheng, Yukun Huang, Deli Zhao, Michael Jordan, Zheng-Jun Zha

arXiv:2206.06072v124.354 citationsh-index: 78

Originality Incremental advance

AI Analysis

This work addresses a fundamental gap in understanding rank dynamics in deep neural networks, which is incremental but may advance theoretical principles for researchers in machine learning.

The authors tackled the problem of understanding the intrinsic mechanism behind low-rank structures in deep neural networks by theoretically establishing a universal monotonic decreasing property of network rank and providing the first empirical analysis of per-layer rank behavior in ResNets, MLPs, and Transformers on ImageNet, with results aligning with their theory.

The rank of neural networks measures information flowing across layers. It is an instance of a key structural condition that applies across broad domains of machine learning. In particular, the assumption of low-rank feature representations leads to algorithmic developments in many architectures. For neural networks, however, the intrinsic mechanism that yields low-rank structures remains vague and unclear. To fill this gap, we perform a rigorous study on the behavior of network rank, focusing particularly on the notion of rank deficiency. We theoretically establish a universal monotonic decreasing property of network rank from the basic rules of differential and algebraic composition, and uncover rank deficiency of network blocks and deep function coupling. By virtue of our numerical tools, we provide the first empirical analysis of the per-layer behavior of network rank in practical settings, i.e., ResNets, deep MLPs, and Transformers on ImageNet. These empirical results are in direct accord with our theory. Furthermore, we reveal a novel phenomenon of independence deficit caused by the rank deficiency of deep networks, where classification confidence of a given category can be linearly decided by the confidence of a handful of other categories. The theoretical results of this work, together with the empirical findings, may advance understanding of the inherent principles of deep neural networks.

View on arXiv PDF

Similar