Power-Law Spectrum of the Random Feature Model

Elliot Paquette, Ke Liang Xiao, Yizhe Zhu

arXiv:2603.1457854.81 citationsh-index: 2

Predicted impact top 22% in ML · last 90 daysOriginality Incremental advance

AI Analysis

This provides theoretical insight into scaling laws in neural networks for researchers in machine learning theory, though it is incremental as it builds on known models.

The paper tackles the question of whether power-law eigenvalue decay in data covariance is preserved when data passes through a random feature model with a monomial activation, proving that the exponent is inherited exactly with only logarithmic corrections depending on the monomial degree.

Scaling laws for neural networks, in which the loss decays as a power-law in the number of parameters, data, and compute, depend fundamentally on the spectral structure of the data covariance, with power-law eigenvalue decay appearing ubiquitously in vision and language tasks. A central question is whether this spectral structure is preserved or destroyed when data passes through the basic building block of a neural network: a random linear projection followed by a nonlinear activation. We study this question for the random feature model: given data $x \sim N(0,H)\in \mathbb{R}^v$ where $H$ has $Î±$-power-law spectrum ($Î»_j(H ) \asymp j^{-Î±}$, $Î±> 1$), a Gaussian sketch matrix $W \in \mathbb{R}^{v\times d}$, and an entrywise monomial $f(y) = y^{p}$, we characterize the eigenvalues of the population random-feature covariance $\mathbb{E}_{x }[\frac{1}{d}f(W^\top x )^{\otimes 2}]$. We prove matching upper and lower bounds: for all $1 \leq j \leq c_1 d \log^{-(p+1)}(d)$, the $j$-th eigenvalue is of order $\left(\log^{p-1}(j+1)/j\right)^Î±$. For $ c_1 d \log^{-(p+1)}(d)\leq j\leq d$, the $j$-th eigenvalue is of order $j^{-Î±}$ up to a polylog factor. That is, the power-law exponent $Î±$ is inherited exactly from the input covariance, modified only by a logarithmic correction that depends on the monomial degree $p$. The proof combines a dyadic head-tail decomposition with Wick chaos expansions for higher-order monomials and random matrix concentration inequalities.

View on arXiv PDF

Similar