Quantitative Gaussian-Process limits of Tensor Programs
For theorists and practitioners of neural networks, this work offers quantitative guarantees for the Gaussian-process approximation, which is incremental as it extends known convergence results to a broader class of architectures.
The paper provides explicit finite-width error bounds (order inverse square-root of widths) for the convergence of random neural networks to their Gaussian-process limits in Wasserstein distance, covering feed-forward, recurrent, and transformer architectures.
We study the infinite-width Gaussian-process limit of random neural networks through the lens of tensor programs, and we provide a quantitative convergence theory in Wasserstein distance. Our main result gives explicit finite-width error bounds, of order inverse square-root of the widths between finite-network executions and their Gaussian-process limits. The framework is architecture-agnostic and covers feed-forward models together with weight-sharing schemes relevant for recurrent and transformer-type architectures.