LGAIJun 16

Conservation Laws for Modern Neural Architectures

arXiv:2606.178165.2
Predicted impact top 77% in LG · last 90 daysOriginality Incremental advance
AI Analysis

Provides theoretical understanding of implicit bias in over-parameterized modern architectures, which is important for explaining their success.

This work develops a unified framework to characterize conservation laws in gradient flow for modern neural architectures, including GELU, SiLU, SwiGLU activations, multihead attention with positional encodings, and Mixture-of-Experts. Experiments validate the predicted invariants.

Understanding gradient descent dynamics is key to explaining the success of over-parameterized models, where implicit bias manifests through conservation laws in gradient flow. While such laws are well understood for linear and ReLU networks, they remain largely unexplored for modern architectures. This work develops a unified framework to characterize conservation laws for contemporary models, including feedforward networks with GELU, SiLU, and SwiGLU activations, multihead attention with sinusoidal and rotary positional encodings, and Mixture-of-Experts architectures under diverse gating designs. Our theoretical findings are supported by experiments that validate the predicted invariants.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes