On the Smallness of the Large Language Models Scaling Exponents
For the LLM community, this work highlights a fundamental energy sustainability issue in scaling laws, but the analysis is largely conceptual and lacks concrete empirical validation.
The paper argues that the small scaling exponents of LLMs indicate an unsustainable energy regime, and shows that accounting for the pedestal effect does not resolve this unsustainability. It also comments on the influence of data smoothness on scaling exponents via an analogy with fluid turbulence.
We discuss reasons why the scaling exponents of current Large Language Models (LLMs) applications are indicating an unsustainable regime in terms of energy resources. We further show that attributing the smallness of such exponents to a numerical bias due to the neglect of a non-zero value of the loss function in the limit of infinite data (``pedestal effect") does not remove the unsustainability issue. Finally, the effects of the smoothness (roughness) of the data on the scaling exponents is commented upon based on an analogy with phenomenological models of fluid turbulence.