MMAIJul 5

Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption

arXiv:2607.0455322.1Has Code
Predicted impact top 5% in MM · last 90 daysOriginality Incremental advance
AI Analysis

It provides a standardized, theory-grounded method for comparing energy efficiency of video generation models, addressing the need for sustainability assessment in this domain.

The paper introduces a bidirectional framework to estimate energy consumption of text-to-video and text-to-video-audio models from architectural principles and generation parameters, achieving below 3% MAPE across six models. It enables sustainability benchmarking without access to model weights.

We present a bidirectional framework for estimating the energy consumption of text-to-video (T2V) and text-to-video-audio (T2VA) models from architectural first principles and observable generation parameters such as resolution and duration, requiring no access to weights, model size, or implementation details. Forward, it predicts energy from generation parameters and architectural principles; backward, it recovers architectural scaling behavior from observed inference times, with accuracy serving as a criterion for architectural validity. Building on the established compute-bound nature of video diffusion models, we demonstrate that each model's energy profile obeys theoretically derived scaling laws, decomposing into quadratic and linear terms whose coefficients directly reflect the underlying architectural complexity. Validated across six open-source models spanning 8.3B-27B parameters and three GPU configurations, this decomposition achieves below 3% MAPE across all architectures. This approach offers a standardized, empirically and theoretically grounded framework for sustainability benchmarking across T2V models and architectures.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes