Zhehang Du, Weijie Su
This work provides a theoretical foundation for an existing optimizer and offers a slight performance improvement for training large language models, but it is incremental in nature.
Mathematical optimization, control theory
Zhehang Du, Weijie Su
This work provides a theoretical foundation for an existing optimizer and offers a slight performance improvement for training large language models, but it is incremental in nature.
Zakhar Shumaylov, Nathaël Da Costa, Peter Zaika et al.
For optimization researchers, this work demystifies the success of Muon, suggesting that geometric narratives may be overemphasized, though the findings are incremental in nature.
Jianghao Lin, Zi Ling, Chenyu Zhou et al.
For practitioners needing reliable optimization modeling from natural language, Agora-Opt provides a practical and extensible foundation that improves over existing methods.
Xiaowen Jiang, Andrei Semenov, Sebastian U. Stich
This addresses training stability and generalization issues in large language models, representing an incremental improvement over existing methods.
Elad Hazan, Karan Singh · princeton
This addresses the problem of robust control under adversarial conditions for researchers and practitioners in control theory and reinforcement learning, representing a foundational shift rather than an incremental improvement.
Dingzhi Yu, Rui Pan, Yuxing Liu et al.
Provides a theoretically grounded, unbiased sign-based optimizer that enables stable low-precision training of large language models, a critical bottleneck for distributed and efficient LLM training.
Dechen Zhang, Xuan Tang, Xinxiang Yin et al.
This system addresses the challenge of automating complex mathematical research for ML theorists, offering a novel approach to developing and verifying theoretical results.
Shi Chen, Zhengjiang Lin, Kaizhao Liu et al.
Provides rigorous statistical guarantees for transformer performance as context length grows, addressing a key theoretical gap for practitioners scaling models.
Kang An, Jiaxiang Li, Donald Goldfarb et al.
For LLM practitioners, this work clarifies the fundamental role of manifold constraints, potentially simplifying training recipes by replacing multiple heuristics with a principled constrained optimization approach.
Carles Domingo-Enrich, Jiequn Han
For researchers working on reward fine-tuning of generative models and sampling from tilted distributions, this work provides a theoretically grounded and practical method for solving stochastic optimal control problems, though it is an incremental theoretical contribution that generalizes existing work.
Hassan Dbouk, Nidham Gazagnadou, Matthias Reisser et al.
For practitioners fine-tuning LLMs under memory constraints, this work provides a more efficient optimizer that reduces memory overhead without sacrificing performance.
Yuhan Chen, Tao Liu, Jie Huang
This solves a long-standing challenge in control theory for multi-agent systems, enabling applications like robotics and autonomous vehicles with unstable dynamics, though it is a foundational breakthrough rather than incremental.
Pierfrancesco Beneventano, Mahmoud Abdelmoneum, Tomaso Poggio
For deep learning practitioners, Muon offers a spectral bias that helps when many directions need to remain active, but it is not universally superior.
Anuj Apte, Pranav Deshpande, Niraj Kumar et al.
For practitioners training large language models, SF-NorMuon eliminates the need for costly learning-rate schedule tuning and horizon commitment, making anytime training practical.
Jihoon Hong, Alice Chan, Qiyue Dai et al.
For practitioners deploying text-to-video models, this provides a minimally invasive steering method that avoids oversteering and content degradation.
Jihoon Hong, Julian Skifstad, Qiyue Dai et al.
For researchers working on robust control with world models, this work provides a mechanistic understanding and practical steering method to improve robustness without retraining.
Ruicheng Ao, Hongyu Chen, Siyang Gao et al.
This addresses the challenge of efficient service system design for industries relying on textual data, offering a practical solution with significant cost savings, though it is incremental in combining existing techniques for a specific bottleneck.
Clément Hongler, Franck Gabriel, Valentin Hartmann et al.
This work addresses the open problem of developing general AI capabilities through automated curriculum learning, potentially impacting the entire field of artificial intelligence if successful.
Begoña García Malaxechebarría, Courtney Paquette, Maryam Fazel et al.
Provides a rigorous theoretical framework for understanding SGD dynamics in a simplified neural network setting, offering explicit non-asymptotic convergence guarantees.
Ian Osband
This addresses inefficiencies in policy gradient methods for reinforcement learning, offering a novel approach to improve training stability and performance, though it appears incremental as it builds on existing gradient techniques.