4.1LGNov 12, 2025
On the Convergence of Overparameterized Problems: Inherent Properties of the Compositional Structure of Neural NetworksArthur Castello Branco de Oliveira, Dhruv Jatkar, Eduardo Sontag
This paper investigates how the compositional structure of neural networks shapes their optimization landscape and training dynamics. We analyze the gradient flow associated with overparameterized optimization problems, which can be interpreted as training a neural network with linear activations. Remarkably, we show that the global convergence properties can be derived for any cost function that is proper and real analytic. We then specialize the analysis to scalar-valued cost functions, where the geometry of the landscape can be fully characterized. In this setting, we demonstrate that key structural features -- such as the location and stability of saddle points -- are universal across all admissible costs, depending solely on the overparameterized representation rather than on problem-specific details. Moreover, we show that convergence can be arbitrarily accelerated depending on the initialization, as measured by an imbalance metric introduced in this work. Finally, we discuss how these insights may generalize to neural networks with sigmoidal activations, showing through a simple example which geometric and dynamical properties persist beyond the linear case.
4.1OCMar 31, 2025
Remarks on the Polyak-Lojasiewicz inequality and the convergence of gradient systemsArthur Castello B. de Oliveira, Leilei Cui, Eduardo D. Sontag
This work explores generalizations of the Polyak-Lojasiewicz inequality (PLI) and their implications for the convergence behavior of gradient flows in optimization problems. Motivated by the continuous-time linear quadratic regulator (CT-LQR) policy optimization problem -- where only a weaker version of the PLI is characterized in the literature -- this work shows that while weaker conditions are sufficient for global convergence to, and optimality of the set of critical points of the cost function, the "profile" of the gradient flow solution can change significantly depending on which "flavor" of inequality the cost satisfies. After a general theoretical analysis, we focus on fitting the CT-LQR policy optimization problem to the proposed framework, showing that, in fact, it can never satisfy a PLI in its strongest form. We follow up our analysis with a brief discussion on the difference between continuous- and discrete-time LQR policy optimization, and end the paper with some intuition on the extension of this framework to optimization problems with L1 regularization and solved through proximal gradient flows.
3.8LGMay 17, 2023
On the ISS Property of the Gradient Flow for Single Hidden-Layer Neural Networks with Linear ActivationsArthur Castello B. de Oliveira, Milad Siami, Eduardo D. Sontag
Recent research in neural networks and machine learning suggests that using many more parameters than strictly required by the initial complexity of a regression problem can result in more accurate or faster-converging models -- contrary to classical statistical belief. This phenomenon, sometimes known as ``benign overfitting'', raises questions regarding in what other ways might overparameterization affect the properties of a learning problem. In this work, we investigate the effects of overfitting on the robustness of gradient-descent training when subject to uncertainty on the gradient estimation. This uncertainty arises naturally if the gradient is estimated from noisy data or directly measured. Our object of study is a linear neural network with a single, arbitrarily wide, hidden layer and an arbitrary number of inputs and outputs. In this paper we solve the problem for the case where the input and output of our neural-network are one-dimensional, deriving sufficient conditions for robustness of our system based on necessary and sufficient conditions for convergence in the undisturbed case. We then show that the general overparametrized formulation introduces a set of spurious equilibria which lay outside the set where the loss function is minimized, and discuss directions of future work that might extend our current results for more general formulations.
9.4ROJun 6, 2020
Thruster-assisted center manifold shaping in bipedal legged locomotionArthur C. B. de Oliveira, Alireza Ramezani
This work tries to contribute to the design of legged robots with capabilities boosted through thruster-assisted locomotion. Our long-term goal is the development of robots capable of negotiating unstructured environments, including land and air, by leveraging legs and thrusters collaboratively. These robots could be used in a broad number of applications including search and rescue operations, space exploration, automated package handling in residential spaces and digital agriculture, to name a few. In all of these examples, the unique capability of thruster-assisted mobility greatly broadens the locomotion designs possibilities for these systems. In an effort to demonstrate thrusters effectiveness in the robustification and efficiency of bipedal locomotion gaits, this work explores their effects on the gait limit cycles and proposes new design paradigms based on shaping these center manifolds with strong foliations. Unilateral contact force feasibility conditions are resolved in an optimal control scheme.