Lin Zhao

RO
h-index20
4papers
46citations
Novelty54%
AI Score26

4 Papers

10.1OCOct 29, 2023
Optimization Landscape of Policy Gradient Methods for Discrete-time Static Output Feedback

Jingliang Duan, Jie Li, Xuyang Chen et al.

In recent times, significant advancements have been made in delving into the optimization landscape of policy gradient methods for achieving optimal control in linear time-invariant (LTI) systems. Compared with state-feedback control, output-feedback control is more prevalent since the underlying state of the system may not be fully observed in many practical settings. This paper analyzes the optimization landscape inherent to policy gradient methods when applied to static output feedback (SOF) control in discrete-time LTI systems subject to quadratic cost. We begin by establishing crucial properties of the SOF cost, encompassing coercivity, L-smoothness, and M-Lipschitz continuous Hessian. Despite the absence of convexity, we leverage these properties to derive novel findings regarding convergence (and nearly dimension-free rate) to stationary points for three policy gradient methods, including the vanilla policy gradient method, the natural policy gradient method, and the Gauss-Newton method. Moreover, we provide proof that the vanilla policy gradient method exhibits linear convergence towards local minima when initialized near such minima. The paper concludes by presenting numerical examples that validate our theoretical findings. These results not only characterize the performance of gradient descent for optimizing the SOF problem but also provide insights into the effectiveness of general policy gradient methods within the realm of reinforcement learning.

12.8ROSep 12, 2021
Encoding Distributional Soft Actor-Critic for Autonomous Driving in Multi-lane Scenarios

Jingliang Duan, Yangang Ren, Fawang Zhang et al.

In this paper, we propose a new reinforcement learning (RL) algorithm, called encoding distributional soft actor-critic (E-DSAC), for decision-making in autonomous driving. Unlike existing RL-based decision-making methods, E-DSAC is suitable for situations where the number of surrounding vehicles is variable and eliminates the requirement for manually pre-designed sorting rules, resulting in higher policy performance and generality. We first develop an encoding distributional policy iteration (DPI) framework by embedding a permutation invariant module, which employs a feature neural network (NN) to encode the indicators of each vehicle, in the distributional RL framework. The proposed DPI framework is proved to exhibit important properties in terms of convergence and global optimality. Next, based on the developed encoding DPI framework, we propose the E-DSAC algorithm by adding the gradient-based update rule of the feature NN to the policy evaluation process of the DSAC algorithm. Then, the multi-lane driving task and the corresponding reward function are designed to verify the effectiveness of the proposed algorithm. Results show that the policy learned by E-DSAC can realize efficient, smooth, and relatively safe autonomous driving in the designed scenario. And the final policy performance learned by E-DSAC is about three times that of DSAC. Furthermore, its effectiveness has also been verified in real vehicle experiments.

7.3ROAug 6, 2021
Differentiable Moving Horizon Estimation for Robust Flight Control

Bingheng Wang, Zhengtian Ma, Shupeng Lai et al.

Estimating and reacting to external disturbances is of fundamental importance for robust control of quadrotors. Existing estimators typically require significant tuning or training with a large amount of data, including the ground truth, to achieve satisfactory performance. This paper proposes a data-efficient differentiable moving horizon estimation (DMHE) algorithm that can automatically tune the MHE parameters online and also adapt to different scenarios. We achieve this by deriving the analytical gradient of the estimated trajectory from MHE with respect to the tuning parameters, enabling end-to-end learning for auto-tuning. Most interestingly, we show that the gradient can be calculated efficiently from a Kalman filter in a recursive form. Moreover, we develop a model-based policy gradient algorithm to learn the parameters directly from the trajectory tracking errors without the need for the ground truth. The proposed DMHE can be further embedded as a layer with other neural networks for joint optimization. Finally, we demonstrate the effectiveness of the proposed method via both simulation and experiments on quadrotors, where challenging scenarios such as sudden payload change and flying in downwash are examined.

1.2SYAug 18, 2017
A Unified Stochastic Hybrid System Approach to Aggregate Modeling of Responsive Loads

Lin Zhao, Wei Zhang

Aggregate load modeling is of fundamental importance for systematic analysis and design of various demand response strategies. Instead of keeping track of the trajectories of individual loads, the aggregate modeling problem focuses on characterizing the density evolution of the load population. Most existing models are only applicable to Thermostatically Controlled Loads (TCL) with first-order linear dynamics. This paper develops a unified aggregate modeling approach that can be used for general TCLs as well as deferrable loads. We propose a deterministic hybrid system model to describe individual load dynamics under demand response rules, and develop a general stochastic hybrid system (SHS) model to capture the population dynamics. We also derive a set of partial differential equations (PDE) that governs the probability density evolution of the SHS. Our results cannot be obtained using the exiting SHS tools in the literature as the proposed SHS model involves both random and deterministic switchings with general switching surfaces in multi-dimensional domains. The derived PDE model includes many existing aggregate load modeling results as special cases and can be used in many other realistic modeling scenarios that have not been studied in the literature.