1.2SYMar 8, 2017
A hierarchical MPC scheme for interconnected systemsMarcello Farina, Xinglong Zhang, Riccardo Scattolini
This paper describes a hierarchical control scheme for interconnected systems. The higher layer of the control structure is designed with robust Model Predictive Control (MPC) based on a reduced order dynamic model of the overall system and is aimed at optimizing long-term performance, while at the lower layer local regulators acting at a higher frequency are designed for the full order models of the subsystems to refine the control action. A simulation experiment concerning the control of the temperature inside a building is reported to witness the potentialities of the proposed approach.
4.1LGNov 12, 2025
Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement LearningYunchang Ma, Tenglong Liu, Yixing Lan et al.
In offline reinforcement learning, value overestimation caused by out-of-distribution (OOD) actions significantly limits policy performance. Recently, diffusion models have been leveraged for their strong distribution-matching capabilities, enforcing conservatism through behavior policy constraints. However, existing methods often apply indiscriminate regularization to redundant actions in low-quality datasets, resulting in excessive conservatism and an imbalance between the expressiveness and efficiency of diffusion modeling. To address these issues, we propose DIffusion policies with Value-conditional Optimization (DIVO), a novel approach that leverages diffusion models to generate high-quality, broadly covered in-distribution state-action samples while facilitating efficient policy improvement. Specifically, DIVO introduces a binary-weighted mechanism that utilizes the advantage values of actions in the offline dataset to guide diffusion model training. This enables a more precise alignment with the dataset's distribution while selectively expanding the boundaries of high-advantage actions. During policy improvement, DIVO dynamically filters high-return-potential actions from the diffusion model, effectively guiding the learned policy toward better performance. This approach achieves a critical balance between conservatism and explorability in offline RL. We evaluate DIVO on the D4RL benchmark and compare it against state-of-the-art baselines. Empirical results demonstrate that DIVO achieves superior performance, delivering significant improvements in average returns across locomotion tasks and outperforming existing methods in the challenging AntMaze domain, where sparse rewards pose a major difficulty.
4.3SYNov 22, 2019
Robust Learning-based Predictive Control for Discrete-time Nonlinear Systems with Unknown Dynamics and State ConstraintsXinglong Zhang, Jiahang Liu, Xin Xu et al.
Robust model predictive control (MPC) is a well-known control technique for model-based control with constraints and uncertainties. In classic robust tube-based MPC approaches, an open-loop control sequence is computed via periodically solving an online nominal MPC problem, which requires prior model information and frequent access to onboard computational resources. In this paper, we propose an efficient robust MPC solution based on receding horizon reinforcement learning, called r-LPC, for unknown nonlinear systems with state constraints and disturbances. The proposed r-LPC utilizes a Koopman operator-based prediction model obtained off-line from pre-collected input-output datasets. Unlike classic tube-based MPC, in each prediction time interval of r-LPC, we use an actor-critic structure to learn a near-optimal feedback control policy rather than a control sequence. The resulting closed-loop control policy can be learned off-line and deployed online or learned online in an asynchronous way. In the latter case, online learning can be activated whenever necessary; for instance, the safety constraint is violated with the deployed policy. The closed-loop recursive feasibility, robustness, and asymptotic stability are proven under function approximation errors of the actor-critic networks. Simulation and experimental results on two nonlinear systems with unknown dynamics and disturbances have demonstrated that our approach has better or comparable performance when compared with tube-based MPC and LQR, and outperforms a recently developed actor-critic learning approach.
1.2SYMay 24, 2017
A hierarchical multirate MPC scheme for interconnected systems - $\textit{extended version}$Marcello Farina, Xinglong Zhang, Riccardo Scattolini
This paper presents a hierarchical control scheme for interconnected linear systems. At the higher layer of the control structure a robust centralized Model Predictive Control (MPC) algorithm based on a reduced order dynamic model of the overall system optimizes a long-term performance index penalizing the deviation of the state and the control input from their nominal values. At the lower layer local MPC regulators, possibly working at different rates, are designed for the full order models of the subsystems to refine the control action computed at the higher layer. A simulation experiment is presented to describe the implementation aspects and the potentialities of the proposed approach.