Piyush Grover

h-index11
2papers
567citations

2 Papers

1.2SYJul 26, 2018
Optimal Transport over Deterministic Discrete-time Nonlinear Systems using Stochastic Feedback Laws

Karthik Elamvazhuthi, Piyush Grover, Spring Berman

This paper considers the relaxed version of the transport problem for general nonlinear control systems, where the objective is to design time-varying feedback laws that transport a given initial probability measure to a target probability measure under the action of the closed-loop system. To make the problem analytically tractable, we consider control laws that are stochastic, i.e., the control laws are maps from the state space of the control system to the space of probability measures on the set of admissible control inputs. Under some controllability assumptions on the control system as defined on the state space, we show that the transport problem, considered as a controllability problem for the lifted control system on the space of probability measures, is well-posed for a large class of initial and target measures. We use this to prove the well-posedness of a fixed-endpoint optimal control problem defined on the space of probability measures, where along with the terminal constraints, the goal is to optimize an objective functional along the trajectory of the control system. This optimization problem can be posed as an infinite-dimensional linear programming problem. This formulation facilitates numerical solutions of the transport problem for low-dimensional control systems, as we show in two numerical examples.

7.1LGJun 13, 2018
Reinforcement Learning with Function-Valued Action Spaces for Partial Differential Equation Control

Yangchen Pan, Amir-massoud Farahmand, Martha White et al.

Recent work has shown that reinforcement learning (RL) is a promising approach to control dynamical systems described by partial differential equations (PDE). This paper shows how to use RL to tackle more general PDE control problems that have continuous high-dimensional action spaces with spatial relationship among action dimensions. In particular, we propose the concept of action descriptors, which encode regularities among spatially-extended action dimensions and enable the agent to control high-dimensional action PDEs. We provide theoretical evidence suggesting that this approach can be more sample efficient compared to a conventional approach that treats each action dimension separately and does not explicitly exploit the spatial regularity of the action space. The action descriptor approach is then used within the deep deterministic policy gradient algorithm. Experiments on two PDE control problems, with up to 256-dimensional continuous actions, show the advantage of the proposed approach over the conventional one.