10.4RONov 11, 2020
Offline Learning of Counterfactual Predictions for Real-World Robotic Reinforcement LearningJun Jin, Daniel Graves, Cameron Haigh et al.
We consider real-world reinforcement learning (RL) of robotic manipulation tasks that involve both visuomotor skills and contact-rich skills. We aim to train a policy that maps multimodal sensory observations (vision and force) to a manipulator's joint velocities under practical considerations. We propose to use offline samples to learn a set of general value functions (GVFs) that make counterfactual predictions from the visual inputs. We show that combining the offline learned counterfactual predictions with force feedbacks in online policy learning allows efficient reinforcement learning given only a terminal (success/failure) reward. We argue that the learned counterfactual predictions form a compact and informative representation that enables sample efficiency and provides auxiliary reward signals that guide online explorations towards contact-rich states. Various experiments in simulation and real-world settings were performed for evaluation. Recordings of the real-world robot training can be found via https://sites.google.com/view/realrl.
9.4ROJan 24, 2020
Perception as prediction using general value functions in autonomous driving applicationsDaniel Graves, Kasra Rezaee, Sean Scheideman
We propose and demonstrate a framework called perception as prediction for autonomous driving that uses general value functions (GVFs) to learn predictions. Perception as prediction learns data-driven predictions relating to the impact of actions on the agent's perception of the world. It also provides a data-driven approach to predict the impact of the anticipated behavior of other agents on the world without explicitly learning their policy or intentions. We demonstrate perception as prediction by learning to predict an agent's front safety and rear safety with GVFs, which encapsulate anticipation of the behavior of the vehicle in front and in the rear, respectively. The safety predictions are learned through random interactions in a simulated environment containing other agents. We show that these predictions can be used to produce similar control behavior to an LQR-based controller in an adaptive cruise control problem as well as provide advanced warning when the vehicle behind is approaching dangerously. The predictions are compact policy-based predictions that support prediction of the long term impact on safety when following a given policy. We analyze two controllers that use the learned predictions in a racing simulator to understand the value of the predictions and demonstrate their use in the real-world on a Clearpath Jackal robot and an autonomous vehicle platform.
1.2NANov 16, 2014
A higher-order finite-volume discretization method for Poisson's equation in cut cell geometriesD. Devendran, D. T. Graves, H. Johansen
We present a method for generating higher-order finite volume discretizations for Poisson's equation on Cartesian cut cell grids in two and three dimensions. The discretization is in flux-divergence form, and stencils for the flux are computed by solving small weighted least-squares linear systems. Weights are the key in generating a stable discretization. We apply the method to solve Poisson's equation on a variety of geometries, and we demonstrate that the method can achieve second and fourth order accuracy in both truncation and solution error for these examples. We also show that the Laplacian operator has only stable eigenvalues for each of these examples.