LG RO SPSep 29, 2021

Formulation and validation of a car-following model based on deep reinforcement learning

Fabian Hart, Ostap Okhrin, Martin Treiber

arXiv:2109.14268v16.530 citations

Originality Incremental advance

AI Analysis

This addresses car-following modeling for traffic simulation and autonomous driving, but it is incremental as it builds on existing deep reinforcement learning and traditional models like IDM.

The authors tackled the problem of car-following in traffic by developing a deep reinforcement learning model that maximizes reward functions for driving regimes, resulting in unconditional string stability, comfort, and crash-free performance across various scenarios, with higher reward and better goodness-of-fit compared to a traditional model.

We propose and validate a novel car following model based on deep reinforcement learning. Our model is trained to maximize externally given reward functions for the free and car-following regimes rather than reproducing existing follower trajectories. The parameters of these reward functions such as desired speed, time gap, or accelerations resemble that of traditional models such as the Intelligent Driver Model (IDM) and allow for explicitly implementing different driving styles. Moreover, they partially lift the black-box nature of conventional neural network models. The model is trained on leading speed profiles governed by a truncated Ornstein-Uhlenbeck process reflecting a realistic leader's kinematics. This allows for arbitrary driving situations and an infinite supply of training data. For various parameterizations of the reward functions, and for a wide variety of artificial and real leader data, the model turned out to be unconditionally string stable, comfortable, and crash-free. String stability has been tested with a platoon of five followers following an artificial and a real leading trajectory. A cross-comparison with the IDM calibrated to the goodness-of-fit of the relative gaps showed a higher reward compared to the traditional model and a better goodness-of-fit.

View on arXiv PDF

Similar