Agent-driven Long-tail Simulation for Autonomous Driving
For autonomous driving researchers, this work provides a more realistic and challenging simulation benchmark to evaluate closed-loop planning in long-tail scenarios.
The paper proposes an agent-driven simulation framework using LLMs to generate diverse, long-tail traffic scenarios for autonomous driving evaluation, and introduces the SemanticPlan benchmark. Results show that current planners struggle with these scenarios, highlighting remaining challenges.
Evaluating autonomous driving systems in closed-loop settings requires realistic and interactive simulation, yet existing simulators largely rely on log replay or rule-based agents, limiting behavioral diversity and long-tail coverage. We propose an agent-driven simulation framework in which surrounding road participants are controlled by instruction-following large language models through a structured action interface, enabling intentional and reactive behaviors while preserving physical plausibility. Furthermore, we introduce SemanticPlan, a benchmark of closed-loop planning in long-tail and semantically rich scenarios that augment real nuPlan scenes with multiple interactive agents following diverse language instructions. Evaluation results show that state-of-the-art planners still struggle to consistently achieve safe and effective task completion, suggesting that these long-tail scenarios remain challenging.