Multi-Agent Embodied Autonomous Driving: From V2X Information Exchange to Shared World Models
For researchers and practitioners in autonomous driving, this survey organizes the fragmented literature on multi-agent coordination and highlights critical open problems for safe deployment.
This survey reviews over 380 publications on multi-agent embodied autonomous driving, focusing on Shared World Models (SWMs) for collaborative perception, intent inference, and coordinated action. It identifies key gaps: evaluation is mostly in simulation, and foundation-model-based coordination lacks real-time safety guarantees.
Autonomous driving is shifting from isolated vehicle intelligence toward multi-agent embodied systems that share perception, infer intent, and coordinate action under uncertainty. This survey examines this transition through the lens of Shared World Models (SWMs): predictive cross-agent representations maintained across vehicles, infrastructure, and other traffic participants. We review more than 380 publications spanning vehicle-to-everything (V2X) communication, collaborative perception, inter-agent cognition, cooperative planning, end-to-end cooperative driving, and simulation and data engines for closed-loop validation. The organizing question is how exchanged observations become aligned state, intent-aware interaction, and coordinated downstream action. Across the surveyed literature, evaluation remains concentrated in simulation, curated benchmarks, and offline protocols. Foundation-model-based coordination also lacks verified real-time safety guarantees in open traffic. These gaps motivate key research priorities for multi-agent embodied autonomous driving (MAEAD): verifiable shared-state maintenance, robust intent and plan alignment, and safe coordinated action under communication, latency, and deployment constraints.