UNCAP: Uncertainty-Guided Planning Using Natural Language Communication for Cooperative Autonomous Vehicles
Neel P. Bhatt, Po-han Li, Kushagra Gupta, Rohan Siva, Daniel Milan, Alexander T. Hogue, Sandeep P. Chinchali, David Fridovich-Keil, Zhangyang Wang, Ufuk Topcu
TL;DR
UNCAP tackles scalable, safe cooperative autonomous driving by enabling inter-vehicle communication through lightweight natural language while explicitly modeling perception and planning uncertainties. It introduces a two-stage protocol, BARE and SPARE, to selectively exchange information and uses uncertainty quantification plus mutual information to fuse signals for VLM-based planning, yielding interpretable plans with associated uncertainty scores. Across CARLA-based OPV2V scenarios, UNCAP achieves substantial bandwidth reductions (up to 63% relative to baselines), improved driving quality (≈31% higher DS), reduced planning uncertainty (≈61%), and markedly larger collision-distance margins in near-miss events, with robustness across different VLMs. The approach demonstrates that uncertainty-aware, language-based coordination can scale to larger CAV fleets with real-time performance and safety guarantees, offering a practical path toward interpretable, low-bandwidth cooperative autonomous driving.
Abstract
Safe large-scale coordination of multiple cooperative connected autonomous vehicles (CAVs) hinges on communication that is both efficient and interpretable. Existing approaches either rely on transmitting high-bandwidth raw sensor data streams or neglect perception and planning uncertainties inherent in shared data, resulting in systems that are neither scalable nor safe. To address these limitations, we propose Uncertainty-Guided Natural Language Cooperative Autonomous Planning (UNCAP), a vision-language model-based planning approach that enables CAVs to communicate via lightweight natural language messages while explicitly accounting for perception uncertainty in decision-making. UNCAP features a two-stage communication protocol: (i) an ego CAV first identifies the subset of vehicles most relevant for information exchange, and (ii) the selected CAVs then transmit messages that quantitatively express their perception uncertainty. By selectively fusing messages that maximize mutual information, this strategy allows the ego vehicle to integrate only the most relevant signals into its decision-making, improving both the scalability and reliability of cooperative planning. Experiments across diverse driving scenarios show a 63% reduction in communication bandwidth with a 31% increase in driving safety score, a 61% reduction in decision uncertainty, and a four-fold increase in collision distance margin during near-miss events. Project website: https://uncap-project.github.io/
