2024

VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision

Xu, Yi, Hu, Yuxin, Zhang, Zaiwei et al.

Understand

Human drivers rely on commonsense reasoning to navigate diverse and dynamic real-world scenarios.

  • Existing end-to-end (E2E) autonomous driving (AD) models are typically optimized to mimic driving patterns observed in data, without capturing the underlying reasoning processes.
  • This limitation constrains their ability to handle challenging driving scenarios.
  • To close this gap, we propose VLM-AD, a method that leverages vision-language models (VLMs) as teachers to enhance training by providing additional supervision that incorporates unstructured reasoning information and structured action labels.

Reading the bibliography…