Fetching the paper…
Reading the bibliography…
The pursuit of autonomous agents capable of temporally coherent planning is hindered by a fundamental flaw in current vision-language models (VLMs): they lack cognitive inertia.
CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving
Arai, H.; Miwa, K.; Sasaki, K.; Watanabe, K.; Yamaguchi, Y.; Aoki, S.; and Yamamoto, I. 2025 · 1943
Earlier work this paper cites.
SGDR: Stochastic Gradient Descent with Warm Restarts
Loshchilov, I.; and Hutter, F. 2016 · 2016
Earlier work this paper cites.
Textual Explanations for Self-Driving Vehicles
Kim, J.; Rohrbach, A.; Darrell, T.; Canny, J.; and Akata, Z. 2018 · 2018
Earlier work this paper cites.
Grounding Human-To-Vehicle Advice for Self-Driving Vehicles
Kim, J.; Misu, T.; Chen, Y.-T.; Tawari, A.; and Canny, J. 2019 · 2019
Earlier work this paper cites.
nuScenes: A Multimodal Dataset for Autonomous Driving
Caesar, H.; Bankiti, V.; Lang, A. H.; Vora, S.; Liong, V. E.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; and Beijbom, O. 2020 · 2020
Earlier work this paper cites.
Explainable Object-Induced Action Decision for Autonomous Vehicles
Xu, Y.; Yang, X.; Gong, L.; Lin, H.-C.; Wu, T.-Y.; Li, Y.; and Vasconcelos, N. 2020 · 2020
Earlier work this paper cites.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Talk2Car: Predicting Physical Trajectories for Natural Language Commands
Deruyttere, T.; Grujicic, D.; Blaschko, M. B.; and Moens, M.-F. 2022 · 2022
Earlier work this paper cites.
ST-P3: End-to-End Vision-Based Autonomous Driving via Spatial-Temporal Feature Learning
Hu, S.; Chen, L.; Wu, P.; Li, H.; Yan, J.; and Tao, D. 2022 · 2022
Earlier work this paper cites.
PETR: Position Embedding Transformation for Multi-view 3D Object Detection
Liu, Y.; Wang, T.; Zhang, X.; and Sun, J. 2022 · 2022
Earlier work this paper cites.
DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries
Wang, Y.; Guizilini, V. C.; Zhang, T.; Wang, Y.; Zhao, H.; and Solomon, J. 2022 · 2022
Earlier work this paper cites.
VAD: Vectorized Scene Representation for Efficient Autonomous Driving
Jiang, B.; Chen, S.; Xu, Q.; Liao, B.; Chen, J.; Zhou, H.; Zhang, Q.; Liu, W.; Huang, C.; and Wang, X. 2023 · 2023
Earlier work this paper cites.
Open-sourced Data Ecosystem in Autonomous Driving: the Present and Future
Li, H.; Li, Y.; Wang, H.; Zeng, J.; Xu, H.; Cai, P.; Chen, L.; Yan, J.; Xu, F.; Xiong, L.; et al. 2023 · 2023
Earlier work this paper cites.
DRAMA: Joint Risk Localization and Captioning in Driving
Malla, S.; Choi, C.; Dwivedi, I.; Choi, J. H.; and Li, J. 2023 · 2023
Earlier work this paper cites.
LingoQA: Visual Question Answering for Autonomous Driving
Marcu, A.; Chen, L.; Hünermann, J.; Karnsund, A.; Hanotte, B.; Chidananda, P.; Nair, S.; Badrinarayanan, V.; Kendall, A.; Shotton, J.; et al. 2023 · 2023
Earlier work this paper cites.
MotionLM: Multi-Agent Motion Forecasting as Language Modeling
Seff, A.; Cera, B.; Chen, D.; Ng, M.; Zhou, A.; Nayakanti, N.; Refaat, K. S.; Al-Rfou, R.; and Sapp, B. 2023 · 2023
Cited alongside, same era.
Sigmoid Loss for Language Image Pre-Training
Zhai, X.; Mustafa, B.; Kolesnikov, A.; and Beyer, L. 2023 · 2023
Cited alongside, same era.
Driving with LLMs: Fusing Object-Level Vector Modality for Explainable Autonomous Driving
Chen, L.; Sinavski, O.; Hünermann, J.; Karnsund, A.; Willmott, A. J.; Birch, D.; Maund, D.; and Shotton, J. 2024a · 2024
Cited alongside, same era.
EVA-02: A visual representation for neon genesis
Fang, Y.; Sun, Q.; Wang, X.; Huang, T.; Wang, X.; and Cao, Y. 2024 · 2024
Cited alongside, same era.
Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving
Jia, X.; Yang, Z.; Li, Q.; Zhang, Z.; and Yan, J. 2024 · 2024
Cited alongside, same era.
Hallucination is Inevitable: An Innate Limitation of Large Language Models
Xu, Z.; Jain, S.; and Kankanhalli, M. 2024 · 2024
Later among the works it cites.
DriveGPT4: Interpretable End-to-End Autonomous Driving Via Large Language Model
Xu, Z.; Zhang, Y.; Xie, E.; Zhao, Z.; Guo, Y.; Wong, K.-Y. K.; Li, Z.; and Zhao, H. 2024 · 2024
Later among the works it cites.
Generalized Predictive Model for Autonomous Driving
Yang, J.; Gao, S.; Qiu, Y.; Chen, L.; Li, T.; Dai, B.; Chitta, K.; Wu, P.; Zeng, J.; Luo, P.; et al. 2024 · 2024
Later among the works it cites.
Yuan, J.; Sun, S.; Omeiza, D.; Zhao, B.; Newman, P.; Kunze, L.; and Gadd, M. 2024 · 2024
Later among the works it cites.
GPT-4V as Traffic Assistant: An In-depth Look at Vision Language Model on Complex Traffic Events
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiang, B.; Chen, S.; Liao, B.; Zhang, X.; Yin, W.; Zhang, Q.; Huang, C.; Liu, W.; and Wang, X. 2024 · 2024
Cited alongside, same era.
Reason2Drive: Towards Interpretable and Chain-Based Reasoning for Autonomous Driving
Nie, M.; Peng, R.; Wang, C.; Cai, X.; Han, J.; Xu, H.; and Zhang, L. 2024 · 2024
Cited alongside, same era.
VLP: Vision Language Planning for Autonomous Driving
Pan, C.; Yaman, B.; Nesti, T.; Mallik, A.; Allievi, A. G.; Velipasalar, S.; and Ren, L. 2024 · 2024
Cited alongside, same era.
CarLLaVA: Vision language models for camera-only closed-loop driving
Renz, K.; Chen, L.; Marcu, A.-M.; Hünermann, J.; Hanotte, B.; Karnsund, A.; Shotton, J.; Arani, E.; and Sinavski, O. 2024 · 2024
Cited alongside, same era.
Rank2Tell: A Multimodal Driving Dataset for Joint Importance Ranking and Reasoning
Sachdeva, E.; Agarwal, N.; Chundi, S.; Roelofs, S.; Li, J.; Kochenderfer, M.; Choi, C.; and Dariush, B. 2024 · 2024
Cited alongside, same era.
LMDrive: Closed-Loop End-to-End Driving with Large Language Models
Shao, H.; Hu, Y.; Wang, L.; Song, G.; Waslander, S. L.; Liu, Y.; and Li, H. 2024 · 2024
Cited alongside, same era.
DriveLM: Driving with Graph Visual Question Answering
Sima, C.; Renz, K.; Chitta, K.; Chen, L.; Zhang, H.; Xie, C.; Beißwenger, J.; Luo, P.; Geiger, A.; and Li, H. 2024 · 2024
Cited alongside, same era.
Zhou, X.; and Knoll, A. C. 2024 · 2024
Later among the works it cites.
Vision Language Models in Autonomous Driving: A Survey and Outlook
Zhou, X.; Liu, M.; Yurtsever, E.; Zagar, B. L.; Zimmer, W.; Cao, H.; and Knoll, A. C. 2024 · 2024
Later among the works it cites.
Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; et al. 2025 · 2025
Closest in time.
Fu, H.; Zhang, D.; Zhao, Z.; Cui, J.; Liang, D.; Zhang, C.; Zhang, D.; Xie, H.; Wang, B.; and Bai, X. 2025 · 2025
Closest in time.
DME-Driver: Integrating Human Decision Logic and 3D Scene Perception in Autonomous Driving
Han, W.; Guo, D.; Xu, C.-Z.; and Shen, J. 2025 · 2025
Closest in time.
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
Wang, S.; Yu, Z.; Jiang, X.; Lan, S.; Shi, M.; Chang, N.; Kautz, J.; Li, Y.; and Alvarez, J. M. 2025 · 2025
Closest in time.
Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Zhang, Y.; Li, Y.; Cui, L.; Cai, D.; Liu, L.; Fu, T.; Huang, X.; Zhao, E.; Zhang, Y.; Chen, Y.; et al. 2025 · 2025
Closest in time.
OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
Zhou, X.; Han, X.; Yang, F.; Ma, Y.; and Knoll, A. C. 2025 · 2025
Closest in time.
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Zhu, J.; Wang, W.; Chen, Z.; Liu, Z.; Ye, S.; Gu, L.; Tian, H.; Duan, Y.; Su, W.; Shao, J.; et al. 2025 · 2025
Closest in time.
Talk2Car: Taking Control of Your Self-Driving Car
Deruyttere, T.; Vandenhende, S.; Grujicic, D.; Van Gool, L.; and Moens, M. F. 2019 · 2098
Closest in time.