Fetching the paper…
Reading the bibliography…
End-to-end autonomous driving systems built on Vision Language Models (VLMs) have shown significant promise, yet their reliance on autoregressive architectures introduces some limitations for real-world applications.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in International conference on machine learning . pmlr, 2015, pp. 2256–2265
2015
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuScenes: A Multimodal Dataset for Autonomous Driving,” in CVPR , 2020, pp. 11 621–11 631
2020
Earlier work this paper cites.
J. Austin, D. D. Johnson, J. Ho, D. Tarlow, and R. Van Den Berg, “Structured denoising diffusion models in discrete state-spaces,” Advances in neural information processing systems , vol. 34, pp. 17 981–17 993, 2021
2021
Earlier work this paper cites.
P. Wu, X. Jia, L. Chen, J. Yan, H. Li, and Y. Qiao, “Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong baseline,” Advances in Neural Information Processing Systems , vol. 35, pp. 6119–6132, 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
OpenAI, “GPT-4 Technical Report,” Mar. 2023, arXiv:2303.08774 [cs.CL]
2023
Cited alongside, same era.
2023
Cited alongside, same era.
X. Tian, J. Gu, B. Li, Y. Liu, C. Hu, Y. Wang, K. Zhan, P. Jia, X. Lang, and H. Zhao, “DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models,” arXiv , 2024
2024
Cited alongside, same era.
C. Cui, Y. Ma, X. Cao, W. Ye, and Z. Wang, “Drive as You Speak: Enabling Human-Like Interaction with Large Language Models in Autonomous Vehicles,” in 2024 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW) . Waikoloa, HI, USA: IEEE, Jan. 2024, pp. 902–909. [Online]. Available: https://ieeexplore.ieee.org/document/10495655/
2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Z. Xu, Y. Zhang, E. Xie, Z. Zhao, Y. Guo, K.-Y. K. Wong, Z. Li, and H. Zhao, “DriveGPT4: Interpretable End-to-End Autonomous Driving Via Large Language Model,” IEEE Robotics and Automation Letters , vol. 9, no. 10, pp. 8186–8193, Oct. 2024, conference Name: IEEE Robotics and Automation Letters. [Online]. Available: https://ieeexplore.ieee.org/document/10629039/?arnumber=10629039
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2025
Closest in time.
S. Xing, C. Qian, Y. Wang, H. Hua, K. Tian, Y. Zhou, and Z. Tu, “Openemma: Open-source multimodal model for end-to-end autonomous driving,” in Proceedings of the Winter Conference on Applications of Computer Vision , 2025, pp. 1001–1009
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.