Fetching the paper…
Reading the bibliography…
Due to the powerful vision-language reasoning and generalization abilities, multimodal large language models (MLLMs) have garnered significant attention in the field of end-to-end (E2E) autonomous driving.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Carla: An open urban driving simulator
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun · 2017
Earlier work this paper cites.
An environment for autonomous driving decision-making
E. Leurent · 2018
Earlier work this paper cites.
Causal confusion in imitation learning
P. De Haan, D. Jayaraman, and S. Levine · 2019
Earlier work this paper cites.
Multi-modal fusion transformer for end-to-end autonomous driving
A. Prakash, K. Chitta, and A. Geiger · 2021
Earlier work this paper cites.
Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong baseline
P. Wu, X. Jia, L. Chen, J. Yan, H. Li, and Y. Qiao · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
Planning-oriented autonomous driving
Y. Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wang, et al · 2023
Earlier work this paper cites.
Vad: Vectorized scene representation for efficient autonomous driving
B. Jiang, S. Chen, Q. Xu, B. Liao, J. Chen, H. Zhou, Q. Zhang, W. Liu, C. Huang, and X. Wang · 2023
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
W. Wang, J. Xie, C. Hu, H. Zou, J. Fan, W. Tong, Y. Wen, S. Wu, H. Deng, Z. Li, et al · 2023
Earlier work this paper cites.
Reasonnet: End-to-end driving with temporal and global reasoning
H. Shao, L. Wang, R. Chen, S. L. Waslander, H. Li, and Y. Liu · 2023
Earlier work this paper cites.
Rethinking the open-loop evaluation of end-to-end autonomous driving in nuscenes
J.-T. Zhai, Z. Feng, J. Du, Y. Mao, J.-J. Liu, Z. Tan, Y. Zhang, X. Ye, and J. Wang · 2023
Earlier work this paper cites.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Earlier work this paper cites.
Vadv2: End-to-end vectorized autonomous driving via probabilistic planning
S. Chen, B. Jiang, H. Gao, B. Liao, Q. Xu, Q. Zhang, C. Huang, W. Liu, and X. Wang · 2024
Earlier work this paper cites.
Sparsedrive: End-to-end autonomous driving via sparse scene representation
W. Sun, X. Lin, Y. Shi, C. Zhang, H. Wu, and S. Zheng · 2024
Cited alongside, same era.
Llava-onevision: Easy visual task transfer
B. Li, Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y. Li, Z. Liu, et al · 2024
Cited alongside, same era.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, et al · 2024
Cited alongside, same era.
Deepseek-vl: Towards real-world vision-language understanding
H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang, et al · 2024
Cited alongside, same era.
Drivevlm: The convergence of autonomous driving and large vision-language models
X. Tian, J. Gu, B. Li, Y. Liu, Y. Wang, Z. Zhao, K. Zhan, P. Jia, X. Lang, and H. Zhao · 2024
Navsim: Data-driven non-reactive autonomous vehicle simulation and benchmarking
D. Dauner, M. Hallgarten, T. Li, X. Weng, Z. Huang, Z. Yang, H. Li, I. Gilitschenski, B. Ivanovic, M. Pavone, et al · 2024
Later among the works it cites.
Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning
H. Shao, S. Qian, H. Xiao, G. Song, Z. Zong, L. Wang, Y. Liu, and H. Li · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, et al · 2024
Later among the works it cites.
Online preference-based reinforcement learning with self-augmented feedback from large language model
S. Tu, J. Sun, Q. Zhang, X. Lan, and D. Zhao · 2024
Later among the works it cites.
Carllava: Vision language models for camera-only closed-loop driving
K. Renz, L. Chen, A.-M. Marcu, J. Hünermann, B. Hanotte, A. Karnsund, J. Shotton, E. Arani, and O. Sinavski · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Senna: Bridging large vision-language models and end-to-end autonomous driving
B. Jiang, S. Chen, B. Liao, X. Zhang, W. Yin, Q. Zhang, C. Huang, W. Liu, and X. Wang · 2024
Cited alongside, same era.
Vlp: Vision language planning for autonomous driving
C. Pan, B. Yaman, T. Nesti, A. Mallik, A. G. Allievi, S. Velipasalar, and L. Ren · 2024
Cited alongside, same era.
Lmdrive: Closed-loop end-to-end driving with large language models
H. Shao, Y. Hu, L. Wang, G. Song, S. L. Waslander, Y. Liu, and H. Li · 2024
Cited alongside, same era.
Tokenize the world into object-level knowledge to address long-tail events in autonomous driving
T. Tian, B. Li, X. Weng, Y. Chen, E. Schmerling, Y. Wang, B. Ivanovic, and M. Pavone · 2024
Cited alongside, same era.
Reason2drive: Towards interpretable and chain-based reasoning for autonomous driving
M. Nie, R. Peng, C. Wang, X. Cai, J. Han, H. Xu, and L. Zhang · 2024
Cited alongside, same era.
Dilu: A knowledge-driven approach to autonomous driving with large language models
L. Wen, D. Fu, X. Li, X. Cai, M. Tao, P. Cai, M. Dou, B. Shi, L. He, and Y. Qiao · 2024
Cited alongside, same era.
Llm4drive: A survey of large language models for autonomous driving
Z. Yang, X. Jia, H. Li, and J. Yan · 2024
Cited alongside, same era.
Later among the works it cites.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Later among the works it cites.
Uncad: Towards safe end-to-end autonomous driving via online map uncertainty
P. Yang, Y. Zheng, Q. Zhang, K. Zhu, Z. Xing, Q. Lin, Y.-F. Liu, Z. Su, and D. Zhao · 2025
Closest in time.
Empowering llm agents with zero-shot optimal decision-making through q-learning
J. Chai, S. Li, Y. Fu, D. Zhao, and Y. Zhu · 2025
Closest in time.
Distilling multi-modal large language models for autonomous driving
D. Hegde, R. Yasarla, H. Cai, S. Han, A. Bhattacharyya, S. Mahajan, L. Liu, R. Garrepalli, V. M. Patel, and F. Porikli · 2025
Closest in time.
Don’t shake the wheel: Momentum-aware planning in end-to-end autonomous driving
Z. Song, C. Jia, L. Liu, H. Pan, Y. Zhang, J. Wang, X. Zhang, S. Xu, L. Yang, and Y. Luo · 2025
Closest in time.
Enhancing end-to-end autonomous driving with latent world model
Y. Li, L. Fan, J. He, Y. Wang, Y. Chen, Z. Zhang, and T. Tan · 2025
Closest in time.
Diffusiondrive: Truncated diffusion model for end-to-end autonomous driving
B. Liao, S. Chen, H. Yin, B. Jiang, C. Wang, S. Yan, X. Zhang, X. Li, Y. Zhang, Q. Zhang, et al · 2025
Closest in time.
Goalflow: Goal-driven flow matching for multimodal trajectories generation in end-to-end autonomous driving
Z. Xing, X. Zhang, Y. Hu, B. Jiang, T. He, Q. Zhang, X. Long, and W. Yin · 2025
Closest in time.
Conrft: A reinforced fine-tuning method for vla models via consistency policy
Y. Chen, S. Tian, S. Liu, Y. Zhou, H. Li, and D. Zhao · 2025
Closest in time.
Simlingo: Vision-only closed-loop autonomous driving with language-action alignment
K. Renz, L. Chen, E. Arani, and O. Sinavski · 2025
Closest in time.