Fetching the paper…
Reading the bibliography…
Recent advancements in Vision-Language Models (VLMs) have demonstrated strong potential for autonomous driving tasks.
Sun rgb-d: A rgb-d scene understanding benchmark suite
Shuran Song, Samuel P. Lichtenberg, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning, 2016
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering, 2017
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering, 2019
Drew A. Hudson and Christopher D. Manning · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi · 2019
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning, 2022
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan · 2022
Earlier work this paper cites.
Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction, 2022
Cheng Sun, Min Sun, and Hwann-Tzong Chen · 2022
Earlier work this paper cites.
Understanding depth map progressively: Adaptive distance interval separation for monocular 3d object detection, 2023
Xianhui Cheng, Shoumeng Qiu, Zhikang Zou, Jian Pu, and Xiangyang Xue · 2023
Earlier work this paper cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Earlier work this paper cites.
Language guided visual question answering: Elevate your multimodal language model using knowledge-enriched prompts, 2023
Deepanway Ghosal, Navonil Majumder, Roy Ka-Wei Lee, Rada Mihalcea, and Soujanya Poria · 2023
Earlier work this paper cites.
Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations, 2023
Yuichi Inoue, Yuki Yada, Kotaro Tanahashi, and Yu Yamaguchi · 2023
Earlier work this paper cites.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Earlier work this paper cites.
Lingoqa: video question answering for autonomous driving (2023)
AM Marcu, L Chen, J Hünermann, A Karnsund, B Hanotte, P Chidananda, S Nair, V Badrinarayanan, A Kendall, J Shotton, et al · 2023
Earlier work this paper cites.
Multi-camera bird’s eye view perception for autonomous driving, 2023
David Unger, Nikhil Gosala, Varun Ravi Kumar, Shubhankar Borse, Abhinav Valada, and Senthil Yogamani · 2023
Cited alongside, same era.
Myvlm: Personalizing vlms for user-specific queries
Yuval Alaluf, Elad Richardson, Sergey Tulyakov, Kfir Aberman, and Daniel Cohen-Or · 2024
Cited alongside, same era.
Covla: Comprehensive vision-language-action dataset for autonomous driving
Hidehisa Arai, Keita Miwa, Kento Sasaki, Yu Yamaguchi, Kohei Watanabe, Shunsuke Aoki, and Issei Yamamoto · 2024
Cited alongside, same era.
Spatialbot: Precise spatial understanding with vision language models, 2024
Wenxiao Cai, Iaroslav Ponomarenko, Jianhao Yuan, Xiaoqi Li, Wankou Yang, Hao Dong, and Bo Zhao · 2024
Cited alongside, same era.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities, 2024
Boyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter, Danny Driess, Pete Florence, Dorsa Sadigh, Leonidas Guibas, and Fei Xia · 2024
Cited alongside, same era.
Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
AI Meta · 2024
Later among the works it cites.
Reason2drive: Towards interpretable and chain-based reasoning for autonomous driving
Ming Nie, Renyuan Peng, Chunwei Wang, Xinyue Cai, Jianhua Han, Hang Xu, and Li Zhang · 2024
Later among the works it cites.
Gpt-4o system card, 2024
OpenAI · 2024
Later among the works it cites.
Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario, 2024
Tianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao, and Yu-Gang Jiang · 2024
Later among the works it cites.
Drivelm: Driving with graph visual question answering
Chonghao Sima, Katrin Renz, Kashyap Chitta, Li Chen, Hanxue Zhang, Chengen Xie, Jens Beißwenger, Ping Luo, Andreas Geiger, and Hongyang Li · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spatialrgpt: Grounded spatial reasoning in vision language models, 2024
An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo, Ruihan Yang, Jan Kautz, Xiaolong Wang, and Sifei Liu · 2024
Cited alongside, same era.
Deep attention driven reinforcement learning (dad-rl) for autonomous decision-making in dynamic environment, 2024
Jayabrata Chowdhury, Venkataramanan Shivaraman, Sumit Dangi, Suresh Sundaram, and P. B. Sujit · 2024
Cited alongside, same era.
Blink: Multimodal large language models can see but not perceive
Xingyu Fu, Yushi Hu, Bangzheng Li, Yu Feng, Haoyu Wang, Xudong Lin, Dan Roth, Noah A Smith, Wei-Chiu Ma, and Ranjay Krishna · 2024
Cited alongside, same era.
Akshay Gopalkrishnan, Ross Greer, and Mohan Trivedi · 2024
Cited alongside, same era.
Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation
Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai, Xiaoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen · 2024
Cited alongside, same era.
Emma: End-to-end multimodal model for autonomous driving
Jyh-Jing Hwang, Runsheng Xu, Hubert Lin, Wei-Chih Hung, Jingwei Ji, Kristy Choi, Di Huang, Tong He, Paul Covington, Benjamin Sapp, et al · 2024
Cited alongside, same era.
Senna: Bridging large vision-language models and end-to-end autonomous driving, 2024
Bo Jiang, Shaoyu Chen, Bencheng Liao, Xingyu Zhang, Wei Yin, Qian Zhang, Chang Huang, Wenyu Liu, and Xinggang Wang · 2024
Cited alongside, same era.
Shihao Wang, Zhiding Yu, Xiaohui Jiang, Shiyi Lan, Min Shi, Nadine Chang, Jan Kautz, Ying Li, and Jose M Alvarez · 2024
Later among the works it cites.
Autotrust: Benchmarking trustworthiness in large vision language models for autonomous driving
Shuo Xing, Hongyuan Hua, Xiangbo Gao, Shenzhe Zhu, Renjie Li, Kexin Tian, Xiaopeng Li, Heng Huang, Tianbao Yang, Zhangyang Wang, et al · 2024
Later among the works it cites.
Jianhao Yuan, Shuyang Sun, Daniel Omeiza, Bo Zhao, Paul Newman, Lars Kunze, and Matthew Gadd · 2024
Later among the works it cites.
A survey of automatic driving environment perception
Hui Zhao, Xin Li, Cheng Xu, Bingxin Xu, and Hongzhe Liu · 2024
Later among the works it cites.
Simplellm4ad: An end-to-end vision-language model with graph visual question answering for autonomous driving, 2024
Peiru Zheng, Yun Zhao, Zhan Gong, Hong Zhu, and Shaohua Wu · 2024
Later among the works it cites.
Investigating prompting techniques for zero- and few-shot visual question answering, 2025
Rabiul Awal, Le Zhang, and Aishwarya Agrawal · 2025
Closest in time.
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, et al · 2025
Closest in time.
Stamp: Scalable task and model-agnostic collaborative perception
Xiangbo Gao, Runsheng Xu, Jiachen Li, Ziran Wang, Zhiwen Fan, and Zhengzhong Tu · 2025
Closest in time.
Openemma: Open-source multimodal model for end-to-end autonomous driving
Shuo Xing, Chengyuan Qian, Yuping Wang, Hongyuan Hua, Kexin Tian, Yang Zhou, and Zhengzhong Tu · 2025
Closest in time.
Multimodal inconsistency reasoning (mmir): A new benchmark for multimodal reasoning models, 2025
Qianqi Yan, Yue Fan, Hongquan Li, Shan Jiang, Yang Zhao, Xinze Guan, Ching-Chen Kuo, and Xin Eric Wang · 2025
Closest in time.