Fetching the paper…
Reading the bibliography…
Vision Large Language Models (VLLMs) have demonstrated impressive capabilities in general visual tasks such as image captioning and visual question answering.
Description of the SHRP 2 naturalistic database and the crash, near-crash, and baseline data sets
Jonathan M Hankey, Miguel A Perez, and Julie A McClafferty. 2016 · 2016
Earlier work this paper cites.
Textual explanations for self-driving vehicles. In Proceedings of the European conference on computer vision (ECCV) . 563–578
Jinkyu Kim, Anna Rohrbach, Trevor Darrell, John Canny, and Zeynep Akata. 2018 · 2018
Earlier work this paper cites.
A framework for automated driving system testable cases and scenarios
Eric Thorn, Shawn C Kimmel, Michelle Chaka, Booz Allen Hamilton, et al · 2018
Earlier work this paper cites.
Talk2car: Taking control of your self-driving car
Thierry Deruyttere, Simon Vandenhende, Dusan Grujicic, Luc Van Gool, and Marie-Francine Moens. 2019 · 2019
Earlier work this paper cites.
Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 2446–2454
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al · 2020
Earlier work this paper cites.
Explainable object-induced action decision for autonomous vehicles. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9523–9532
Yiran Xu, Xiaoyin Yang, Lihang Gong, Hsuan-Chu Lin, Tz-Ying Wu, Yunsheng Li, and Nuno Vasconcelos. 2020 · 2020
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Earlier work this paper cites.
DriveScenify: Boosting Driving Scene Understanding with Advanced Vision-Language Models
Xiaowei Gao, Pengxiang Li, xinke Jiang, James Haworth, Jonathan Cardoso-Silva, and Ming Li. 2023 · 2023
Earlier work this paper cites.
Challenges and applications of large language models
Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy. 2023 · 2023
Earlier work this paper cites.
Video-llava: Learning united visual representation by alignment before projection
Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan. 2023 · 2023
Earlier work this paper cites.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan. 2023 · 2023
Earlier work this paper cites.
Drama: Joint risk localization and captioning in driving. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 1043–1052
Srikanth Malla, Chiho Choi, Isht Dwivedi, Joon Hee Choi, and Jiachen Li. 2023 · 2023
Earlier work this paper cites.
Gpt-driver: Learning to drive with gpt
Jiageng Mao, Yuxi Qian, Junjie Ye, Hang Zhao, and Yue Wang. 2023 · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al · 2023
Earlier work this paper cites.
Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 11975–11986
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. 2023 · 2023
Earlier work this paper cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023 · 2023
Cited alongside, same era.
Toolqa: A dataset for llm question answering with external tools
Yuchen Zhuang, Yue Yu, Kuan Wang, Haotian Sun, and Chao Zhang. 2023 · 2023
Cited alongside, same era.
Premise order matters in reasoning with large language models
Xinyun Chen, Ryan A Chi, Xuezhi Wang, and Denny Zhou. 2024 · 2024
Cited alongside, same era.
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, et al · 2024
Cited alongside, same era.
Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 7513–7522
Enna Sachdeva, Nakul Agarwal, Suhas Chundi, Sean Roelofs, Jiachen Li, Mykel Kochenderfer, Chiho Choi, and Behzad Dariush. 2024 · 2024
Later among the works it cites.
ScVLM: a Vision-Language Model for Driving Safety Critical Event Understanding
Liang Shi, Boyu Jiang, and Feng Guo. 2024 · 2024
Later among the works it cites.
Enhancing Traffic Safety with Parallel Dense Video Captioning for End-to-End Event Analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops . 7125–7133
Maged Shoman, Dongdong Wang, Armstrong Aboah, and Mohamed Abdel-Aty. 2024 · 2024
Later among the works it cites.
Drivelm: Driving with graph visual question answering. In European Conference on Computer Vision . Springer, 256–274
Chonghao Sima, Katrin Renz, Kashyap Chitta, Li Chen, Hanxue Zhang, Chengen Xie, Jens Beißwenger, Ping Luo, Andreas Geiger, and Hongyang Li. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xianda Guo, Ruijun Zhang, Yiqun Duan, Yuhang He, Chenming Zhang, Shuai Liu, and Long Chen. 2024 · 2024
Cited alongside, same era.
Dme-driver: Integrating human decision logic and 3d scene perception in autonomous driving
Wencheng Han, Dongqian Guo, Cheng-Zhong Xu, and Jianbing Shen. 2024 · 2024
Cited alongside, same era.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Cited alongside, same era.
Semantic Understanding of Traffic Scenes with Large Vision Language Models. In 2024 IEEE Intelligent Vehicles Symposium (IV) . IEEE, 1580–1587
Sandesh Jain, Surendrabikram Thapa, Kuan-Ting Chen, A Lynn Abbott, and Abhijit Sarkar. 2024 · 2024
Cited alongside, same era.
Llava-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. 2024b · 2024
Cited alongside, same era.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024a · 2024
Cited alongside, same era.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024b · 2024
Cited alongside, same era.
Can LVLMs Obtain a Driver’s License? A Benchmark Towards Reliable AGI for Autonomous Driving
Yuhang Lu, Yichen Yao, Jiadong Tu, Jiangnan Shao, Yuexin Ma, and Xinge Zhu. 2024 · 2024
Cited alongside, same era.
Drivevlm: The convergence of autonomous driving and large vision-language models
Xiaoyu Tian, Junru Gu, Bailin Li, Yicheng Liu, Yang Wang, Zhiyong Zhao, Kun Zhan, Peng Jia, Xianpeng Lang, and Hang Zhao. 2024 · 2024
Later among the works it cites.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al · 2024
Later among the works it cites.
Pllava: Parameter-free llava extension from images to videos for video dense captioning
Lin Xu, Yilin Zhao, Daquan Zhou, Zhijie Lin, See Kiong Ng, and Jiashi Feng. 2024b · 2024
Later among the works it cites.
Drivegpt4: Interpretable end-to-end autonomous driving via large language model
Zhenhua Xu, Yujia Zhang, Enze Xie, Zhen Zhao, Yong Guo, Kwan-Yee K Wong, Zhenguo Li, and Hengshuang Zhao. 2024a · 2024
Later among the works it cites.
Minicpm-v: A gpt-4v level mllm on your phone
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al · 2024
Later among the works it cites.
LLaVA-NeXT: A Strong Zero-shot Video Understanding Model
Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li. 2024 · 2024
Later among the works it cites.
Llama-vid: An image is worth 2 tokens in large language models. In European Conference on Computer Vision . Springer, 323–340
Yanwei Li, Chengyao Wang, and Jiaya Jia. 2025 · 2025
Closest in time.
Mmbench: Is your multi-modal model an all-around player?. In European Conference on Computer Vision . Springer, 216–233
Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al · 2025
Closest in time.
Reason2drive: Towards interpretable and chain-based reasoning for autonomous driving. In European Conference on Computer Vision . Springer, 292–308
Ming Nie, Renyuan Peng, Chunwei Wang, Xinyue Cai, Jianhua Han, Hang Xu, and Li Zhang. 2025 · 2025
Closest in time.
Shaoyuan Xie, Lingdong Kong, Yuhao Dong, Chonghao Sima, Wenwei Zhang, Qi Alfred Chen, Ziwei Liu, and Liang Pan. 2025 · 2025
Closest in time.