Fetching the paper…
Reading the bibliography…
The rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language understanding, and control within a single policy.
Autonomous driving in traffic: Boss and the urban challenge
Chris Urmson, Chris Baker, John Dolan, Paul Rybski, Bryan Salesky, William “Red” Whittaker, Dave Ferguson, and Michael Darms · 2009
Earlier work this paper cites.
A survey on motion prediction and risk assessment for intelligent vehicles
Stéphanie Lefèvre, Dizan Vasquez, and Christian Laugier · 2014
Earlier work this paper cites.
A review of motion planning techniques for automated vehicles
David González, Joshué Pérez, Vicente Milanés, and Fawzi Nashashibi · 2015
Earlier work this paper cites.
Visual object recognition with 3d-aware features in kitti urban scenes
J Javier Yebes, Luis M Bergasa, and Miguel Ángel García-Garrido · 2015
Earlier work this paper cites.
A survey of motion planning and control techniques for self-driving urban vehicles
Brian Paden, Michal Čáp, Sze Zheng Yong, Dmitry Yershov, and Emilio Frazzoli · 2016
Earlier work this paper cites.
An lstm network for highway trajectory prediction
Florent Altché and Arnaud de La Fortelle · 2017
Earlier work this paper cites.
Carla: An open urban driving simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun · 2017
Earlier work this paper cites.
Textual explanations for self-driving vehicles
Jinkyu Kim, Z. Li, B. Floyd, et al · 2018
Earlier work this paper cites.
Fast and furious: Real time end-to-end 3d detection, tracking and motion forecasting with a single convolutional net
Wenjie Luo, Bin Yang, and Raquel Urtasun · 2018
Earlier work this paper cites.
Planning and decision-making for autonomous vehicles
Wilko Schwarting, Javier Alonso-Mora, and Daniela Rus · 2018
Earlier work this paper cites.
Bdd100k: A diverse driving video database with scalable annotation tooling
Fisher Yu, Wenqi Xian, Yingying Chen, Fangchen Liu, Mike Liao, Vashisht Madhavan, Trevor Darrell, et al · 2018
Earlier work this paper cites.
Learning to drive in a day
Alex Kendall, Jeffrey Hawke, David Janz, Przemyslaw Mazur, Daniele Reda, John-Mark Allen, Vinh-Dieu Lam, Alex Bewley, and Amar Shah · 2019
Earlier work this paper cites.
End-to-end interpretable neural motion planner
Wenyuan Zeng, Wenjie Luo, Simon Suo, Abbas Sadat, Bin Yang, Sergio Casas, and Raquel Urtasun · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Alex Bankiti, Orien Lang, et al · 2020
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom · 2020
Earlier work this paper cites.
Urban driving with conditional imitation learning
Jeffrey Hawke, Richard Shen, Corina Gurau, Siddharth Sharma, Daniele Reda, Nikolay Nikolov, Przemysław Mazur, Sean Micklethwaite, Nicolas Griffiths, Amar Shah, et al · 2020
Earlier work this paper cites.
Pnpnet: End-to-end perception and prediction with tracking in the loop
Ming Liang, Bin Yang, Wenyuan Zeng, Yun Chen, Rui Hu, Sergio Casas, and Raquel Urtasun · 2020
Earlier work this paper cites.
Perceive, predict, and plan: Safe motion planning through interpretable semantic representations
Abbas Sadat, Sergio Casas, Mengye Ren, Xinyu Wu, Pranaab Dhawan, and Raquel Urtasun · 2020
Earlier work this paper cites.
Scalability in perception for autonomous driving: Waymo open dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al · 2020
Earlier work this paper cites.
What data do we need for training an av motion planner?
Long Chen, Lukas Platinsky, Stefanie Speichert, Błażej Osiński, Oliver Scheel, Yawei Ye, Hugo Grimmett, Luca Del Pero, and Peter Ondruska · 2021
Earlier work this paper cites.
Neat: Neural attention fields for end-to-end autonomous driving
Kashyap Chitta, Aditya Prakash, and Andreas Geiger · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Earlier work this paper cites.
Learning from all vehicles
Dian Chen and Philipp Krähenbühl · 2022
Earlier work this paper cites.
Transfuser: Imitation with transformer-based sensor fusion for autonomous driving
Kashyap Chitta, Aditya Prakash, Bernhard Jaeger, Zehao Yu, Katrin Renz, and Andreas Geiger · 2022
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2022
Earlier work this paper cites.
A survey on trajectory-prediction methods for autonomous driving
Yanjun Huang, Jiatong Du, Ziru Yang, Zewei Zhou, Lin Zhang, and Hong Chen · 2022
Earlier work this paper cites.
Plant: Explainable planning transformers via object-level representations
Katrin Renz, Kashyap Chitta, Otniel-Bogdan Mercea, A Koepke, Zeynep Akata, and Andreas Geiger · 2022
Earlier work this paper cites.
Beverse: Unified perception and prediction in birds-eye-view for vision-centric autonomous driving
Yunpeng Zhang, Zheng Zhu, Wenzhao Zheng, Junjie Huang, Guan Huang, Jie Zhou, and Jiwen Lu · 2022
Earlier work this paper cites.
Fine-grained affective processing capabilities emerging from large language models
Joost Broekens, Bernhard Hilpert, Suzan Verberne, Kim Baraka, Patrick Gebhard, and Aske Plaat · 2023
Earlier work this paper cites.
Planning-oriented autonomous driving
Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, et al · 2023
Earlier work this paper cites.
Planning-oriented autonomous driving
Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, et al · 2023
Earlier work this paper cites.
Yolo-v1 to yolo-v8, the rise of yolo and its complementary nature toward digital manufacturing and industrial defect detection
Muhammad Hussain · 2023
Earlier work this paper cites.
Adriver-i: A general world model for autonomous driving
Fan Jia, Weixin Mao, Yingfei Liu, Yucheng Zhao, Yuqing Wen, Chi Zhang, Xiangyu Zhang, and Tiancai Wang · 2023
Earlier work this paper cites.
Think twice before driving: Towards scalable decoders for end-to-end autonomous driving
Xiaosong Jia, Penghao Wu, Li Chen, Jiangwei Xie, Conghui He, Junchi Yan, and Hongyang Li · 2023
Earlier work this paper cites.
Vad: Vectorized scene representation for efficient autonomous driving
Bo Jiang, Shaoyu Chen, Qing Xu, Bencheng Liao, Jiajie Chen, Helong Zhou, Qian Zhang, Wenyu Liu, Chang Huang, and Xinggang Wang · 2023
Earlier work this paper cites.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Earlier work this paper cites.
Mtd-gpt: A multi-task decision-making gpt model for autonomous driving at unsignalized intersections
Jiaqi Liu, Peng Hang, Xiao Qi, Jianqiang Wang, and Jian Sun · 2023
Earlier work this paper cites.
Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation
Zhijian Liu, Haotian Tang, Alexander Amini, Xinyu Yang, Huizi Mao, Daniela L Rus, and Song Han · 2023
Earlier work this paper cites.
Gpt-driver: Learning to drive with gpt
Jiageng Mao, Yuxi Qian, Junjie Ye, Hang Zhao, and Yue Wang · 2023
Earlier work this paper cites.
A language agent for autonomous driving
Jiageng Mao, Junjie Ye, Yuxi Qian, Marco Pavone, and Yue Wang · 2023
Earlier work this paper cites.
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Earlier work this paper cites.
Safety-enhanced autonomous driving using interpretable sensor fusion transformer
Hao Shao, Letian Wang, Ruobing Chen, Hongsheng Li, and Yu Liu · 2023
Earlier work this paper cites.
Reasonnet: End-to-end driving with temporal and global reasoning
Hao Shao, Letian Wang, Ruobing Chen, Steven L Waslander, Hongsheng Li, and Yu Liu · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Earlier work this paper cites.
Wenhai Wang, Jiangwei Xie, ChuanYang Hu, Haoming Zou, Jianan Fan, Wenwen Tong, Yang Wen, Silei Wu, Hanming Deng, Zhiqi Li, et al · 2023
Earlier work this paper cites.
S4tp: Social-suitable and safety-sensitive trajectory planning for autonomous vehicles
Xiao Wang, Ke Tang, Xingyuan Dai, Jintao Xu, Quancheng Du, Rui Ai, Yuxiao Wang, and Weihao Gu · 2023
Earlier work this paper cites.
Dilu: A knowledge-driven approach to autonomous driving with large language models
Licheng Wen, Daocheng Fu, Xin Li, Xinyu Cai, Tao Ma, Pinlong Cai, Min Dou, Botian Shi, Liang He, and Yu Qiao · 2023
Earlier work this paper cites.
Argoverse 2: Next generation datasets for self-driving perception and forecasting
Benjamin Wilson, William Qi, Tanmay Agarwal, John Lambert, Jagjeet Singh, Siddhesh Khandelwal, Bowen Pan, Ratnesh Kumar, Andrew Hartnett, Jhony Kaesemodel Pontes, et al · 2023
Earlier work this paper cites.
Convnext v2: Co-designing and scaling convnets with masked autoencoders
Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu, In So Kweon, and Saining Xie · 2023
Earlier work this paper cites.
Fusionad: Multi-modality fusion for prediction and planning tasks of autonomous driving
Tengju Ye, Wei Jing, Chunyong Hu, Shikun Huang, Lingping Gao, Fangzhen Li, Jingke Wang, Ke Guo, Wencong Xiao, Weibo Mao, et al · 2023
Earlier work this paper cites.
pi0: A vision-language-action flow model for general robot control
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al · 2024
Earlier work this paper cites.
End-to-end autonomous driving: Challenges and frontiers
Li Chen, Penghao Wu, Kashyap Chitta, Bernhard Jaeger, Andreas Geiger, and Hongyang Li · 2024
Cited alongside, same era.
Vadv2: End-to-end vectorized autonomous driving via probabilistic planning
Shaoyu Chen, Bo Jiang, Hao Gao, Bencheng Liao, Qing Xu, Qian Zhang, Chang Huang, Wenyu Liu, and Xinggang Wang · 2024
Cited alongside, same era.
Asynchronous large language model enhanced planner for autonomous driving
Yuan Chen, Zi-han Ding, Ziqin Wang, Yan Wang, Lijun Zhang, and Si Liu · 2024
Cited alongside, same era.
Ppad: Iterative interactions of prediction and planning for end-to-end autonomous driving
Zhili Chen, Maosheng Ye, Shuangjie Xu, Tongyi Cao, and Qifeng Chen · 2024
Cited alongside, same era.
Talk2bev: Language-enhanced bird’s-eye view maps for autonomous driving
Tushar Choudhary, Vikrant Dewangan, Shivam Chandhok, Shubham Priyadarshan, Anushka Jain, Arun K Singh, Siddharth Srivastava, Krishna Murthy Jatavallabhula, and K Madhava Krishna · 2024
Cited alongside, same era.
Enming Zhang, Xingyuan Dai, Yisheng Lv, and Qinghai Miao · 2024
Later among the works it cites.
Instruct large language models to drive like humans
Ruijun Zhang, Xianda Guo, Wenzhao Zheng, Chenming Zhang, Kurt Keutzer, and Long Chen · 2024
Later among the works it cites.
Weize Zhang, Mohammed Elmahgiubi, Kasra Rezaee, Behzad Khamidehi, Hamidreza Mirkhani, Fazel Arasteh, Chunlin Li, Muhammad Ahsan Kaleem, Eduardo R Corral-Soto, Dhruv Sharma, et al · 2024
Later among the works it cites.
Graphad: Interaction scene graph for end-to-end autonomous driving
Yunpeng Zhang, Deheng Qian, Ding Li, Yifeng Pan, Yong Chen, Zhenbao Liang, Zhiyao Zhang, Shurui Zhang, Hongxu Li, Maolei Fu, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A survey on multimodal large language models for autonomous driving
Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, Yang Zhou, Kaizhao Liang, Jintai Chen, Juanwu Lu, Zichong Yang, Kuei-Da Liao, et al · 2024
Cited alongside, same era.
Hint-ad: Holistically aligned interpretability in end-to-end autonomous driving
Kairui Ding, Boyuan Chen, Yuchen Su, Huan-ang Gao, Bu Jin, Chonghao Sima, Wuqiang Zhang, Xiaohui Li, Paul Barsch, Hongyang Li, et al · 2024
Cited alongside, same era.
Dualad: Disentangling the dynamic and static world for end-to-end driving
Simon Doll, Niklas Hanselmann, Lukas Schneider, Richard Schulz, Marius Cordts, Markus Enzweiler, and Hendrik Lensch · 2024
Cited alongside, same era.
On the road to portability: Compressing end-to-end motion planner for autonomous driving
Kaituo Feng, Changsheng Li, Dongchun Ren, Ye Yuan, and Guoren Wang · 2024
Cited alongside, same era.
Polarpoint-bev: Bird-eye-view perception in polar points for explainable end-to-end autonomous driving
Yuchao Feng and Yuxiang Sun · 2024
Cited alongside, same era.
Drive like a human: Rethinking autonomous driving with large language models
Daocheng Fu, Xin Li, Licheng Wen, Min Dou, Pinlong Cai, Botian Shi, and Yu Qiao · 2024
Cited alongside, same era.
A survey for foundation models in autonomous driving
Haoxiang Gao, Zhongruo Wang, Yaqian Li, Kaiwen Long, Ming Yang, and Yiqing Shen · 2024
Cited alongside, same era.
3d-vla: A 3d vision-language-action generative world model
Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan · 2024
Later among the works it cites.
Genad: Generative end-to-end autonomous driving
Wenzhao Zheng, Ruiqi Song, Xianda Guo, Chenming Zhang, and Long Chen · 2024
Later among the works it cites.
Yupeng Zheng, Zhongpu Xia, Qichao Zhang, Teng Zhang, Ben Lu, Xiaochuang Huo, Chao Han, Yixian Li, Mengjie Yu, Bu Jin, et al · 2024
Later among the works it cites.
Enhance planning with physics-informed safety controller for end-to-end autonomous driving
Hang Zhou, Haichao Liu, Hongliang Lu, Jun Ma, and Yiding Ji · 2024
Later among the works it cites.
Vision language models in autonomous driving: A survey and outlook
Xingcheng Zhou, Mingyu Liu, Ekim Yurtsever, Bare Luka Zagar, Walter Zimmer, Hu Cao, and Alois C Knoll · 2024
Later among the works it cites.
Vavim and vavam: Autonomous driving through video generative modeling
Florent Bartoccioni, Elias Ramzi, Victor Besnier, Shashanka Venkataramanan, Tuan-Hung Vu, Yihong Xu, Loick Chambon, Spyros Gidaris, Serkan Odabas, David Hurych, et al · 2025
Closest in time.
Insight: Enhancing autonomous driving safety through vision-language models on context-aware hazard detection and edge case evaluation
Dianwei Chen, Zifan Zhang, Yuchen Liu, and Xianfeng Terry Yang · 2025
Closest in time.
Ts-vlm: Text-guided softsort pooling for vision-language models in multi-view driving reasoning
Lihong Chen, Hossein Hassani, and Soodeh Nikan · 2025
Closest in time.
CoVLA: Comprehensive vision-language-action dataset for autonomous driving
Haohan Chi, Huan-ang Gao, Ziming Liu, et al · 2025
Closest in time.
Impromptu vla: Open weights and open data for driving vision-language-action models
Haohan Chi, Huan-ang Gao, Ziming Liu, et al · 2025
Closest in time.
Chain-of-thought for autonomous driving: A comprehensive survey and future prospects
Yixin Cui, Haotian Lin, Shuo Yang, Yixiao Wang, Yanjun Huang, and Hong Chen · 2025
Closest in time.
Haoyu Fu, Diankun Zhang, Zongchuang Zhao, Jianfeng Cui, Dingkang Liang, Chong Zhang, Dingyuan Zhang, Hongwei Xie, Bing Wang, and Xiang Bai · 2025
Closest in time.
Haoyu Fu, Diankun Zhang, Zongchuang Zhao, Jianfeng Cui, Dingkang Liang, Chong Zhang, Dingyuan Zhang, Hongwei Xie, Bing Wang, and Xiang Bai · 2025
Closest in time.
Langcoop: Collaborative driving with language
Xiangbo Gao, Yuheng Wu, Rujia Wang, et al · 2025
Closest in time.
ipad: Iterative proposal-centric end-to-end autonomous driving
Ke Guo, Haochen Liu, Xiaojun Wu, Jia Pan, and Chen Lv · 2025
Closest in time.
Dme-driver: Integrating human decision logic and 3d scene perception in autonomous driving
Wencheng Han, Dongqian Guo, Cheng-Zhong Xu, and Jianbing Shen · 2025
Closest in time.
Driveaction: A benchmark for exploring human-like driving decisions in vla models
Yuhan Hao, Zhengning Li, Lei Sun, Weilong Wang, Naixin Yi, Sheng Song, Caihong Qin, Mofan Zhou, Yifei Zhan, Peng Jia, et al · 2025
Closest in time.
Xinmeng Hou, Wuqi Wang, Long Yang, Hao Lin, Jinglun Feng, Haigen Min, and Xiangmo Zhao · 2025
Closest in time.
Nora: A small open-sourced generalist vision language action model for embodied tasks
Chia-Yu Hung, Qi Sun, Pengfei Hong, Amir Zadeh, Chuan Li, U Tan, Navonil Majumder, Soujanya Poria, et al · 2025
Closest in time.
Ayesha Ishaq, Jean Lahoud, Ketan More, Omkar Thawakar, Ritesh Thawkar, Dinura Dissanayake, Noor Ahsan, Yuhao Li, Fahad Shahbaz Khan, Hisham Cholakkal, et al · 2025
Closest in time.
Drivetransformer: Unified transformer for scalable end-to-end autonomous driving
Xiaosong Jia, Junqi You, Zhiyuan Zhang, and Junchi Yan · 2025
Closest in time.
Diffvla: Vision-language guided diffusion planning for autonomous driving
Anqing Jiang, Yu Gao, Zhigang Sun, et al · 2025
Closest in time.
Fine-tuning vision-language-action models: Optimizing speed and success
Moo Jin Kim, Chelsea Finn, and Percy Liang · 2025
Closest in time.
Pointvla: Injecting the 3d world into vision-language-action models
Chengmeng Li, Junjie Wen, Yan Peng, Yaxin Peng, Feifei Feng, and Yichen Zhu · 2025
Closest in time.
Recogdrive: A reinforced cognitive framework for end-to-end autonomous driving, 2025
Yongkang Li, Kaixin Xiong, Xiangyu Guo, Fang Li, Sixu Yan, Gangwei Xu, Lijun Zhou, Long Chen, Haiyang Sun, Bing Wang, Guang Chen, Hangjun Ye, Wenyu Liu, and Xinggang Wang · 2025
Closest in time.
Generalized trajectory scoring for end-to-end multimodal planning
Zhenxin Li, Wenhao Yao, Zi Wang, Xinglong Sun, Joshua Chen, Nadine Chang, Maying Shen, Zuxuan Wu, Shiyi Lan, and Jose M Alvarez · 2025
Closest in time.
Lin Liu, Ziying Song, Hongyu Pan, Lei Yang, and Caiyan Jia · 2025
Closest in time.
Vlm-e2e: Enhancing end-to-end autonomous driving with multimodal driver attention fusion
Pei Liu, Haipeng Liu, Haichao Liu, Xin Liu, Jinxin Ni, and Jun Ma · 2025
Closest in time.
Reasonplan: Unified scene prediction and decision reasoning for closed-loop autonomous driving
Xueyi Liu, Zuodong Zhong, Yuxin Guo, Yun-Fu Liu, Zhiguo Su, Qichao Zhang, Junli Wang, Yinfeng Gao, Yupeng Zheng, Qiao Lin, et al · 2025
Closest in time.
Leapvad: A leap in autonomous driving via cognitive perception and dual-process thinking
Yukai Ma, Tiantian Wei, Naiting Zhong, Jianbiao Mei, Tao Hu, Licheng Wen, Xuemeng Yang, Botian Shi, and Yong Liu · 2025
Closest in time.
Data scaling laws for end-to-end autonomous driving
Alexander Naumann, Xunjiang Gu, Tolga Dimlioglu, Mariusz Bojarski, Alperen Degirmenci, Alexander Popov, Devansh Bisla, Marco Pavone, Urs Müller, and Boris Ivanovic · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine · 2025
Closest in time.
Kangan Qian, Sicong Jiang, Yang Zhong, Ziang Luo, Zilin Huang, Tianze Zhu, Kun Jiang, Mengmeng Yang, Zheng Fu, Jinyu Miao, et al · 2025
Closest in time.
Kangan Qian, Ziang Luo, Sicong Jiang, Zilin Huang, Jinyu Miao, Zhikun Ma, Tianze Zhu, Jiayin Li, Yangfan He, Zheng Fu, et al · 2025
Closest in time.
Spatialvla: Exploring spatial representations for visual-language-action model
Delin Qu, Haoming Song, Qizhi Chen, Yuanqi Yao, Xinyi Ye, Yan Ding, Zhigang Wang, JiaYuan Gu, Bin Zhao, Dong Wang, et al · 2025
Closest in time.
Simlingo: Vision-only closed-loop autonomous driving with language-action alignment
Katrin Renz, Long Chen, Elahe Arani, and Oleg Sinavski · 2025
Closest in time.
Carllava: Vision language models for camera-only closed-loop driving
Katrin Renz, Long Chen, Ana-Maria Marcu, Jamie Shotton, et al · 2025
Closest in time.
Vision-language-action models: Concepts, progress, applications and challenges
Ranjan Sapkota, Yang Cao, Konstantinos I Roumeliotis, and Manoj Karkee · 2025
Closest in time.
Bevdriver: Leveraging bev maps in llms for robust closed-loop driving
Katharina Winter, Mark Azer, and Fabian B Flohr · 2025
Closest in time.
Shaoyuan Xie, Lingdong Kong, Yuhao Dong, Chonghao Sima, Wenwei Zhang, Qi Alfred Chen, Ziwei Liu, and Liang Pan · 2025
Closest in time.
Chatbev: A visual language model that understands bev maps
Qingyao Xu, Siheng Chen, Guang Chen, Yanfeng Wang, and Ya Zhang · 2025
Closest in time.
Lidar-llm: Exploring the potential of large language models for 3d lidar understanding
Senqiao Yang, Jiaming Liu, Renrui Zhang, Mingjie Pan, Ziyu Guo, Xiaoqi Li, Zehui Chen, Peng Gao, Hongsheng Li, Yandong Guo, et al · 2025
Closest in time.
Drivemoe: Mixture-of-experts for vision-language-action model in end-to-end autonomous driving
Zhenjie Yang, Yilin Chai, Xiaosong Jia, Yuqian Shao, et al · 2025
Closest in time.
Drivesuprim: Towards precise trajectory selection for end-to-end planning
Wenhao Yao, Zhenxin Li, Shiyi Lan, Zi Wang, Xinglong Sun, Jose M Alvarez, and Zuxuan Wu · 2025
Closest in time.
World knowledge-enhanced reasoning using instruction-guided interactor in autonomous driving
Mingliang Zhai, Cheng Li, Zengyuan Guo, Ningrui Yang, Xiameng Qin, Sanyuan Zhao, Junyu Han, Ji Tao, Yuwei Wu, and Yunde Jia · 2025
Closest in time.
Safeauto: Knowledge-enhanced safe autonomous driving with multimodal foundation models
Jiawei Zhang, Xuan Yang, Taiqi Wang, Yu Yao, Aleksandr Petiushko, and Bo Li · 2025
Closest in time.
Mpdrive: Improving spatial understanding with marker-based prompt learning for autonomous driving
Zhiyuan Zhang, Xiaofan Li, Zhihao Xu, Wenjie Peng, Zijian Zhou, Miaojing Shi, and Shuangping Huang · 2025
Closest in time.
Sce2drivex: A generalized mllm framework for scene-to-drive learning
Rui Zhao, Qirui Yuan, Jinyu Li, Haofeng Hu, Yun Li, Chengyuan Zheng, and Fei Gao · 2025
Closest in time.
Extending large vision-language model for diverse interactive tasks in autonomous driving
Zongcai Zhao, Yue Zhao, et al · 2025
Closest in time.
Opendrivevla: Towards end-to-end autonomous driving with large vision language action model
Xingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma, and Alois C Knoll · 2025
Closest in time.
Dynrsl-vlm: Enhancing autonomous driving perception with dynamic resolution vision-language models
Xirui Zhou, Lianlei Shan, and Xiaolin Gui · 2025
Closest in time.
Autovla: A vision-language-action model for end-to-end autonomous driving with adaptive reasoning and reinforcement fine-tuning, 2025
Zewei Zhou, Tianhui Cai, Seth Z. Zhao, Yun Zhang, Zhiyu Huang, Bolei Zhou, and Jiaqi Ma · 2025
Closest in time.