Fetching the paper…
Reading the bibliography…
Foundation models have indeed made a profound impact on various fields, emerging as pivotal components that significantly shape the capabilities of intelligent systems.
An approach to environmental psychology
Albert Mehrabian and James A Russell · 1974
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Are we ready for autonomous driving? the kitti vision benchmark suite
Andreas Geiger, Philip Lenz, and Raquel Urtasun · 2012
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell · 2016
Earlier work this paper cites.
Situational driving anger, driving performance and allocation of visual attention
Tingru Zhang, Alan HS Chan, Yutao Ba, and Wei Zhang · 2016
Earlier work this paper cites.
CARLA: An open urban driving simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H Lampert · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deeppicar: A low-cost deep neural network-based autonomous car
Michael G Bechtel, Elise McEllhiney, Minje Kim, and Heechul Yun · 2018
Earlier work this paper cites.
Using control synthesis to generate corner cases: A case study on autonomous driving
Glen Chou, Yunus Emre Sahin, Liren Yang, Kwesi J. Rutledge, Petter Nilsson, and Necmiye Ozay · 2018
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
Alex Kendall, Yarin Gal, and Roberto Cipolla · 2018
Earlier work this paper cites.
Textual explanations for self-driving vehicles
Jinkyu Kim, Anna Rohrbach, Trevor Darrell, John Canny, and Zeynep Akata · 2018
Earlier work this paper cites.
Modeling task relationships in multi-task learning with multi-gate mixture-of-experts
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H Chi · 2018
Earlier work this paper cites.
Multi-task learning as multi-objective optimization
Ozan Sener and Vladlen Koltun · 2018
Earlier work this paper cites.
Multinet: Real-time joint semantic reasoning for autonomous driving
Marvin Teichmann, Michael Weber, Marius Zoellner, Roberto Cipolla, and Raquel Urtasun · 2018
Earlier work this paper cites.
End-to-end multi-modal multi-task vehicle control for self-driving cars with visual perceptions
Zhengyuan Yang, Yixuan Zhang, Jerry Yu, Junjie Cai, and Jiebo Luo · 2018
Earlier work this paper cites.
Talk2car: Taking control of your self-driving car
Thierry Deruyttere, Simon Vandenhende, Dusan Grujicic, Luc Van Gool, and Marie-Francine Moens · 2019
Earlier work this paper cites.
Pareto multi-task learning
Xi Lin, Hui-Ling Zhen, Zhenhua Li, Qing-Fu Zhang, and Sam Kwong · 2019
Earlier work this paper cites.
End-to-end multi-task learning with attention
Shikun Liu, Edward Johns, and Andrew J Davison · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Actively semi-supervised deep rule-based classifier applied to adverse driving scenarios
Eduardo Soares, Plamen Angelov, Bruno Costa, and Marcos Castro · 2019
Earlier work this paper cites.
Training a binary weight object detector by knowledge transfer for autonomous driving
Jiaolong Xu, Yiming Nie, Peng Wang, and Antonio M López · 2019
Earlier work this paper cites.
Bidirectional encoder representations from transformers (bert): A sentiment analysis odyssey
Shivaji Alaparthi and Manit Mishra · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom · 2020
Earlier work this paper cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Earlier work this paper cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Earlier work this paper cites.
Big self-supervised models are strong semi-supervised learners
Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E Hinton · 2020
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges
Di Feng, Christian Haase-Schütz, Lars Rosenbaum, Heinz Hertlein, Claudius Glaeser, Fabian Timm, Werner Wiesbeck, and Klaus Dietmayer · 2020
Earlier work this paper cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Earlier work this paper cites.
Efficient continuous pareto exploration in multi-task learning
Pingchuan Ma, Tao Du, and Wojciech Matusik · 2020
Earlier work this paper cites.
Robust and efficient post-processing for video object detection
Alberto Sabater, Luis Montesano, and Ana C Murillo · 2020
Earlier work this paper cites.
Frustratingly simple few-shot object detection
Xin Wang, Thomas E Huang, Trevor Darrell, Joseph E Gonzalez, and Fisher Yu · 2020
Earlier work this paper cites.
Beit: Bert pre-training of image transformers. arxiv 2021
H Bao, L Dong, and F Wei · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Multi-task learning with attention for end-to-end autonomous driving
Keishi Ishihara, Anssi Kanervisto, Jun Miura, and Ville Hautamaki · 2021
Earlier work this paper cites.
Autonomy 2.0: Why is self-driving always 5 years away?
Ashesh Jain, Luca Del Pero, Hugo Grimmett, and Peter Ondruska · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Earlier work this paper cites.
Multi-modal fusion transformer for end-to-end autonomous driving
Aditya Prakash, Kashyap Chitta, and Andreas Geiger · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Trafficsim: Learning to simulate realistic multi-agent behaviors
Simon Suo, Sebastian Regalado, Sergio Casas, and Raquel Urtasun · 2021
Earlier work this paper cites.
Center-based 3d object detection and tracking
Tianwei Yin, Xingyi Zhou, and Philipp Krahenbuhl · 2021
Earlier work this paper cites.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Earlier work this paper cites.
Learning pedestrian group representations for multi-modal trajectory prediction
Inhwan Bae, Jin-Hwi Park, and Hae-Gon Jeon · 2022
Earlier work this paper cites.
Transfusion: Robust lidar-camera fusion for 3d object detection with transformers
Xuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang, Yilun Chen, Hongbo Fu, and Chiew-Lan Tai · 2022
Earlier work this paper cites.
Quantized convolutional neural networks through the lens of partial differential equations
Ido Ben-Yair, Gil Ben Shalom, Moshe Eliasof, and Eran Treister · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al · 2022
Earlier work this paper cites.
Transfuser: Imitation with transformer-based sensor fusion for autonomous driving
Kashyap Chitta, Aditya Prakash, Bernhard Jaeger, Zehao Yu, Katrin Renz, and Andreas Geiger · 2022
Earlier work this paper cites.
Davit: Dual attention vision transformers
Mingyu Ding, Bin Xiao, Noel Codella, Ping Luo, Jingdong Wang, and Lu Yuan · 2022
Earlier work this paper cites.
King: Generating safety-critical driving scenarios for robust imitation via kinematics gradients
Niklas Hanselmann, Katrin Renz, Kashyap Chitta, Apratim Bhattacharyya, and Andreas Geiger · 2022
Earlier work this paper cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Earlier work this paper cites.
Good: Exploring geometric cues for detecting objects in an open world
Haiwen Huang, Andreas Geiger, and Dan Zhang · 2022
Earlier work this paper cites.
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27
Yann LeCun · 2022
Earlier work this paper cites.
Coda: A real-world road corner case dataset for object detection in autonomous driving
Kaican Li, Kai Chen, Haoyu Wang, Lanqing Hong, Chaoqiang Ye, Jianhua Han, Yukuai Chen, Wei Zhang, Chunjing Xu, Dit-Yan Yeung, et al · 2022
Earlier work this paper cites.
Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers
Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie, Chonghao Sima, Tong Lu, Yu Qiao, and Jifeng Dai · 2022
Earlier work this paper cites.
Edge yolo: Real-time intelligent object detection system based on edge-cloud cooperation in autonomous vehicles
Siyuan Liang, Hao Wu, Li Zhen, Qiaozhi Hua, Sahil Garg, Georges Kaddoum, Mohammad Mehedi Hassan, and Keping Yu · 2022
Earlier work this paper cites.
Brain-inspired domain-incremental adaptive detection for autonomous driving
Weihao Liang, Lu Gan, Pengfei Wang, and Wei Meng · 2022
Earlier work this paper cites.
Effective adaptation in multi-task co-training for unified autonomous driving
Xiwen Liang, Yangxin Wu, Jianhua Han, Hang Xu, Chunjing Xu, and Xiaodan Liang · 2022
Earlier work this paper cites.
An efficient domain-incremental learning approach to drive in all weather conditions
M Jehanzeb Mirza, Marc Masana, Horst Possegger, and Horst Bischof · 2022
Earlier work this paper cites.
Chatgpt: Optimizing language models for dialogue. openai, 2022
TB OpenAI · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
A survey of autonomous driving scenarios and scenario databases
Hongping Ren, Hui Gao, He Chen, and Guangzhen Liu · 2022
Earlier work this paper cites.
K-lite: Learning transferable visual models with external knowledge
Sheng Shen, Chunyuan Li, Xiaowei Hu, Yujia Xie, Jianwei Yang, Pengchuan Zhang, Zhe Gan, Lijuan Wang, Lu Yuan, Ce Liu, et al · 2022
Earlier work this paper cites.
Masked feature prediction for self-supervised visual pre-training
Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Yolop: You only look once for panoptic driving perception
Dong Wu, Man-Wen Liao, Wei-Tian Zhang, Xing-Gang Wang, Xiang Bai, Wen-Qing Cheng, and Wen-Yu Liu · 2022
Earlier work this paper cites.
Simmim: A simple framework for masked image modeling
Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu · 2022
Earlier work this paper cites.
Cae v2: Context autoencoder with clip target
Xinyu Zhang, Jiahui Chen, Junkun Yuan, Qiang Chen, Jian Wang, Xiaodi Wang, Shumin Han, Xiaokang Chen, Jimin Pi, Kun Yao, et al · 2022
Earlier work this paper cites.
Drivemlm: Aligning multi-modal large language models with behavioral planning states for autonomous driving
Anonymous · 2023
Earlier work this paper cites.
Self-supervised learning from images with a joint-embedding predictive architecture
Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas · 2023
Earlier work this paper cites.
Explaining autonomous driving actions with visual question answering
Shahin Atakishiyev, Mohammad Salameh, Housam Babiker, and Randy Goebel · 2023
Earlier work this paper cites.
Openflamingo: An open-source framework for training large autoregressive vision-language models
Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, et al · 2023
Earlier work this paper cites.
Sequential modeling enables scalable learning for large vision models
Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan Yuille, Trevor Darrell, Jitendra Malik, and Alexei A Efros · 2023
Earlier work this paper cites.
Neeraja Bhide, Nanami Hashimoto, Kazimierz Dokurno, Chris Van der Hoorn, Sascha Hoogendoorn-Lanser, and Sina Nordhoff · 2023
Earlier work this paper cites.
Exploring the potential of world models for anomaly detection in autonomous driving
Daniel Bogdoll, Lukas Bosch, Tim Joseph, Helen Gremmelmaier, Yitian Yang, and J Marius Zöllner · 2023
Earlier work this paper cites.
Muvo: A multimodal generative world model for autonomous driving with geometric representations
Daniel Bogdoll, Yitian Yang, and J Marius Zöllner · 2023
Cited alongside, same era.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al · 2023
Cited alongside, same era.
Less is more: Removing text-regions improves clip training efficiency and robustness
Liangliang Cao, Bowen Zhang, Chen Chen, Yinfei Yang, Xianzhi Du, Wencong Zhang, Zhiyun Lu, and Yantao Zheng · 2023
Cited alongside, same era.
End-to-end autonomous driving: Challenges and frontiers
Li Chen, Penghao Wu, Kashyap Chitta, Bernhard Jaeger, Andreas Geiger, and Hongyang Li · 2023
Cited alongside, same era.
Cpsor-gcn: A vehicle trajectory prediction method powered by emotion and cognitive theory
Lanyue Tang, Y Li, Jinghui Yuan, A Fu, and J Sun · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Multi-camera bird’s eye view perception for autonomous driving
David Unger, Nikhil Gosala, Varun Ravi Kumar, Shubhankar Borse, Abhinav Valada, and Senthil Yogamani · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Long Chen, Oleg Sinavski, Jan Hünermann, Alice Karnsund, Andrew James Willmott, Danny Birch, Daniel Maund, and Jamie Shotton · 2023
Cited alongside, same era.
Clip2scene: Towards label-efficient 3d scene understanding by clip
Runnan Chen, Youquan Liu, Lingdong Kong, Xinge Zhu, Yuexin Ma, Yikang Li, Yuenan Hou, Yu Qiao, and Wenping Wang · 2023
Cited alongside, same era.
Context autoencoder for self-supervised representation learning
Xiaokang Chen, Mingyu Ding, Xiaodi Wang, Ying Xin, Shentong Mo, Yunhao Wang, Shumin Han, Ping Luo, Gang Zeng, and Jingdong Wang · 2023
Cited alongside, same era.
Voxelnext: Fully sparse voxelnet for 3d object detection and tracking
Yukang Chen, Jianhui Liu, Xiangyu Zhang, Xiaojuan Qi, and Jiaya Jia · 2023
Cited alongside, same era.
Large model based referring camouflaged object detection
Shupeng Cheng, Ge-Peng Ji, Pengda Qin, Deng-Ping Fan, Bowen Zhou, and Peng Xu · 2023
Cited alongside, same era.
Language-guided 3d object detection in point cloud for autonomous driving
Wenhao Cheng, Junbo Yin, Wei Li, Ruigang Yang, and Jianbing Shen · 2023
Cited alongside, same era.
idet3d: Towards efficient interactive object detection for lidar point clouds
Dongmin Choi, Wonwoo Cho, Kangyeol Kim, and Jaegul Choo · 2023
Cited alongside, same era.
Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, and Ziran Wang · 2023
Cited alongside, same era.
Mathijs R van Geerenstein, Felicia Ruppel, Klaus Dietmayer, and Dariu M Gavrila · 2023
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2023
Later among the works it cites.
You only look at once for real-time and generic multi-task
Jiayuan Wang, QM Wu, and Ning Zhang · 2023
Later among the works it cites.
Pengqin Wang, Meixin Zhu, Hongliang Lu, Hui Zhong, Xianda Chen, Shaojie Shen, Xuesong Wang, and Yinhai Wang · 2023
Later among the works it cites.
Drive anywhere: Generalizable end-to-end autonomous driving with multi-modal foundation models
Tsun-Hsuan Wang, Alaa Maalouf, Wei Xiao, Yutong Ban, Alexander Amini, Guy Rosman, Sertac Karaman, and Daniela Rus · 2023
Later among the works it cites.
Internimage: Exploring large-scale vision foundation models with deformable convolutions
Wenhai Wang, Jifeng Dai, Zhe Chen, Zhenhang Huang, Zhiqi Li, Xizhou Zhu, Xiaowei Hu, Tong Lu, Lewei Lu, Hongsheng Li, et al · 2023
Later among the works it cites.
Drivedreamer: Towards real-world-driven world models for autonomous driving
Xiaofeng Wang, Zheng Zhu, Guan Huang, Xinze Chen, and Jiwen Lu · 2023
Later among the works it cites.
Empowering autonomous driving with large language models: A safety perspective
Yixuan Wang, Ruochen Jiao, Chengtian Lang, Sinong Simon Zhan, Chao Huang, Zhaoran Wang, Zhuoran Yang, and Qi Zhu · 2023
Later among the works it cites.
Panoocc: Unified occupancy representation for camera-based 3d panoptic segmentation
Yuqi Wang, Yuntao Chen, Xingyu Liao, Lue Fan, and Zhaoxiang Zhang · 2023
Later among the works it cites.
Yuqi Wang, Jiawei He, Lue Fan, Hongxin Li, Yuntao Chen, and Zhaoxiang Zhang · 2023
Later among the works it cites.
Lingo-1: Exploring natural language for autonomous driving
Wayve · 2023
Later among the works it cites.
Dilu: A knowledge-driven approach to autonomous driving with large language models
Licheng Wen, Daocheng Fu, Xin Li, Xinyu Cai, Tao Ma, Pinlong Cai, Min Dou, Botian Shi, Liang He, and Yu Qiao · 2023
Later among the works it cites.
On the road with gpt-4v (ision): Early explorations of visual-language model on autonomous driving
Licheng Wen, Xuemeng Yang, Daocheng Fu, Xiaofeng Wang, Pinlong Cai, Xin Li, Tao Ma, Yingxuan Li, Linran Xu, Dengke Shang, et al · 2023
Later among the works it cites.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan · 2023
Later among the works it cites.
Referring multi-object tracking
Dongming Wu, Wencheng Han, Tiancai Wang, Xingping Dong, Xiangyu Zhang, and Jianbing Shen · 2023
Later among the works it cites.
Language prompt for autonomous driving
Dongming Wu, Wencheng Han, Tiancai Wang, Yingfei Liu, Xiangyu Zhang, and Jianbing Shen · 2023
Later among the works it cites.
Mars: An instance-aware, modular and realistic simulator for autonomous driving
Zirui Wu, Tianyu Liu, Liyi Luo, Zhide Zhong, Jianteng Chen, Hongmin Xiao, Chao Hou, Haozhe Lou, Yuantao Chen, Runyi Yang, et al · 2023
Later among the works it cites.
Calibformer: A transformer-based automatic lidar-camera calibration network
Yuxuan Xiao, Yao Li, Chengzhen Meng, Xingchen Li, and Yanyong Zhang · 2023
Later among the works it cites.
Drivinggaussian: Composite gaussian splatting for surrounding dynamic autonomous driving scenes
Zhou Xiaoyu, Lin Zhiwei, Shan Xiaojun, Wang Yongtao, Sun Deqing, and Yang Ming-Hsuan · 2023
Later among the works it cites.
Learning to adapt sam for segmenting cross-domain point clouds
Peng Xidong, Chen Runnan, Qiao Feng, Kong Lingdong, Liu Youquan, Wang Tai, Zhu Xinge, and Ma Yuexin · 2023
Later among the works it cites.
Sparsefusion: Fusing multi-modal sparse representations for multi-sensor 3d object detection
Yichen Xie, Chenfeng Xu, Marie-Julie Rakotosaona, Patrick Rim, Federico Tombari, Kurt Keutzer, Masayoshi Tomizuka, and Wei Zhan · 2023
Later among the works it cites.
Qinghua Xu, Tao Yue, Shaukat Ali, and Maite Arratibel · 2023
Later among the works it cites.
Frnet: Frustum-range networks for scalable lidar segmentation
Xiang Xu, Lingdong Kong, Hui Shuai, and Qingshan Liu · 2023
Later among the works it cites.
Drivegpt4: Interpretable end-to-end autonomous driving via large language model
Zhenhua Xu, Yujia Zhang, Enze Xie, Zhen Zhao, Yong Guo, Kenneth KY Wong, Zhenguo Li, and Hengshuang Zhao · 2023
Later among the works it cites.
Traffic sign interpretation in real road scene
Chuang Yang, Kai Zhuang, Mulin Chen, Haozhao Ma, Xu Han, Tao Han, Changxing Guo, Han Han, Bingxuan Zhao, and Qi Wang · 2023
Later among the works it cites.
Traffic sign interpretation in real road scene
Chuang Yang, Kai Zhuang, Mulin Chen, Haozhao Ma, Xu Han, Tao Han, Changxing Guo, Han Han, Bingxuan Zhao, and Qi Wang · 2023
Later among the works it cites.
Unipad: A universal pre-training paradigm for autonomous driving
Honghui Yang, Sha Zhang, Di Huang, Xiaoyang Wu, Haoyi Zhu, Tong He, Shixiang Tang, Hengshuang Zhao, Qibo Qiu, Binbin Lin, et al · 2023
Later among the works it cites.
Lidar-llm: Exploring the potential of large language models for 3d lidar understanding
Senqiao Yang, Jiaming Liu, Ray Zhang, Mingjie Pan, Zoey Guo, Xiaoqi Li, Zehui Chen, Peng Gao, Yandong Guo, and Shanghang Zhang · 2023
Later among the works it cites.
Unisim: A neural closed-loop sensor simulator
Ze Yang, Yun Chen, Jingkang Wang, Sivabalan Manivasagam, Wei-Chiu Ma, Anqi Joyce Yang, and Raquel Urtasun · 2023
Later among the works it cites.
A survey of large language models for autonomous driving
Zhenjie Yang, Xiaosong Jia, Hongyang Li, and Junchi Yan · 2023
Later among the works it cites.
Taskprompter: Spatial-channel multi-task prompting for dense scene understanding
Hanrong Ye and Dan Xu · 2023
Later among the works it cites.
Contextual object detection with multimodal large language models
Yuhang Zang, Wei Li, Jun Han, Kaiyang Zhou, and Chen Change Loy · 2023
Later among the works it cites.
Bo Zhang, Xinyu Cai, Jiakang Yuan, Donglin Yang, Jianfei Guo, Renqiu Xia, Botian Shi, Min Dou, Tao Chen, Si Liu, et al · 2023
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Hang Zhang, Xin Li, and Lidong Bing · 2023
Later among the works it cites.
Opensight: A simple open-vocabulary framework for lidar-based object detection
Hu Zhang, Jianhua Xu, Tao Tang, Haiyang Sun, Xin Yu, Zi Huang, and Kaicheng Yu · 2023
Later among the works it cites.
Nerf-lidar: Generating realistic lidar point clouds with neural radiance fields
Junge Zhang, Feihu Zhang, Shaochen Kuang, and Li Zhang · 2023
Later among the works it cites.
Adaptive budget allocation for parameter-efficient fine-tuning
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao · 2023
Later among the works it cites.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao · 2023
Later among the works it cites.
Trafficgpt: Viewing, processing and interacting with traffic foundation models
Siyao Zhang, Daocheng Fu, Zhao Zhang, Bin Yu, and Pinlong Cai · 2023
Later among the works it cites.
Meta-transformer: A unified framework for multimodal learning
Yiyuan Zhang, Kaixiong Gong, Kaipeng Zhang, Hongsheng Li, Yu Qiao, Wanli Ouyang, and Xiangyu Yue · 2023
Later among the works it cites.
TrafficBots: Towards world models for autonomous driving simulation and motion prediction
Zhejun Zhang, Alexander Liniger, Dengxin Dai, Fisher Yu, and Luc Van Gool · 2023
Later among the works it cites.
Occworld: Learning a 3d occupancy world model for autonomous driving
Wenzhao Zheng, Weiliang Chen, Yuanhui Huang, Borui Zhang, Yueqi Duan, and Jiwen Lu · 2023
Later among the works it cites.
Vision language models in autonomous driving and intelligent transportation systems
Xingcheng Zhou, Mingyu Liu, Bare Luka Zagar, Ekim Yurtsever, and Alois C Knoll · 2023
Later among the works it cites.
Openannotate3d: Open-vocabulary auto-labeling system for multi-modal 3d data
Yijie Zhou, Likun Cai, Xianhui Cheng, Zhongxue Gan, Xiangyang Xue, and Wenchao Ding · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Later among the works it cites.
Open world object detection in the era of foundation models
Orr Zohar, Alejandro Lozano, Shelly Goel, Serena Yeung, and Kuan-Chieh Wang · 2023
Later among the works it cites.
Multimodal prompt transformer with hybrid contrastive learning for emotion recognition in conversation
Shihao Zou, Xianying Huang, and Xudong Shen · 2023
Later among the works it cites.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Homes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Wing Yin Ng, Ricky Wang, and Aditya Ramesh · 2024
Closest in time.
S-nerf++: Autonomous driving simulation via neural reconstruction and generation
Yurui Chen, Junge Zhang, Ziyang Xie, Wenye Li, Feihu Zhang, Jiachen Lu, and Li Zhang · 2024
Closest in time.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2024
Closest in time.
A survey for foundation models in autonomous driving
Haoxiang Gao, Yaqian Li, Kaiwen Long, Ming Yang, and Yiqing Shen · 2024
Closest in time.
Vanishing-point-guided video semantic segmentation of driving scenes
Diandian Guo, Deng-Ping Fan, Tongyu Lu, Christos Sakaridis, and Luc Van Gool · 2024
Closest in time.
Lynn Vonder Haar, Timothy Elvira, Luke Newcomb, and Omar Ochoa · 2024
Closest in time.
Lidarclip or: How i learned to talk to point clouds
Georg Hess, Adam Tonderski, Christoffer Petersson, Kalle Åström, and Lennart Svensson · 2024
Closest in time.
Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations
Yuichi Inoue, Yuki Yada, Kotaro Tanahashi, and Yu Yamaguchi · 2024
Closest in time.
Real-time traffic object detection for autonomous driving
Abdul Hannan Khan, Syed Tahseen Raza Rizvi, and Andreas Dengel · 2024
Closest in time.
Bayesian multi-task transfer learning for soft prompt tuning
Haeju Lee, Minchan Jeong, Se-Young Yun, and Kee-Eung Kim · 2024
Closest in time.
World model on million-length video and language with ringattention
Hao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel · 2024
Closest in time.
Driveworld: 4d pre-trained scene understanding via world models for autonomous driving
Chen Min, Dawei Zhao, Liang Xiao, Jian Zhao, Xinli Xu, Zheng Zhu, Lei Jin, Jianshu Li, Yulan Guo, Junliang Xing, Liping Jing, Yiming Nie, and Bin Dai · 2024
Closest in time.
Rank2tell: A multimodal driving dataset for joint importance ranking and reasoning
Enna Sachdeva, Nakul Agarwal, Suhas Chundi, Sean Roelofs, Jiachen Li, Mykel Kochenderfer, Chiho Choi, and Behzad Dariush · 2024
Closest in time.
Chain-of-instructions: Compositional instruction tuning on large language models
Hayati Shirley Anugrah, Jung Taehee, Bodding-Long Tristan, Kar Sudipta, Sethy Abhinav, Kim Joo-Kyung, and Kang Dongyeop · 2024
Closest in time.
Lip-loc: Lidar image pretraining for cross-modal localization
Sai Shubodh, Mohammad Omama, Husain Zaidi, Udit Singh Parihar, and Madhava Krishna · 2024
Closest in time.
Ziying Song, Guoxin Zhang, Jun Xie, Lin Liu, Caiyan Jia, Shaoqing Xu, and Zhepeng Wang · 2024
Closest in time.
Robofusion: Towards robust multi-modal 3d obiect detection via sam
Ziying Song, Guoxing Zhang, Lin Liu, Lei Yang, Shaoqing Xu, Caiyan Jia, Feiyang Jia, and Li Wang · 2024
Closest in time.
Text2street: Controllable text-to-image generation for street views
Jinming Su, Songen Gu, Yiting Duan, Xingyue Chen, and Junfeng Luo · 2024
Closest in time.
Drivevlm: The convergence of autonomous driving and large vision-language models
Xiaoyu Tian, Junru Gu, Bailin Li, Yicheng Liu, Chenxu Hu, Yang Wang, Kun Zhan, Peng Jia, Xianpeng Lang, and Hang Zhao · 2024
Closest in time.
Off-road lidar intensity based semantic segmentation
Kasi Viswanath, Peng Jiang, Sujit PB, and Srikanth Saripalli · 2024
Closest in time.
Revisiting the power of prompt for visual tuning
Yuzhu Wang, Lechao Cheng, Chaowei Fang, Dingwen Zhang, Manni Duan, and Meng Wang · 2024
Closest in time.
Bev-clip: Multi-modal bev retrieval methodology for complex scene in autonomous driving
Dafeng Wei, Tian Gao, Zhengyu Jia, Changwei Cai, Chengkai Hou, Peng Jia, Fu Liu, Kun Zhan, Jingchen Fan, Yixing Zhao, et al · 2024
Closest in time.
Editable scene simulation for autonomous driving via collaborative llm-agents
Yuxi Wei, Zi Wang, Yifan Lu, Chenxin Xu, Changxing Liu, Hao Zhao, Siheng Chen, and Yanfeng Wang · 2024
Closest in time.
Liu Weiwei, Hu Wenxuan, Jing Wei, Lei Lanxin, Gao Lingping, and Liu Yong · 2024
Closest in time.
Xu Yan, Haiming Zhang, Yingjie Cai, Jingming Guo, Weichao Qiu, Bin Gao, Kaiqiang Zhou, Yue Zhao, Huan Jin, Jiantao Gao, et al · 2024
Closest in time.
Brian Yang, Huangyuan Su, Nikolaos Gkanatsios, Tsung-Wei Ke, Ayush Jain, Jeff Schneider, and Katerina Fragkiadaki · 2024
Closest in time.
Mixsup: Mixed-grained supervision for label-efficient lidar-based 3d object detection
Yuxue Yang, Lue Fan, and Zhaoxiang Zhang · 2024
Closest in time.
Low-resource vision challenges for foundation models
Yunhua Zhang, Hazel Doughty, and Cees GM Snoek · 2024
Closest in time.
Genad: Generative end-to-end autonomous driving
Wenzhao Zheng, Ruiqi Song, Xianda Guo, and Long Chen · 2024
Closest in time.
3d lane detection from front or surround-view using joint-modeling & matching
Haibin Zhou, Jun Chang, Tao Lu, and Huabing Zhou · 2024
Closest in time.
Lidar-ptq: Post-training quantization for point cloud 3d object detection
Sifan Zhou, Liang Li, Xinyu Zhang, Bo Zhang, Shipeng Bai, Miao Sun, Ziyu Zhao, Xiaobo Lu, and Xiangxiang Chu · 2024
Closest in time.
Embodied understanding of driving scenarios
Yunsong Zhou, Linyan Huang, Qingwen Bu, Jia Zeng, Tianyu Li, Hang Qiu, Hongzi Zhu, Minyi Guo, Yu Qiao, and Hongyang Li · 2024
Closest in time.
Rsud20k: A dataset for road scene understanding in autonomous driving
Hasib Zunair, Shakib Khan, and A Ben Hamza · 2024
Closest in time.