Fetching the paper…
Reading the bibliography…
The rise of multi-modal large language models(MLLMs) has spurred their applications in autonomous driving.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T.; Razeghi, Y.; Logan IV, R. L.; Wallace, E.; and Singh, S. 2020 · 2010
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
Ha, D.; and Schmidhuber, J. 2018 · 2018
Earlier work this paper cites.
Pointpillars: Fast encoders for object detection from point clouds
Lang, A. H.; Vora, S.; Caesar, H.; Zhou, L.; Yang, J.; and Beijbom, O. 2019 · 2019
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Caesar, H.; Bankiti, V.; Lang, A. H.; Vora, S.; Liong, V. E.; Xu, Q.; Krishnan, A.; Pan, Y.; Baldan, G.; and Beijbom, O. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P.; Rombach, R.; and Ommer, B. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; and Guo, B. 2021 · 2021
Earlier work this paper cites.
Dall-e: Creating images from text
Reddy, M. D. M.; Basha, M. S. M.; Hari, M. M. C.; and Penchalaiah, M. N. 2021 · 2021
Earlier work this paper cites.
Differentiable raycasting for self-supervised occupancy forecasting
Khurana, T.; Hu, P.; Dave, A.; Ziglar, J.; Held, D.; and Ramanan, D. 2022 · 2022
Earlier work this paper cites.
Muvo: A multimodal generative world model for autonomous driving with geometric representations
Bogdoll, D.; Yang, Y.; and Zöllner, J. M. 2023 · 2023
Earlier work this paper cites.
Talk2BEV: Language-enhanced Bird’s-eye View Maps for Autonomous Driving
Dewangan, V.; Choudhary, T.; Chandhok, S.; Priyadarshan, S.; Jain, A.; Singh, A. K.; Srivastava, S.; Jatavallabhula, K. M.; and Krishna, K. M. 2023 · 2023
Cited alongside, same era.
Ding, X.; Han, J.; Xu, H.; Zhang, W.; and Li, X. 2023 · 2023
Cited alongside, same era.
Keysan, A.; Look, A.; Kosman, E.; Gürsun, G.; Wagner, J.; Yu, Y.; and Rakitsch, B. 2023 · 2023
Cited alongside, same era.
Point cloud forecasting as a proxy for 4d occupancy forecasting
Khurana, T.; Hu, P.; Held, D.; and Ramanan, D. 2023 · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
Dilu: A knowledge-driven approach to autonomous driving with large language models
Wen, L.; Fu, D.; Li, X.; Cai, X.; Ma, T.; Cai, P.; Dou, M.; Shi, B.; He, L.; and Qiao, Y. 2023 · 2023
Later among the works it cites.
Lidar-llm: Exploring the potential of large language models for 3d lidar understanding
Yang, S.; Liu, J.; Zhang, R.; Pan, M.; Guo, Z.; Li, X.; Chen, Z.; Gao, P.; Guo, Y.; and Zhang, S. 2023 · 2023
Later among the works it cites.
Occworld: Learning a 3d occupancy world model for autonomous driving
Zheng, W.; Chen, W.; Huang, Y.; Zhang, B.; Duan, Y.; and Lu, J. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, X.; Zhang, Y.; and Ye, X. 2023 · 2023
Cited alongside, same era.
Fb-occ: 3d occupancy prediction based on forward-backward view transformation
Li, Z.; Yu, Z.; Austin, D.; Fang, M.; Lan, S.; Kautz, J.; and Alvarez, J. M. 2023 · 2023
Cited alongside, same era.
Mtd-gpt: A multi-task decision-making gpt model for autonomous driving at unsignalized intersections
Liu, J.; Hang, P.; Qi, X.; Wang, J.; and Sun, J. 2023 · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W.; and Xie, S. 2023 · 2023
Cited alongside, same era.
Drivelm: Driving with graph visual question answering
Sima, C.; Renz, K.; Chitta, K.; Chen, L.; Zhang, H.; Xie, C.; Luo, P.; Geiger, A.; and Li, H. 2023 · 2023
Cited alongside, same era.
Occ3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous Driving
Tian, X.; Jiang, T.; Yun, L.; Mao, Y.; Yang, H.; Wang, Y.; Wang, Y.; and Zhao, H. 2023 · 2023
Cited alongside, same era.
Scene as occupancy
Tong, W.; Sima, C.; Wang, T.; Chen, L.; Wu, S.; Deng, H.; Gu, Y.; Lu, L.; Luo, P.; Lin, D.; et al. 2023 · 2023
Cited alongside, same era.
Gaia-1: A generative world model for autonomous driving
Hu, A.; Russell, L.; Yeo, H.; Murez, Z.; Fedoseev, G.; Kendall, A.; Shotton, J.; and Corrado, G. 2023a
Cited in the paper.
Driving with llms: Fusing object-level vector modality for explainable autonomous driving
Chen, L.; Sinavski, O.; Hünermann, J.; Karnsund, A.; Willmott, A. J.; Birch, D.; Maund, D.; and Shotton, J. 2024 · 2024
Closest in time.
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2024 · 2024
Closest in time.
Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario
Qian, T.; Chen, J.; Zhuo, L.; Jiao, Y.; and Jiang, Y.-G. 2024 · 2024
Closest in time.
CarLLaVA: Vision language models for camera-only closed-loop driving
Renz, K.; Chen, L.; Marcu, A.-M.; Hünermann, J.; Hanotte, B.; Karnsund, A.; Shotton, J.; Arani, E.; and Sinavski, O. 2024 · 2024
Closest in time.
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Sun, P.; Jiang, Y.; Chen, S.; Zhang, S.; Peng, B.; Luo, P.; and Yuan, Z. 2024 · 2024
Closest in time.
Drivegpt4: Interpretable end-to-end autonomous driving via large language model
Xu, Z.; Zhang, Y.; Xie, E.; Zhao, Z.; Guo, Y.; Wong, K.-Y. K.; Li, Z.; and Zhao, H. 2024 · 2024
Closest in time.
BEVWorld: A Multimodal World Model for Autonomous Driving via Unified BEV Latent Space
Zhang, Y.; Gong, S.; Xiong, K.; Ye, X.; Tan, X.; Wang, F.; Huang, J.; Wu, H.; and Wang, H. 2024 · 2024
Closest in time.