Fetching the paper…
Reading the bibliography…
We study how vision-language models (VLMs) trained on web-scale data can be integrated into end-to-end driving systems to boost generalization and enable interactivity with human users.
Jelinek, F., Lafferty, J.D., Mercer, R.L.: Basic methods of probabilistic context free grammars. Springer, Berlin, Heidelberg (1992)
1992
Earlier work this paper cites.
Toomarian, N.B., Barhen, J.: Learning a trajectory using adjoint functions and teacher forcing. Neural networks (1992)
1992
Earlier work this paper cites.
Treiber, M., Hennecke, A., Helbing, D.: Congested traffic states in empirical observations and microscopic simulations. Physical Review E 62
2000
Earlier work this paper cites.
Treiber, M., Hennecke, A., Helbing, D.: Congested traffic states in empirical observations and microscopic simulations. Physical review E (2000)
2000
Earlier work this paper cites.
Papineni, K., Roukos, S., Ward, T., Zhu, W.J.: Bleu: a method for automatic evaluation of machine translation. In: ACL (2002)
2002
Earlier work this paper cites.
Macadam, C.C.: Understanding and modeling the human driver. Veh. Syst. Dyn (2003)
2003
Earlier work this paper cites.
Lin, C.Y.: Rouge: A package for automatic evaluation of summaries. In: ACL Workshop (2004)
2004
Earlier work this paper cites.
Lavie, A., Agarwal, A.: METEOR: An automatic metric for MT evaluation with high levels of correlation with human judgments. In: ACL Workshop (2007)
2007
Earlier work this paper cites.
Spelke, E.S., Kinzler, K.D.: Core knowledge. Dev Sci (2007)
2007
Earlier work this paper cites.
Marr, D.: Vision: A computational investigation into the human representation and processing of visual information. The MIT Press (2010)
2010
Earlier work this paper cites.
Groeger, J.A.: Understanding driving: Applying cognitive psychology to a complex everyday task. Routledge (2013)
2013
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: ECCV (2014)
2014
Earlier work this paper cites.
Vedantam, R., Lawrence Zitnick, C., Parikh, D.: Cider: Consensus-based image description evaluation. In: CVPR (2015)
2015
Earlier work this paper cites.
Anderson, P., Fernando, B., Johnson, M., Gould, S.: Spice: Semantic propositional image caption evaluation. In: ECCV (2016)
2016
Earlier work this paper cites.
Lamb, A.M., ALIAS PARTH GOYAL, A.G., Zhang, Y., Zhang, S., Courville, A.C., Bengio, Y.: Professor forcing: A new algorithm for training recurrent networks. In: NeurIPS (2016)
2016
Earlier work this paper cites.
Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., Koltun, V.: CARLA: An open urban driving simulator. In: CoRL (2017)
2017
Earlier work this paper cites.
Kim, J., Rohrbach, A., Darrell, T., Canny, J., Akata, Z.: Textual explanations for self-driving vehicles. In: ECCV (2018)
2018
Earlier work this paper cites.
Chen, Y., Rohrbach, M., Yan, Z., Shuicheng, Y., Feng, J., Kalantidis, Y.: Graph-based global reasoning networks. In: CVPR (2019)
2019
Earlier work this paper cites.
Hudson, D.A., Manning, C.D.: Gqa: A new dataset for real-world visual reasoning and compositional question answering. In: CVPR (2019)
2019
Earlier work this paper cites.
Kim, J., Misu, T., Chen, Y.T., Tawari, A., Canny, J.: Grounding human-to-vehicle advice for self-driving vehicles. In: CVPR (2019)
2019
Earlier work this paper cites.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I.: Language models are unsupervised multitask learners. OpenAI blog (2019)
2019
Earlier work this paper cites.
Shi, J., Zhang, H., Li, J.: Explainable and explicit visual reasoning over scene graphs. In: CVPR (2019)
2019
Earlier work this paper cites.
Wang, X., Wang, D., Xu, C., He, X., Cao, Y., Chua, T.S.: Explainable reasoning over knowledge graphs for recommendation. In: AAAI (2019)
2019
Earlier work this paper cites.
Akhauri, S., Zheng, L.Y., Lin, M.C.: Enhanced transfer learning for autonomous driving with systematic accident simulation. In: IROS (2020)
2020
Earlier work this paper cites.
Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., Beijbom, O.: nuScenes: A multimodal dataset for autonomous driving. In: CVPR (2020)
2020
Earlier work this paper cites.
Chen, X., Jia, S., Xiang, Y.: A review: Knowledge reasoning over knowledge graph. Expert Syst. Appl (2020)
2020
Earlier work this paper cites.
Floridi, L., Chiriatti, M.: GPT-3: Its nature, scope, limits, and consequences. MIND MACH (2020)
2020
Earlier work this paper cites.
Sun, P., Kretzschmar, H., Dotiwalla, X., Chouard, A., Patnaik, V., Tsui, P., Guo, J., Zhou, Y., Chai, Y., Caine, B., et al.: Scalability in perception for autonomous driving: Waymo open dataset. In: CVPR (2020)
2020
Earlier work this paper cites.
Tampuu, A., Matiisen, T., Semikin, M., Fishman, D., Muhammad, N.: A survey of end-to-end driving: Architectures and training methods. IEEE T-NNLS (2020)
2020
Earlier work this paper cites.
Caesar, H., Kabzan, J., Tan, K.S., Fong, W.K., Wolff, E.M., Lang, A.H., Fletcher, L., Beijbom, O., Omari, S.: nuplan: A closed-loop ml-based planning benchmark for autonomous vehicles. In: CVPR Workshops (2021)
2021
Earlier work this paper cites.
Chen, D., Koltun, V., Krähenbühl, P.: Learning to drive from a world on rails. In: ICCV (2021)
2021
Earlier work this paper cites.
Chitta, K., Prakash, A., Geiger, A.: Neat: Neural attention fields for end-to-end autonomous driving. In: ICCV (2021)
2021
Earlier work this paper cites.
Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.: LoRA: Low-rank adaptation of large language models. In: CoRL (2021)
2021
Earlier work this paper cites.
Prakash, A., Chitta, K., Geiger, A.: Multi-modal fusion transformer for end-to-end autonomous driving. In: CVPR (2021)
2021
Earlier work this paper cites.
Suo, S., Regalado, S., Casas, S., Urtasun, R.: Trafficsim: Learning to simulate realistic multi-agent behaviors. In: CVPR (2021)
2021
Earlier work this paper cites.
Wang, J., Pun, A., Tu, J., Manivasagam, S., Sadat, A., Casas, S., Ren, M., Urtasun, R.: Advsim: Generating safety-critical scenarios for self-driving vehicles. In: CVPR (2021)
2021
Earlier work this paper cites.
Wen, C., Lin, J., Qian, J., Gao, Y., Jayaraman, D.: Keyframe-focused visual imitation learning. In: ICML (2021)
2021
Cited alongside, same era.
Zareian, A., Rosa, K.D., Hu, D.H., Chang, S.F.: Open-vocabulary object detection using captions. In: CVPR (2021)
2021
Cited alongside, same era.
Chen, D., Krähenbühl, P.: Learning from all vehicles. In: CVPR (2022)
2022
Cited alongside, same era.
Database, A.I.: Incident 293: Cruise’s self-driving car involved in a multiple-injury collision at an san francisco intersection. https://incidentdatabase.ai/cite/293/ (2022)
2022
Cited alongside, same era.
Hanselmann, N., Renz, K., Chitta, K., Bhattacharyya, A., Geiger, A.: King: Generating safety-critical driving scenarios for robust imitation via kinematics gradients. In: ECCV (2022)
2022
Cited alongside, same era.
Li, Z., Yu, Z., Lan, S., Li, J., Kautz, J., Lu, T., Alvarez, J.M.: Is ego status all you need for open-loop end-to-end autonomous driving? (2023)
2023
Closest in time.
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., Zeng, A.: Code as policies: Language model programs for embodied control. In: ICRA (2023)
2023
Closest in time.
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. In: NeurIPS (2023)
2023
Closest in time.
Malla, S., Choi, C., Dwivedi, I., Choi, J.H., Li, J.: DRAMA: Joint risk localization and captioning in driving. In: WACV (2023)
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hu, S., Chen, L., Wu, P., Li, H., Yan, J., Tao, D.: St-p3: End-to-end vision-based autonomous driving via spatial-temporal feature learning. In: ECCV (2022)
2022
Cited alongside, same era.
Huang, W., Xia, F., Xiao, T., Chan, H., Liang, J., Florence, P., Zeng, A., Tompson, J., Mordatch, I., et al.: Inner monologue: Embodied reasoning through planning with language models. In: CoRL (2022)
2022
Cited alongside, same era.
Kojima, T., Gu, S.S., Reid, M., Matsuo, Y., Iwasawa, Y.: Large language models are zero-shot reasoners. In: NeurIPS (2022)
2022
Cited alongside, same era.
Li, K., Chen, K., Wang, H., Hong, L., Ye, C., Han, J., Chen, Y., Zhang, W., Xu, C., Yeung, D.Y., et al.: CODA: A real-world road corner case dataset for object detection in autonomous driving. In: ECCV (2022)
2022
Cited alongside, same era.
Li, Z., Wang, W., Li, H., Xie, E., Sima, C., Lu, T., Qiao, Y., Dai, J.: Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. In: ECCV (2022)
2022
Cited alongside, same era.
OpenAI: OpenAI: Introducing ChatGPT. https://openai.com/blog/chatgpt (2022)
2022
Cited alongside, same era.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., et al.: Training language models to follow instructions with human feedback. In: NeurIPS (2022)
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Rana, K., Haviland, J., Garg, S., Abou-Chakra, J., Reid, I., Suenderhauf, N.: SayPlan: Grounding large language models using 3d scene graphs for scalable robot task planning. In: CoRL (2023)
2023
Closest in time.
2023
Closest in time.
Seff, A., Cera, B., Chen, D., Ng, M., Zhou, A., Nayakanti, N., Refaat, K.S., Al-Rfou, R., Sapp, B.: MotionLM: Multi-agent motion forecasting as language modeling. In: ICCV (2023)
2023
Closest in time.
2023
Closest in time.
Shi, D., Tao, C., Rao, A., Yang, Z., Yuan, C., Wang, J.: Crossget: Cross-guided ensemble of tokens for accelerating vision-language transformers (2023)
2023
Closest in time.
Teng, S., Hu, X., Deng, P., Li, B., Li, Y., Ai, Y., Yang, D., Li, L., Xuanyuan, Z., Zhu, F., et al.: Motion planning for autonomous driving: The state of the art and future perspectives. IEEE T-IV (2023)
2023
Closest in time.
2023
Closest in time.
Wang, H., Li, T., Li, Y., Chen, L., Sima, C., Liu, Z., Wang, B., Jia, P., Wang, Y., Jiang, S., et al.: OpenLane-V2: A topology reasoning benchmark for unified 3d HD mapping. In: NeurIPS Datasets and Benchmarks (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., Zhou, D.: Self-Consistency improves chain of thought reasoning in language models. In: ICLR (2023)
2023
Closest in time.
Wayve: Lingo-1. https://wayve.ai/thinking/lingo-natural-language-autonomous-driving/ (2023)
2023
Closest in time.
Wu, D., Han, W., Wang, T., Dong, X., Zhang, X., Shen, J.: Referring Multi-Object tracking. In: CVPR (2023)
2023
Closest in time.
2023
Closest in time.
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., Han, S.: SmoothQuant: Accurate and efficient post-training quantization for large language models. In: Proceedings of the 40th International Conference on Machine Learning (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Zitkovich, B., Yu, T., Xu, S., Xu, P., Xiao, T., Xia, F., Wu, J., Wohlhart, P., Welker, S., et al.: RT-2: Vision-language-action models transfer web knowledge to robotic control. In: CoRL (2023)
2023
Closest in time.
Beißwenger, J.: PDM-Lite: A rule-based planner for carla leaderboard 2.0. https://github.com/OpenDriveLab/DriveLM/blob/DriveLM-CARLA/docs/report.pdf (2024)
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Xinpeng, D., Jinahua, H., Hang, X., Xiaodan, L., Xu, H., Wei, Z., Xiaomeng, L.: Holistic autonomous driving understanding by bird’s-eye-view injected multi-modal large models. In: CVPR (2024)
2024
Closest in time.