Fetching the paper…
Reading the bibliography…
With the emergence of Large Language Models (LLMs) and Vision Foundation Models (VFMs), multimodal AI systems benefiting from large models have the potential to equally perceive the real world, make decisions, and control tools as humans.
Language as an Abstraction for Hierarchical Deep Reinforcement Learning, Nov. 2019
Yiding Jiang, Shixiang Gu, Kevin Murphy, and Chelsea Finn · 1906
Earlier work this paper cites.
A Survey of Reinforcement Learning Informed by Natural Language, June 2019
Jelena Luketina, Nantas Nardelli, Gregory Farquhar, Jakob Foerster, Jacob Andreas, Edward Grefenstette, Shimon Whiteson, and Tim Rocktäschel · 1906
Earlier work this paper cites.
Scalability in Perception for Autonomous Driving: Waymo Open Dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, Vijay Vasudevan, Wei Han, Jiquan Ngiam, Hang Zhao, Aleksei Timofeev, Scott Ettinger, Maxim Krivokon, Amy Gao, Aditya Joshi, Sheng Zhao, Shuyang Cheng, Yu Zhang, Jonathon Shlens, Zhifeng Chen, and Dragomir Anguelov · 1912
Earlier work this paper cites.
The present status of automatic translation of languages
Yehoshua Bar-Hillel · 1960
Earlier work this paper cites.
Man-to-machine communication and automatic code translation
Anatol W Holt and WJ Turanski · 1960
Earlier work this paper cites.
Automatic language translation: Lexical and technical aspects, with particular reference to Russian
Anthony G Oettinger · 1960
Earlier work this paper cites.
Procedures as a Representation for Data in a Computer Program for Understanding Natural Language
Terry Winograd · 1971
Earlier work this paper cites.
Autonomous land vehicle project at CMU
Takeo Kanade, Chuck Thorpe, and William Whittaker · 1986
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1988
Earlier work this paper cites.
Class-based n-gram models of natural language
Peter F Brown, Vincent J Della Pietra, Peter V Desouza, Jennifer C Lai, and Robert L Mercer · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K Paliwal · 1997
Earlier work this paper cites.
The hierarchical hidden markov model: Analysis and applications
Shai Fine, Yoram Singer, and Naftali Tishby · 1998
Earlier work this paper cites.
Blocks world revisited
John Slaney and Sylvie Thiébaux · 2001
Earlier work this paper cites.
Language Conditioned Imitation Learning over Unstructured Data, July 2021
Corey Lynch and Pierre Sermanet · 2005
Earlier work this paper cites.
Stanley: The robot that won the darpa grand challenge
Sebastian Thrun, Mike Montemerlo, Hendrik Dahlkamp, David Stavens, Andrei Aron, James Diebel, Philip Fong, John Gale, Morgan Halpenny, Gabriel Hoffmann, Kenny Lau, Celia Oakley, Mark Palatucci, Vaughan Pratt, Pascal Stang, Sven Strohband, Cedric Dupont, Lars-Erik Jendrossek, Christian Koelen, Charles Markey, Carlo Rummel, Joe van Niekerk, Eric Jensen, Philippe Alessandrini, Gary Bradski, Bob Davies, Scott Ettinger, Adrian Kaehler, Ara Nefian, and Pamela Mahoney · 2006
Earlier work this paper cites.
Prasoon Goyal, Scott Niekum, and Raymond J. Mooney · 2007
Earlier work this paper cites.
Toward understanding natural language directions
Thomas Kollar, Stefanie Tellex, Deb Roy, and Nicholas Roy · 2010
Earlier work this paper cites.
Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge · 2010
Earlier work this paper cites.
Towards fully autonomous driving: Systems and algorithms
Jesse Levinson, Jake Askeland, Jan Becker, Jennifer Dolson, David Held, Soeren Kammel, J. Zico Kolter, Dirk Langer, Oliver Pink, Vaughan Pratt, Michael Sokolsky, Ganymed Stanek, David Stavens, Alex Teichman, Moritz Werling, and Sebastian Thrun · 2011
Earlier work this paper cites.
Are we ready for autonomous driving? The KITTI vision benchmark suite
A. Geiger, P. Lenz, and R. Urtasun · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Three decades of driver assistance systems: Review and future perspectives
Klaus Bengler, Klaus Dietmayer, Berthold Farber, Markus Maurer, Christoph Stiller, and Hermann Winner · 2014
Earlier work this paper cites.
Aspects of the Theory of Syntax
Noam Chomsky · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems
On-Road Automated Driving (ORAD) Committee · 2014
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, and Alan Yuille · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
An open approach to autonomous vehicles
Shinpei Kato, Eijiro Takeuchi, Yoshio Ishiguro, Yoshiki Ninomiya, Kazuya Takeda, and Tsuyoshi Hamada · 2015
Earlier work this paper cites.
Deep multimodal learning for audio-visual speech recognition
Youssef Mroueh, Tom Sercu, and Vaibhava Goel · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
Online Prediction of Driver Distraction Based on Brain Activity Patterns
Shouyi Wang, Yiqi Zhang, Changxu Wu, Felix Darvas, and Wanpracha Art Chaovalitwongse · 2015
Earlier work this paper cites.
EEG-Based Attention Tracking During Distracted Driving
Yu-Kai Wang, Tzyy-Ping Jung, and Chin-Teng Lin · 2015
Earlier work this paper cites.
Volumetric and multi-view cnns for object classification on 3d data
Charles R Qi, Hao Su, Matthias Niessner, Angela Dai, Mengyuan Yan, and Leonidas J Guibas · 2016
Earlier work this paper cites.
Learning with Latent Language, 2017
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelović and Andrew Zisserman · 2017
Earlier work this paper cites.
Mapping Instructions and Visual Observations to Actions with Reinforcement Learning, July 2017
Dipendra Misra, John Langford, and Yoav Artzi · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation, 2017
Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Autoware on board: Enabling autonomous vehicles with embedded systems
Shinpei Kato, Shota Tokunaga, Yuya Maruyama, Seiya Maeda, Manato Hirabayashi, Yuki Kitsukawa, Abraham Monrroy, Tomohito Ando, Yusuke Fujii, and Takuya Azumi · 2018
Earlier work this paper cites.
Textual explanations for self-driving vehicles
Jinkyu Kim, Anna Rohrbach, Trevor Darrell, John Canny, and Zeynep Akata · 2018
Earlier work this paper cites.
An Environment for Autonomous Driving Decision-Making, 2018
Edouard Leurent · 2018
Earlier work this paper cites.
Cooperation of v2i/p2i communication and roadside radar perception for the safety of vulnerable road users
Weijie Liu, Shintaro Muramatsu, and Yoshiyuki Okubo · 2018
Earlier work this paper cites.
Spatial as deep: Spatial cnn for traffic scene understanding
Xingang Pan, Jianping Shi, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles
SAE On-Road Automated Vehicle Standards Committee and others · 2018
Earlier work this paper cites.
Generating adversarial driving scenarios in high-fidelity simulators
Yasasa Abeysirigoonawardena, Florian Shkurti, and Gregory Dudek · 2019
Earlier work this paper cites.
Argoverse: 3D Tracking and Forecasting With Rich Maps
Ming-Fang Chang, John Lambert, Patsorn Sangkloy, Jagjeet Singh, Slawomir Bak, Andrew Hartnett, De Wang, Peter Carr, Simon Lucey, Deva Ramanan, and James Hays · 2019
Earlier work this paper cites.
A review on safety failures, security attacks, and available countermeasures for autonomous vehicles
Jin Cui, Lin Shen Liew, Giedre Sabaliauskaite, and Fengjun Zhou · 2019
Earlier work this paper cites.
Talk2Car: Taking Control of Your Self-Driving Car
Thierry Deruyttere, Simon Vandenhende, Dusan Grujicic, Luc Van Gool, and Marie-Francine Moens · 2019
Earlier work this paper cites.
Learning to drive in a day
Alex Kendall, Jeffrey Hawke, David Janz, Przemyslaw Mazur, Daniele Reda, John-Mark Allen, Vinh-Dieu Lam, Alex Bewley, and Amar Shah · 2019
Earlier work this paper cites.
Pointpillars: Fast encoders for object detection from point clouds, 2019
Alex H. Lang, Sourabh Vora, Holger Caesar, Lubing Zhou, Jiong Yang, and Oscar Beijbom · 2019
Earlier work this paper cites.
VisualBERT: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Earlier work this paper cites.
VilBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Talk to the vehicle: Language conditioned autonomous navigation of self driving cars
N. N. Sriram, Tirth Maniar, Jayaganesh Kalyanasundaram, Vineet Gandhi, Brojeshwar Bhowmick, and K Madhava Krishna · 2019
Earlier work this paper cites.
Driver Activity Recognition for Intelligent Vehicles: A Deep Learning Approach
Yang Xing, Chen Lv, Huaji Wang, Dongpu Cao, Efstathios Velenis, and Fei-Yue Wang · 2019
Earlier work this paper cites.
Self-supervised multimodal versatile networks
Jean-Baptiste Alayrac, Adrià Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Zeeuw, Hervé Jégou, and Andrew Zisserman · 2020
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
nuScenes: A Multimodal Dataset for Autonomous Driving
Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom · 2020
Earlier work this paper cites.
A survey of deep learning techniques for autonomous driving
Sorin Grigorescu, Bogdan Trasnea, Tiberiu Cocias, and Gigel Macesanu · 2020
Earlier work this paper cites.
Occuseg: Occupancy-aware 3d instance segmentation
Xinyu Han, Jianhui Lai, Kuiyuan Yang, Xiaojuan Li, Yujun Zhang, Dahua Lin, and Hao Zeng · 2020
Earlier work this paper cites.
Computer vision for autonomous vehicles: Problems, datasets and state of the art
Joel Janai, Fatma Güney, Aseem Behl, and Andreas Geiger · 2020
Earlier work this paper cites.
Models and algorithms for the exploration of the space of scenarios: toward the validation of the autonomous vehicle
Marc Nabhan · 2020
Earlier work this paper cites.
A v2x-based approach for avoiding potential blind-zone collisions between right-turning vehicles and pedestrians at intersections
Ying Ni, Shihan Wang, Liuyan Xin, Yiwei Meng, Juyuan Yin, and Jian Sun · 2020
Earlier work this paper cites.
Enhancing Driver Distraction Recognition Using Generative Adversarial Networks
Chaojie Ou and Fakhri Karray · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Robots That Use Language
Stefanie Tellex, Nakul Gopalan, Hadas Kress-Gazit, and Cynthia Matuszek · 2020
Cited alongside, same era.
A survey on cooperative longitudinal motion control of multiple connected and automated vehicles
Ziran Wang, Yougang Bian, Steven E. Shladover, Guoyuan Wu, Shengbo Eben Li, and Matthew J. Barth · 2020
Cited alongside, same era.
A survey of autonomous driving: Common practices and emerging technologies
Ekim Yurtsever, Jacob Lambert, Alexander Carballo, and Kazuya Takeda · 2020
Cited alongside, same era.
Actbert: Learning global-local video-text representations
Opt-iml: Scaling language model instruction meta learning through the lens of generalization, 2023
Srinivasan Iyer, Xi Victoria Lin, Ramakanth Pasunuru, Todor Mihaylov, Daniel Simig, Ping Yu, Kurt Shuster, Tianlu Wang, Qing Liu, Punit Singh Koura, Xian Li, Brian O’Horo, Gabriel Pereyra, Jeff Wang, Christopher Dewan, Asli Celikyilmaz, Luke Zettlemoyer, and Ves Stoyanov · 2023
Closest in time.
Ye Jin, Xiaoxi Shen, Huiling Peng, Xiaoan Liu, Jingli Qin, Jiayang Li, Jintao Xie, Peizhong Gao, Guyue Zhou, and Jiangtao Gong · 2023
Closest in time.
Aishwarya Kamath, Peter Anderson, Su Wang, Jing Yu Koh, Alexander Ku, Austin Waters, Yinfei Yang, Jason Baldridge, and Zarana Parekh · 2023
Closest in time.
Imagined subgoals for hierarchical goal-conditioned policies
Xuhui Kang, Wenqian Ye, and Yen-Ling Kuo · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Linjie Zhu, Jieyu Xu, Yi Yang, and Alexander G Hauptmann · 2020
Cited alongside, same era.
Multimodal safety-critical scenarios generation for decision-making algorithms evaluation
Wenhao Ding, Baiming Chen, Bo Li, Kim Ji Eun, and Ding Zhao · 2021
Cited alongside, same era.
Explainable AI: current status and future directions, July 2021
Prashant Gohel, Priyanka Singh, and Manoranjan Mohanty · 2021
Cited alongside, same era.
A survey of deep learning applications to autonomous vehicle control
Sampo Kuutti, Richard Bowden, Yaochu Jin, Phil Barber, and Saber Fallah · 2021
Cited alongside, same era.
Video prediction recalling long-term motion context via memory alignment learning
Sangmin Lee, Hak Gu Kim, Dae Hwi Choi, Hyung-Il Kim, and Yong Man Ro · 2021
Cited alongside, same era.
Prevent the language model from being overconfident in neural machine translation
Mengqi Miao, Fandong Meng, Yijin Liu, Xiao-Hua Zhou, and Jie Zhou · 2021
Cited alongside, same era.
Suraj Nair, Eric Mitchell, Kevin Chen, Brian Ichter, Silvio Savarese, and Chelsea Finn · 2021
Cited alongside, same era.
Can you text what is happening? integrating pre-trained language encoders into trajectory prediction models for autonomous driving, 2023
Ali Keysan, Andreas Look, Eitan Kosman, Gonca Gürsun, Jörg Wagner, Yu Yao, and Barbara Rakitsch · 2023
Closest in time.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al · 2023
Closest in time.
Large-scale text-to-image generation models for visual artists’ creative works
Hyung-Kwon Ko, Gwanmo Park, Hyeon Jeon, Jaemin Jo, Juho Kim, and Jinwook Seo · 2023
Closest in time.
In the eye of transformer: Global–local correlation for egocentric gaze estimation and beyond
Bolin Lai, Miao Liu, Fiona Ryan, and James M Rehg · 2023
Closest in time.
Listen to look into the future: Audio-visual egocentric gaze anticipation
Bolin Lai, Fiona Ryan, Wenqi Jia, Miao Liu, and James M Rehg · 2023
Closest in time.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Code as Policies: Language Model Programs for Embodied Control
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng · 2023
Closest in time.
Pie: Simulating disease progression via progressive image editing, 2023
Kaizhao Liang, Xu Cao, Kuei-Da Liao, Tianren Gao, Wenqian Ye, Zhengyu Chen, Jianguo Cao, Tejas Nama, and Jimeng Sun · 2023
Closest in time.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Closest in time.
Chameleon: Plug-and-play compositional reasoning with large language models
Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao · 2023
Closest in time.
ViT-DD: Multi-Task Vision Transformer for Semi-Supervised Driver Distraction Detection
Yunsheng Ma and Ziran Wang · 2023
Closest in time.
CEMFormer: Learning to Predict Driver Intentions from In-Cabin and External Cameras via Spatial-Temporal Transformers
Yunsheng Ma, Wenqian Ye, Xu Cao, Amr Abdelraouf, Kyungtae Han, Rohit Gupta, and Ziran Wang · 2023
Closest in time.
M2DAR: Multi-View Multi-Scale Driver Action Recognition with Vision Transformer
Yunsheng Ma, Liangqi Yuan, Amr Abdelraouf, Kyungtae Han, Rohit Gupta, Zihao Li, and Ziran Wang · 2023
Closest in time.
Drama: Joint risk localization and captioning in driving
Srikanth Malla, Chiho Choi, Isht Dwivedi, Joon Hee Choi, and Jiachen Li · 2023
Closest in time.
GPT-Driver: Learning to Drive with GPT, Oct. 2023
Jiageng Mao, Yuxi Qian, Hang Zhao, and Yue Wang · 2023
Closest in time.
The Waymo Open Sim Agents Challenge, July 2023
Nico Montali, John Lambert, Paul Mougin, Alex Kuefler, Nick Rhinehart, Michelle Li, Cole Gulino, Tristan Emrich, Zoey Yang, Shimon Whiteson, Brandyn White, and Dragomir Anguelov · 2023
Closest in time.
Model S Owner’s Manual [Online]
Tesla Motors · 2023
Closest in time.
Alt-pilot: Autonomous navigation with language augmented topometric maps, 2023
Mohammad Omama, Pranav Inani, Pranjal Paul, Sarat Chandra Yellapragada, Krishna Murthy Jatavallabhula, Sandeep Chinchali, and Madhava Krishna · 2023
Closest in time.
ChatGPT, 2023
OpenAI · 2023
Closest in time.
GPT-4 Technical Report, Mar. 2023
OpenAI · 2023
Closest in time.
Gpt-4v(ision) system card
OpenAI · 2023
Closest in time.
Proto-CLIP: Vision-Language Prototypical Network for Few-Shot Learning, July 2023
Jishnu Jaykumar P, Kamalesh Palanisamy, Yu-Wei Chao, Xinya Du, and Yu Xiang · 2023
Closest in time.
Ai art in architecture
Joern Ploennigs and Markus Berger · 2023
Closest in time.
Nuscenes-qa: A multi-modal visual question answering benchmark for autonomous driving scenario
Tianwen Qian, Jingjing Chen, Linhai Zhuo, Yang Jiao, and Yu-Gang Jiang · 2023
Closest in time.
Saynav: Grounding large language models for dynamic planning to navigation in new environments
Abhinav Rajvanshi, Karan Sikka, Xiao Lin, Bhoram Lee, Han-Pang Chiu, and Alvaro Velasquez · 2023
Closest in time.
Visual chain of thought: Bridging logical gaps with multimodal infillings, 2023
Daniel Rose, Vaishnavi Himakunthala, Andy Ouyang, Ryan He, Alex Mei, Yujie Lu, Michael Saxon, Chinmay Sonar, Diba Mirza, and William Yang Wang · 2023
Closest in time.
Motionlm: Multi-agent motion forecasting as language modeling
Ari Seff, Brian Cera, Dian Chen, Mason Ng, Aurick Zhou, Nigamaa Nayakanti, Khaled S Refaat, Rami Al-Rfou, and Benjamin Sapp · 2023
Closest in time.
Languagempc: Large language models as decision makers for autonomous driving
Hao Sha, Yao Mu, Yuxuan Jiang, Li Chen, Chenfeng Xu, Ping Luo, Shengbo Eben Li, Masayoshi Tomizuka, Wei Zhan, and Mingyu Ding · 2023
Closest in time.
Navigation with large language models: Semantic guesswork as a heuristic for planning
Dhruv Shah, Michael Equi, Blazej Osinski, Fei Xia, Brian Ichter, and Sergey Levine · 2023
Closest in time.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
Dhruv Shah, Błażej Osiński, brian ichter, and Sergey Levine · 2023
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang · 2023
Closest in time.
Progprompt: Generating situated robot task plans using large language models
Ishika Singh, Valts Blukis, Arsalan Mousavian, Ankit Goyal, Danfei Xu, Jonathan Tremblay, Dieter Fox, Jesse Thomason, and Animesh Garg · 2023
Closest in time.
Thma: Tencent hd map ai system for creating hd map annotations
Kun Tang, Xu Cao, Zhipeng Cao, Tong Zhou, Erlong Li, Ao Liu, Shengtao Zou, Chang Liu, Shuqi Mei, Elena Sizikova, et al · 2023
Closest in time.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality, 2023
The Vicuna Team · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models, Feb. 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Closest in time.
The Robot Hall of Fame
Carnegie Mellon University · 2023
Closest in time.
ChatGPT for Robotics: Design Principles and Model Abilities, July 2023
Sai Vemprala, Rogerio Bonatti, Arthur Bucker, and Ashish Kapoor · 2023
Closest in time.
Voyager: An Open-Ended Embodied Agent with Large Language Models, May 2023
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2023
Closest in time.
Caption anything: Interactive image description with diverse multimodal controls, 2023
Teng Wang, Jinrui Zhang, Junjie Fei, Hao Zheng, Yunlong Tang, Zhe Li, Mingqi Gao, and Shanshan Zhao · 2023
Closest in time.
LINGO-1: Exploring Natural Language for Autonomous Driving, Sept. 2023
Wayve · 2023
Closest in time.
Dilu: A knowledge-driven approach to autonomous driving with large language models
Licheng Wen, Daocheng Fu, Xin Li, Xinyu Cai, Tao Ma, Pinlong Cai, Min Dou, Botian Shi, Liang He, and Yu Qiao · 2023
Closest in time.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Chenfei Wu, Shengming Yin, Weizhen Qi, Xiaodong Wang, Zecheng Tang, and Nan Duan · 2023
Closest in time.
Language prompt for autonomous driving
Dongming Wu, Wencheng Han, Tiancai Wang, Yingfei Liu, Xiangyu Zhang, and Jianbing Shen · 2023
Closest in time.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Miao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li, Jie Fu, Junxian He, and Bryan Hooi · 2023
Closest in time.
V2V4Real: A Real-world Large-scale Dataset for Vehicle-to-Vehicle Cooperative Perception
Runsheng Xu, Xin Xia, Jinlong Li, Hanzhao Li, Shuo Zhang, Zhengzhong Tu, Zonglin Meng, Hao Xiang, Xiaoyu Dong, Rui Song, Hongkai Yu, Bolei Zhou, and Jiaqi Ma · 2023
Closest in time.
DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model, Oct. 2023
Zhenhua Xu, Yujia Zhang, Enze Xie, Zhen Zhao, Yong Guo, Kwan-Yee. K. Wong, Zhenguo Li, and Hengshuang Zhao · 2023
Closest in time.
Learning Interactive Real-World Simulators, Oct. 2023
Mengjiao Yang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Dale Schuurmans, and Pieter Abbeel · 2023
Closest in time.
Mm-react: Prompting chatgpt for multimodal reasoning and action, 2023
Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang · 2023
Closest in time.
mplug-owl: Modularization empowers large language models with multimodality, 2023
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, Chenliang Li, Yuanhong Xu, Hehong Chen, Junfeng Tian, Qian Qi, Ji Zhang, and Fei Huang · 2023
Closest in time.
Mitigating Transformer Overconfidence via Lipschitz Regularization
Wenqian Ye, Yunsheng Ma, Xu Cao, and Kun Tang · 2023
Closest in time.
A survey on multimodal large language models
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen · 2023
Closest in time.
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Hang Zhang, Xin Li, and Lidong Bing · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala · 2023
Closest in time.
Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners, 2023
Renrui Zhang, Xiangfei Hu, Bohao Li, Siyuan Huang, Hanqiu Deng, Hongsheng Li, Yu Qiao, and Peng Gao · 2023
Closest in time.
Multimodal chain-of-thought reasoning in language models, 2023
Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola · 2023
Closest in time.
High-definition map automatic annotation system based on active learning
Chao Zheng, Xu Cao, Kun Tang, Zhipeng Cao, Elena Sizikova, Tong Zhou, Erlong Li, Ao Liu, Shengtao Zou, Xinrui Yan, and Shuqi Mei · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny · 2023
Closest in time.
Drive as you speak: Enabling human-like interaction with large language models in autonomous vehicles
Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, and Ziran Wang · 2024
Closest in time.
Drive like a human: Rethinking autonomous driving with large language models
Daocheng Fu, Xin Li, Licheng Wen, Pinlong Cai, Botian Shi, and Yu Qiao · 2024
Closest in time.
Nuscenes-mqa: Integrated evaluation of captions and qa for autonomous driving datasets using markup annotations
Yuichi Inoue, Yuki Yada, Kotaro Tanahashi, and Yu Yamaguchi · 2024
Closest in time.
Vlaad: Vision and language assistant for autonomous driving
SungYeon Park, MinJae Lee, JiHyuk Kang, Hahyeon Choi, Yoonah Park, Juhwan Cho, Adam Lee, and Dong-Kyu Kim · 2024
Closest in time.
Lip-loc: Lidar image pretraining for cross-modal localization
Sai Shubodh, Mohammad Omama, Husain Zaidi, Udit Singh Parihar, and Madhava Krishna · 2024
Closest in time.
Human-centric autonomous systems with llms for user command reasoning
Yi Yang, Qingwen Zheng, Ci Li, Daniel L.S. Marta, Nazre Batool, and John Folkesson · 2024
Closest in time.
Latency driven spatially sparse optimization for multi-branch cnns for semantic segmentation
Giorgos Zampokas, Christos-Savvas Bouganis, and Dimitrios Tzovaras · 2024
Closest in time.
Safer vision-based autonomous planning system for quadrotor uavs with dynamic obstacle trajectory prediction
Jiageng Zhong, Ming Li, Yinliang Chen, Zihang Wei, Fan Yang, and Haoran Shen · 2024
Closest in time.