Fetching the paper…
Reading the bibliography…
Vision-and-Language Navigation (VLN) has gained increasing attention over recent years and many approaches have emerged to advance their development.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 1908
Earlier work this paper cites.
Cognitive maps in rats and men
Edward C Tolman · 1948
Earlier work this paper cites.
The hippocampus as a spatial map: preliminary evidence from unit activity in the freely-moving rat
John O’Keefe and Jonathan Dostrovsky · 1971
Earlier work this paper cites.
The hippocampus as a cognitive map
John O’keefe and Lynn Nadel · 1978
Earlier work this paper cites.
The organization of learning
Charles R Gallistel · 1990
Earlier work this paper cites.
The hcrc map task corpus
Anne H Anderson, Miles Bader, Ellen Gurman Bard, Elizabeth Boyle, Gwyneth Doherty, Simon Garrod, Stephen Isard, Jacqueline Kowtko, Jan McAllister, Jim Miller, et al · 1991
Earlier work this paper cites.
Navigational strategies and models
T Rodrigo · 2002
Earlier work this paper cites.
Cognitive load of navigating without vision when guided by virtual sound versus spatial language
Roberta L Klatzky, James R Marston, Nicholas A Giudice, Reginald G Golledge, and Jack M Loomis · 2006
Earlier work this paper cites.
Walk the talk: Connecting language, knowledge, and action in route instructions
Matt MacMahon, Brian Stankiewicz, and Benjamin Kuipers · 2006
Earlier work this paper cites.
Evidence from an emerging sign language reveals that language supports spatial cognition
Jennie E Pyers, Anna Shusterman, Ann Senghas, Elizabeth S Spelke, and Karen Emmorey · 2010
Earlier work this paper cites.
Children’s spatial thinking: Does talk about the spatial world matter?
Shannon M Pruden, Susan C Levine, and Janellen Huttenlocher · 2011
Earlier work this paper cites.
Cognitive effects of language on human navigation
Anna Shusterman, Sang Ah Lee, and Elizabeth S Spelke · 2011
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Crawling predicts infants’ understanding of agents’ navigation of obstacles
Rebecca J Brand, Kelly Escobar, Adrien Baranes, and Amanda Albu · 2015
Earlier work this paper cites.
The cognitive map in humans: spatial navigation and beyond
Russell A Epstein, Eva Zita Patai, Joshua B Julian, and Hugo J Spiers · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel · 2018
Earlier work this paper cites.
Navigating cognition: Spatial codes for human thinking
Jacob LS Bellmund, Peter Gärdenfors, Edvard I Moser, and Christian F Doeller · 2018
Earlier work this paper cites.
Matterport3d: Learning from rgb-d data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niebner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang · 2018
Earlier work this paper cites.
Talk the walk: Navigating new york city through grounded dialogue
Harm De Vries, Kurt Shuster, Dhruv Batra, Devi Parikh, Jason Weston, and Douwe Kiela · 2018
Earlier work this paper cites.
Speaker-follower models for vision-and-language navigation
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Earlier work this paper cites.
Using virtual environments to investigate wayfinding in 8-to 12-year-olds and adults
Jamie Lingwood, Mark Blades, Emily K Farran, Yannick Courbois, and Danielle Matthews · 2018
Earlier work this paper cites.
Learning to navigate in cities without a map
Piotr Mirowski, Matt Grimes, Mateusz Malinowski, Karl Moritz Hermann, Keith Anderson, Denis Teplyashin, Karen Simonyan, Andrew Zisserman, Raia Hadsell, et al · 2018
Earlier work this paper cites.
Mapping instructions to actions in 3d environments with visual goal prediction
Dipendra Misra, Andrew Bennett, Valts Blukis, Eyvind Niklasson, Max Shatkhin, and Yoav Artzi · 2018
Earlier work this paper cites.
Gibson env: Real-world perception for embodied agents
Fei Xia, Amir R Zamir, Zhiyang He, Alexander Sax, Jitendra Malik, and Silvio Savarese · 2018
Earlier work this paper cites.
Chasing ghosts: Instruction following as bayesian state tracking
Peter Anderson, Ayush Shrivastava, Devi Parikh, Dhruv Batra, and Stefan Lee · 2019
Earlier work this paper cites.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi · 2019
Earlier work this paper cites.
A comprehensive study for robot navigation techniques
Faiza Gul, Wan Rahiman, and Syed Sahal Nazli Alhady · 2019
Earlier work this paper cites.
General evaluation for instruction conditioned navigation using dynamic time warping
Gabriel Ilharco, Vihan Jain, Alexander Ku, Eugene Ie, and Jason Baldridge · 2019
Earlier work this paper cites.
Stay on the path: Instruction fidelity in vision-and-language navigation
Vihan Jain, Gabriel Magalhaes, Alexander Ku, Ashish Vaswani, Eugene Ie, and Jason Baldridge · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Earlier work this paper cites.
Robust navigation with language pretraining and stochastic sampling
Xiujun Li, Chunyuan Li, Qiaolin Xia, Yonatan Bisk, Asli Çelikyilmaz, Jianfeng Gao, Noah A. Smith, and Yejin Choi · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Self-monitoring navigation agent via auxiliary progress estimation
Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan AlRegib, Zsolt Kira, Richard Socher, and Caiming Xiong · 2019
Earlier work this paper cites.
Help, anna! visual navigation with natural multimodal assistance via retrospective curiosity-encouraging imitation learning
Khanh Nguyen and Hal Daumé III · 2019
Earlier work this paper cites.
Vision-based navigation with language-based assistance via imitation learning with indirect intervention
Khanh Nguyen, Debadeepta Dey, Chris Brockett, and Bill Dolan · 2019
Earlier work this paper cites.
RUN through the streets: A new dataset and baseline models for realistic urban navigation
Tzuf Paz-Argaman and Reut Tsarfaty · 2019
Earlier work this paper cites.
Habitat: A platform for embodied AI research
Manolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, and Vladlen Koltun · 2019
Earlier work this paper cites.
Talk to the vehicle: Language conditioned autonomous navigation of self driving cars
NN Sriram, Tirth Maniar, Jayaganesh Kalyanasundaram, Vineet Gandhi, Brojeshwar Bhowmick, and K Madhava Krishna · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Earlier work this paper cites.
Learning to navigate unseen environments: Back translation with environmental dropout
Hao Tan, Licheng Yu, and Mohit Bansal · 2019
Earlier work this paper cites.
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang · 2019
Earlier work this paper cites.
Non-euclidean navigation
William H Warren · 2019
Earlier work this paper cites.
Visual entailment: A novel task for fine-grained image understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Object goal navigation using goal-oriented semantic exploration
Devendra Singh Chaplot, Dhiraj Prakashchand Gandhi, Abhinav Gupta, and Russ R Salakhutdinov · 2020
Earlier work this paper cites.
Just ask: An interactive learning framework for vision and language navigation
Ta-Chung Chi, Minmin Shen, Mihail Eric, Seokhwan Kim, and Dilek Hakkani-Tur · 2020
Earlier work this paper cites.
Semantic information for robot navigation: A survey
Jonathan Crespo, Jose Carlos Castillo, Oscar Martinez Mozos, and Ramon Barber · 2020
Earlier work this paper cites.
Evolving graphical planner: Contextual global planning for vision-and-language navigation
Zhiwei Deng, Karthik Narasimhan, and Olga Russakovsky · 2020
Earlier work this paper cites.
Probing the invariant structure of spatial knowledge: Support for the cognitive graph hypothesis
Jonathan D Ericson and William H Warren · 2020
Earlier work this paper cites.
Towards learning a generic agent for vision-and-language navigation via pre-training
Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, and Jianfeng Gao · 2020
Earlier work this paper cites.
Learning to follow directions in street view
Karl Moritz Hermann, Mateusz Malinowski, Piotr Mirowski, Andras Banki-Horvath, Keith Anderson, and Raia Hadsell · 2020
Earlier work this paper cites.
Sub-instruction aware vision-and-language navigation
Yicong Hong, Cristian Rodriguez, Qi Wu, and Stephen Gould · 2020
Earlier work this paper cites.
Beyond the nav-graph: Vision-and-language navigation in continuous environments
Jacob Krantz, Erik Wijmans, Arjun Majumdar, Dhruv Batra, and Stefan Lee · 2020
Earlier work this paper cites.
Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding
Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge · 2020
Earlier work this paper cites.
Generative language-grounded policy in vision-and-language navigation with bayes’ rule
Shuhei Kurita and Kyunghyun Cho · 2020
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al · 2020
Earlier work this paper cites.
Improving vision-and-language navigation with image-text pairs from the web
Arjun Majumdar, Ayush Shrivastava, Stefan Lee, Peter Anderson, Devi Parikh, and Dhruv Batra · 2020
Earlier work this paper cites.
Conditional driving from natural language instructions
Junha Roh, Chris Paxton, Andrzej Pronobis, Ali Farhadi, and Dieter Fox · 2020
Earlier work this paper cites.
Rmm: A recursive mental model for dialogue navigation
Homero Roman Roman, Yonatan Bisk, Jesse Thomason, Asli Celikyilmaz, and Jianfeng Gao · 2020
Earlier work this paper cites.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 2020
Earlier work this paper cites.
Vision-and-dialog navigation
Jesse Thomason, Michael Murray, Maya Cakmak, and Luke Zettlemoyer · 2020
Earlier work this paper cites.
Multi-view learning for vision-and-language navigation
Qiaolin Xia, Xiujun Li, Chunyuan Li, Yonatan Bisk, Zhifang Sui, Jianfeng Gao, Yejin Choi, and Noah A Smith · 2020
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel · 2020
Earlier work this paper cites.
Vision-language navigation with self-supervised auxiliary reasoning tasks
Fengda Zhu, Yi Zhu, Xiaojun Chang, and Xiaodan Liang · 2020
Earlier work this paper cites.
Neighbor-view enhanced model for vision and language navigation
Dong An, Yuankai Qi, Yan Huang, Qi Wu, Liang Wang, and Tieniu Tan · 2021
Earlier work this paper cites.
Sim-to-real transfer for vision-and-language navigation
Peter Anderson, Ayush Shrivastava, Joanne Truong, Arjun Majumdar, Devi Parikh, Dhruv Batra, and Stefan Lee · 2021
Earlier work this paper cites.
The robotslang benchmark: Dialog-guided robot localization and navigation
Shurjo Banerjee, Jesse Thomason, and Jason Corso · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Landmark-rxr: Solving vision-and-language navigation with fine-grained alignment supervision
Keji He, Yan Huang, Qi Wu, Jianhua Yang, Dong An, Shuanglin Sima, and Liang Wang · 2021
Cited alongside, same era.
A recurrent vision-and-language BERT for navigation
Yicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez Opazo, and Stephen Gould · 2021
Cited alongside, same era.
Hierarchical cross-modal agent for robotics vision-and-language navigation
Muhammad Zubair Irshad, Chih-Yao Ma, and Zsolt Kira · 2021
Cited alongside, same era.
Ndh-full: Learning and evaluating navigational agents on full-length dialogue
Hyounghun Kim, Jialu Li, and Mohit Bansal · 2021
Cited alongside, same era.
Pathdreamer: A world model for indoor navigation
Jing Yu Koh, Honglak Lee, Yinfei Yang, Jason Baldridge, and Peter Anderson · 2021
Cited alongside, same era.
Waypoint models for instruction-guided navigation in continuous environments
Planning-oriented autonomous driving
Yihan Hu, Jiazhi Yang, Li Chen, Keyu Li, Chonghao Sima, Xizhou Zhu, Siqi Chai, Senyao Du, Tianwei Lin, Wenhai Wang, Lewei Lu, Xiaosong Jia, Qiang Liu, Jifeng Dai, Yu Qiao, and Hongyang Li · 2023
Later among the works it cites.
Language models, agent models, and world models: The law for machine reasoning and planning
Zhiting Hu and Tianmin Shu · 2023
Later among the works it cites.
Conceptfusion: Open-set multimodal 3d mapping
Krishna Murthy Jatavallabhula, Alihusein Kuwajerwala, Qiao Gu, Mohd Omama, Tao Chen, Alaa Maalouf, Shuang Li, Ganesh Iyer, Soroush Saryazdi, Nikhil Keetha, et al · 2023
Later among the works it cites.
Syntax-guided transformers: Elevating compositional generalization and grounding in multimodal environments
Danial Kamali and Parisa Kordjamshidi · 2023
Later among the works it cites.
A new path: Scaling vision-and-language navigation with synthetic instructions and imitation learning
Aishwarya Kamath, Peter Anderson, Su Wang, Jing Yu Koh, Alexander Ku, Austin Waters, Yinfei Yang, Jason Baldridge, and Zarana Parekh · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Krantz, Aaron Gokaslan, Dhruv Batra, Stefan Lee, and Oleksandr Maksymets · 2021
Cited alongside, same era.
Improving cross-modal alignment in vision language navigation via syntactic information
Jialu Li, Hao Tan, and Mohit Bansal · 2021
Cited alongside, same era.
Scene-intuitive agent for remote embodied visual grounding
Xiangru Lin, Guanbin Li, and Yizhou Yu · 2021
Cited alongside, same era.
Vision-language navigation with random environmental mixup
Chong Liu, Fengda Zhu, Xiaojun Chang, Xiaodan Liang, Zongyuan Ge, and Yi-Dong Shen · 2021
Cited alongside, same era.
Crossmap transformer: A crossmodal masked path transformer using double back-translation for vision-and-language navigation
Aly Magassouba, Komei Sugiura, and Hisashi Kawai · 2021
Cited alongside, same era.
Film: Following instructions in language with modular methods
So Yeon Min, Devendra Singh Chaplot, Pradeep Kumar Ravikumar, Yonatan Bisk, and Ruslan Salakhutdinov · 2021
Cited alongside, same era.
A survey on human-aware robot navigation
Ronja Möller, Antonino Furnari, Sebastiano Battiato, Aki Härmä, and Giovanni Maria Farinella · 2021
Cited alongside, same era.
Simple and effective synthesis of indoor 3d scenes
Jing Yu Koh, Harsh Agrawal, Dhruv Batra, Richard Tucker, Austin Waters, Honglak Lee, Yinfei Yang, Jason Baldridge, and Peter Anderson · 2023
Later among the works it cites.
Improving vision-and-language navigation by generating future-view image semantics
Jialu Li and Mohit Bansal · 2023
Later among the works it cites.
Evaluating object hallucination in large vision-language models
Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen · 2023
Later among the works it cites.
Code as policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng · 2023
Later among the works it cites.
Thinkbot: Embodied instruction following with thought chain reasoning
Guanxing Lu, Ziwei Wang, Changliu Liu, Jiwen Lu, and Yansong Tang · 2023
Later among the works it cites.
Towards a holistic landscape of situated theory of mind in large language models
Ziqiao Ma, Jacob Sansom, Run Peng, and Joyce Chai · 2023
Later among the works it cites.
Gpt-driver: Learning to drive with gpt
Jiageng Mao, Yuxi Qian, Junjie Ye, Hang Zhao, and Yue Wang · 2023
Later among the works it cites.
Evaluating cognitive maps and planning in large language models with cogeval
Ida Momennejad, Hosein Hasanbeig, Felipe Vieira Frujeri, Hiteshi Sharma, Nebojsa Jojic, Hamid Palangi, Robert Ness, and Jonathan Larson · 2023
Later among the works it cites.
Emma: A foundation model for embodied, interactive, multimodal task completion in 3d environments
Amit Parekh, Malvina Nikandrou, Georgios Pantazopoulos, Bhathiya Hemanthage, Arash Eshghi, Ioannis Konstas, Oliver Lemon, and Alessandro Suglia · 2023
Later among the works it cites.
Visual language navigation: A survey and open challenges
Sang-Min Park and Young-Gab Kim · 2023
Later among the works it cites.
Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning
Krishan Rana, Jesse Haviland, Sourav Garg, Jad Abou-Chakra, Ian Reid, and Niko Suenderhauf · 2023
Later among the works it cites.
Robots that ask for help: Uncertainty alignment for large language model planners
Allen Z Ren, Anushri Dixit, Alexandra Bodrova, Sumeet Singh, Stephen Tu, Noah Brown, Peng Xu, Leila Takayama, Fei Xia, Jake Varley, et al · 2023
Later among the works it cites.
Open-ended instructable embodied agents with memory-augmented large language models
Gabriel Sarch, Yue Wu, Michael Tarr, and Katerina Fragkiadaki · 2023
Later among the works it cites.
Languagempc: Large language models as decision makers for autonomous driving
Hao Sha, Yao Mu, Yuxuan Jiang, Li Chen, Chenfeng Xu, Ping Luo, Shengbo Eben Li, Masayoshi Tomizuka, Wei Zhan, and Mingyu Ding · 2023
Later among the works it cites.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
Dhruv Shah, Błażej Osiński, Sergey Levine, et al · 2023
Later among the works it cites.
Drivelm: Driving with graph visual question answering
Chonghao Sima, Katrin Renz, Kashyap Chitta, Li Chen, Hanxue Zhang, Chengen Xie, Ping Luo, Andreas Geiger, and Hongyang Li · 2023
Later among the works it cites.
Llm-planner: Few-shot grounded planning for embodied agents with large language models
Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M Sadler, Wei-Lun Chao, and Yu Su · 2023
Later among the works it cites.
Target-grounded graph-aware transformer for aerial vision-and-dialog navigation
Yifei Su, Dong An, Yuan Xu, Kehan Chen, and Yan Huang · 2023
Later among the works it cites.
Dilu: A knowledge-driven approach to autonomous driving with large language models
Licheng Wen, Daocheng Fu, Xin Li, Xinyu Cai, Tao Ma, Pinlong Cai, Min Dou, Botian Shi, Liang He, and Yu Qiao · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Qinghao Ye, Haiyang Xu, Guohai Xu, Jiabo Ye, Ming Yan, Yiyang Zhou, Junyang Wang, Anwen Hu, Pengcheng Shi, Yaya Shi, et al · 2023
Later among the works it cites.
Homerobot: Open-vocabulary mobile manipulation
Sriram Yenamandra, Arun Ramachandran, Karmesh Yadav, Austin S Wang, Mukul Khanna, Theophile Gervet, Tsung-Yen Yang, Vidhi Jain, Alexander Clegg, John M Turner, et al · 2023
Later among the works it cites.
Seagull: An embodied agent for instruction following through situated dialog
Yichi Zhang, Jianing Yang, Keunwoo Yu, Yinpei Dai, Shane Storks, Yuwei Bao, Jiayi Pan, Nikhil Devraj, Ziqiao Ma, and Joyce Chai · 2023
Later among the works it cites.
VLN-Trans: Translator for the vision and language navigation agent
Yue Zhang and Parisa Kordjamshidi · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Later among the works it cites.
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, et al · 2023
Later among the works it cites.
Tree-structured trajectory encoding for vision-and-language navigation
Xinzhe Zhou and Yadong Mu · 2023
Later among the works it cites.
Vision language navigation with knowledge-driven environmental dreamer
Fengda Zhu, Vincent CS Lee, Xiaojun Chang, and Xiaodan Liang · 2023
Later among the works it cites.
Etpnav: Evolving topological planning for vision-language navigation in continuous environments
Dong An, Hanqing Wang, Wenguan Wang, Zun Wang, Yan Huang, Keji He, and Liang Wang · 2024
Closest in time.
Goat: Go to any thing
Matthew Chang, Theophile Gervet, Mukul Khanna, Sriram Yenamandra, Dhruv Shah, So Yeon Min, Kavit Shah, Chris Paxton, Saurabh Gupta, Dhruv Batra, et al · 2024
Closest in time.
Driving with llms: Fusing object-level vector modality for explainable autonomous driving
Long Chen, Oleg Sinavski, Jan Hünermann, Alice Karnsund, Andrew James Willmott, Danny Birch, Daniel Maund, and Jamie Shotton · 2024
Closest in time.
Shine: Saliency-aware hierarchical negative ranking for compositional temporal grounding
Zixu Cheng, Yujiang Pu, Shaogang Gong, Parisa Kordjamshidi, and Yu Kong · 2024
Closest in time.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2024
Closest in time.
A survey on multimodal large language models for autonomous driving
Can Cui, Yunsheng Ma, Xu Cao, Wenqian Ye, Yang Zhou, Kaizhao Liang, Jintai Chen, Juanwu Lu, Zichong Yang, Kuei-Da Liao, et al · 2024
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, DONGXU LI, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi · 2024
Closest in time.
A survey for foundation models in autonomous driving
Haoxiang Gao, Yaqian Li, Kaiwen Long, Ming Yang, and Yiqing Shen · 2024
Closest in time.
Spatially-aware speaker for vision-and-language navigation instruction generation
Muraleekrishna Gopinathan, Martin Masek, Jumana Abu-Khalaf, and David Suter · 2024
Closest in time.
Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning
Qiao Gu, Ali Kuwajerwala, Sacha Morin, Krishna Murthy Jatavallabhula, Bipasha Sen, Aditya Agarwal, Corban Rivera, William Paul, Kirsty Ellis, Rama Chellappa, et al · 2024
Closest in time.
Learning human-to-humanoid real-time whole-body teleoperation
Tairan He, Zhengyi Luo, Wenli Xiao, Chong Zhang, Kris Kitani, Changliu Liu, and Guanya Shi · 2024
Closest in time.
LRM: Large reconstruction model for single image to 3d
Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan · 2024
Closest in time.
Drivlme: Enhancing llm-based autonomous driving agents with embodied and social experiences
Yidong Huang, Jacob Sansom, Ziqiao Ma, Felix Gervits, and Joyce Chai · 2024
Closest in time.
Panogen: Text-conditioned panoramic environment generation for vision-and-language navigation
Jialu Li and Mohit Bansal · 2024
Closest in time.
Volumetric environment representation for vision-language navigation
Rui Liu, Wenguan Wang, and Yi Yang · 2024
Closest in time.
Discuss before moving: Visual language navigation via multi-expert discussions
Yuxing Long, Xiaoqi Li, Wenzhe Cai, and Hao Dong · 2024
Closest in time.
Embodiedgpt: Vision-language pre-training via embodied chain of thought
Yao Mu, Qinglong Zhang, Mengkang Hu, Wenhai Wang, Mingyu Ding, Jun Jin, Bin Wang, Jifeng Dai, Yu Qiao, and Ping Luo · 2024
Closest in time.
Langnav: Language as a perceptual representation for navigation
Bowen Pan, Rameswar Panda, SouYoung Jin, Rogerio Feris, Aude Oliva, Phillip Isola, and Yoon Kim · 2024
Closest in time.
Habitat 3.0: A co-habitat for humans, avatars, and robots
Xavier Puig, Eric Undersander, Andrew Szot, Mikael Dallaire Cote, Tsung-Yen Yang, Ruslan Partsey, Ruta Desai, Alexander Clegg, Michal Hlavac, So Yeon Min, et al · 2024
Closest in time.
Llm as copilot for coarse-grained vision-and-language navigation
Yanyuan Qiao, Qianyi Liu, Jiajun Liu, Jing Liu, and Qi Wu · 2024
Closest in time.
Saynav: Grounding large language models for dynamic planning to navigation in new environments
Abhinav Rajvanshi, Karan Sikka, Xiao Lin, Bhoram Lee, Han-Pang Chiu, and Alvaro Velasquez · 2024
Closest in time.
Lmdrive: Closed-loop end-to-end driving with large language models
Hao Shao, Yuxuan Hu, Letian Wang, Guanglu Song, Steven L Waslander, Yu Liu, and Hongsheng Li · 2024
Closest in time.
Drivevlm: The convergence of autonomous driving and large vision-language models
Xiaoyu Tian, Junru Gu, Bailin Li, Yicheng Liu, Chenxu Hu, Yang Wang, Kun Zhan, Peng Jia, Xianpeng Lang, and Hang Zhao · 2024
Closest in time.
Bootstrapping llm-based task-oriented dialogue agents via self-talk
Dennis Ulmer, Elman Mansimov, Kaixiang Lin, Justin Sun, Xibin Gao, and Yi Zhang · 2024
Closest in time.
Vision-language navigation: a survey and taxonomy
Wansen Wu, Tao Chang, Xinmeng Li, Quanjun Yin, and Yue Hu · 2024
Closest in time.
Language models meet world models: Embodied experiences enhance language models
Jiannan Xiang, Tianhua Tao, Yi Gu, Tianmin Shu, Zirui Wang, Zichao Yang, and Zhiting Hu · 2024
Closest in time.
Xu Yan, Haiming Zhang, Yingjie Cai, Jingming Guo, Weichao Qiu, Bin Gao, Kaiqiang Zhou, Yue Zhao, Huan Jin, Jiantao Gao, et al · 2024
Closest in time.
Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent
Jianing Yang, Xuweiyi Chen, Shengyi Qian, Nikhil Madaan, Madhavan Iyengar, David F. Fouhey, and Joyce Chai · 2024
Closest in time.
Jianhao Yuan, Shuyang Sun, Daniel Omeiza, Bo Zhao, Paul Newman, Lars Kunze, and Matthew Gadd · 2024
Closest in time.
Fine-tuning large vision-language models as decision-making agents via reinforcement learning
Yuexiang Zhai, Hao Bai, Zipeng Lin, Jiayi Pan, Shengbang Tong, Yifei Zhou, Alane Suhr, Saining Xie, Yann LeCun, Yi Ma, and Sergey Levine · 2024
Closest in time.
Narrowing the gap between vision and action in navigation
Yue Zhang and Parisa Kordjamshidi · 2024
Closest in time.
NavHint: Vision and language navigation agent with a hint generator
Yue Zhang, Quan Guo, and Parisa Kordjamshidi · 2024
Closest in time.
Do vision-language models represent space and how? evaluating spatial frame of reference under ambiguities
Zheyuan Zhang, Fengyuan Hu, Jayjun Lee, Freda Shi, Parisa Kordjamshidi, Joyce Chai, and Ziqiao Ma · 2024
Closest in time.
Octopus: Embodied vision-language programmer from environmental feedback
Jingkang Yang, Yuhao Dong, Shuai Liu, Bo Li, Ziyue Wang, Haoran Tan, Chencheng Jiang, Jiamu Kang, Yuanhan Zhang, Kaiyang Zhou, et al · 2025
Closest in time.
Common sense reasoning for deepfake detection
Yue Zhang, Ben Colman, Xiao Guo, Ali Shahriyari, and Gaurav Bharaj · 2025
Closest in time.
Navgpt-2: Unleashing navigational reasoning capability for large vision-language models
Gengze Zhou, Yicong Hong, Zun Wang, Xin Eric Wang, and Qi Wu · 2025
Closest in time.