Fetching the paper…
Reading the bibliography…
Incremental decision making in real-world environments is one of the most challenging tasks in embodied artificial intelligence.
Language Models are Few-Shot Learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 1901
Earlier work this paper cites.
Driving Semantic Parsing from the World’s Response
Clarke, J.; Goldwasser, D.; Chang, M.-W.; and Roth, D. 2010 · 2010
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
Ross, S.; Gordon, G. J.; and Bagnell, J. A. 2011 · 2011
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Anderson, P.; Wu, Q.; Teney, D.; Bruce, J.; Johnson, M.; Sünderhauf, N.; Reid, I.; Gould, S.; and Van Den Hengel, A. 2018 · 2018
Earlier work this paper cites.
Speaker-Follower Models for Vision-and-Language Navigation
Fried, D.; Hu, R.; Cirik, V.; Rohrbach, A.; Andreas, J.; Morency, L.-P.; Berg-Kirkpatrick, T.; Saenko, K.; Klein, D.; and Darrell, T. 2018 · 2018
Earlier work this paper cites.
TOUCHDOWN: Natural Language Navigation and Spatial Reasoning in Visual Street Environments
Chen, H.; Suhr, A.; Misra, D.; Snavely, N.; and Artzi, Y. 2019 · 2019
Earlier work this paper cites.
Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout
Tan, H.; Yu, L.; and Bansal, M. 2019 · 2019
Earlier work this paper cites.
Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation
Wang, X.; Huang, Q.; Celikyilmaz, A.; Gao, J.; Shen, D.; Wang, Y.-F.; Wang, W. Y.; and Zhang, L. 2019 · 2019
Earlier work this paper cites.
Counterfactual vision-and-language navigation via adversarial path sampler
Fu, T.-J.; Wang, X. E.; Peterson, M. F.; Grafton, S. T.; Eckstein, M. P.; and Wang, W. Y. 2020 · 2020
Earlier work this paper cites.
Learning to Follow Directions in Street View
Hermann, K. M.; Malinowski, M.; Mirowski, P.; Banki-Horvath, A.; Anderson, K.; and Hadsell, R. 2020 · 2020
Earlier work this paper cites.
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding
Ku, A.; Anderson, P.; Patel, R.; Ie, E.; and Baldridge, J. 2020 · 2020
Earlier work this paper cites.
Retouchdown: Releasing Touchdown on StreetLearn as a Public Resource for Language Grounding Tasks in Street View
Mehta, H.; Artzi, Y.; Baldridge, J.; Ie, E.; and Mirowski, P. 2020 · 2020
Earlier work this paper cites.
Reverie: Remote embodied visual referring expression in real indoor environments
Qi, Y.; Wu, Q.; Anderson, P.; Wang, X.; Wang, W. Y.; Shen, C.; and Hengel, A. v. d. 2020 · 2020
Cited alongside, same era.
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Shridhar, M.; Thomason, J.; Gordon, D.; Bisk, Y.; Han, W.; Mottaghi, R.; Zettlemoyer, L.; and Fox, D. 2020 · 2020
Cited alongside, same era.
Learning to Stop: A Simple yet Effective Approach to Urban Vision-Language Navigation
Xiang, J.; Wang, X.; and Wang, W. Y. 2020 · 2020
Cited alongside, same era.
BabyWalk: Going Farther in Vision-and-Language Navigation by Taking Baby Steps
Zhu, W.; Hu, H.; Chen, J.; Deng, Z.; Jain, V.; Ie, E.; and Sha, F. 2020 · 2020
Cited alongside, same era.
History aware multimodal transformer for vision-and-language navigation
Chen, S.; Guhur, P.-L.; Schmid, C.; and Laptev, I. 2021 · 2021
Cited alongside, same era.
Simple but Effective: CLIP Embeddings for Embodied AI
Khandelwal, A.; Weihs, L.; Mottaghi, R.; and Kembhavi, A. 2022 · 2022
Later among the works it cites.
Envedit: Environment editing for vision-and-language navigation
Li, J.; Tan, H.; and Bansal, M. 2022 · 2022
Later among the works it cites.
LAION-5B: An open large-scale dataset for training next generation image-text models
Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C. W.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; Schramowski, P.; Kundurthy, S. R.; Crowson, K.; Schmidt, L.; Kaczmarczyk, R.; and Jitsev, J. 2022 · 2022
Later among the works it cites.
Analyzing Generalization of Vision and Language Navigation to Unseen Outdoor Areas
Schumann, R.; and Riezler, S. 2022 · 2022
Later among the works it cites.
LM-Nav: Robotic Navigation with Large Pre-Trained Models of Language, Vision, and Action
Shah, D.; Osinski, B.; Ichter, B.; and Levine, S. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hong, Y.; Wu, Q.; Qi, Y.; Rodriguez-Opazo, C.; and Gould, S. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
Ramakrishnan, S. K.; Gokaslan, A.; Wijmans, E.; Maksymets, O.; Clegg, A.; Turner, J. M.; Undersander, E.; Galuba, W.; Westbury, A.; Chang, A. X.; Savva, M.; Zhao, Y.; and Batra, D. 2021 · 2021
Cited alongside, same era.
Generating Landmark Navigation Instructions from Maps as a Graph-to-Text Problem
Schumann, R.; and Riezler, S. 2021 · 2021
Cited alongside, same era.
SILG: The Multi-domain Symbolic Interactive Language Grounding Benchmark
Zhong, V.; Hanjie, A. W.; Wang, S.; Narasimhan, K.; and Zettlemoyer, L. 2021 · 2021
Cited alongside, same era.
Multimodal Text Style Transfer for Outdoor Vision-and-Language Navigation
Zhu, W.; Wang, X.; Fu, T.-J.; Yan, A.; Narayana, P.; Sone, K.; Basu, S.; and Wang, W. Y. 2021 · 2021
Cited alongside, same era.
Clip-nav: Using clip for zero-shot vision-and-language navigation
Dorbala, V. S.; Sigurdsson, G.; Piramuthu, R.; Thomason, J.; and Sukhatme, G. S. 2022 · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; et al. 2022 · 2022
Later among the works it cites.
A Priority Map for Vision-and-Language Navigation with Trajectory Plans and Feature-Location Cues
Armitage, J.; Impett, L.; and Sennrich, R. 2023 · 2023
Closest in time.
Mixtral of Experts: A High Quality Sparse Mixture-of-Experts
Mistral AI Team. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Outdoor Vision-and-Language Navigation Needs Object-Level Alignment
Sun, Y.; Qiu, Y.; Aoki, Y.; and Kataoka, H. 2023 · 2023
Closest in time.
Voyager: An Open-Ended Embodied Agent with Large Language Models
Wang, G.; Xie, Y.; Jiang, Y.; Mandlekar, A.; Xiao, C.; Zhu, Y.; Fan, L.; and Anandkumar, A. 2023 · 2023
Closest in time.
NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models
Zhou, G.; Hong, Y.; and Wu, Q. 2023 · 2023
Closest in time.
ESC: Exploration with Soft Commonsense Constraints for Zero-shot Object Navigation
Zhou, K.; Zheng, K.; Pryor, C.; Shen, Y.; Jin, H.; Getoor, L.; and Wang, X. E. 2023 · 2023
Closest in time.