Fetching the paper…
Reading the bibliography…
Recent methods for embodied instruction following are typically trained end-to-end using imitation learning.
A fast marching level set method for monotonically advancing fronts
J A Sethian · 1996
Earlier work this paper cites.
Learning to explore using active neural slam
Devendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta, and Ruslan Salakhutdinov · 2004
Earlier work this paper cites.
Modeling the world from internet photo collections
Noah Snavely, Steven M Seitz, and Richard Szeliski · 2008
Earlier work this paper cites.
Kinectfusion: real-time 3d reconstruction and interaction using a moving depth camera
Shahram Izadi, David Kim, Otmar Hilliges, David Molyneaux, Richard Newcombe, Pushmeet Kohli, Jamie Shotton, Steve Hodges, Dustin Freeman, Andrew Davison, et al · 2011
Earlier work this paper cites.
Visual simultaneous localization and mapping: a survey
Jorge Fuentes-Pacheco, José Ruiz-Ascencio, and Juan Manuel Rendón-Mancha · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Cognitive mapping and planning for visual navigation
Saurabh Gupta, James Davidson, Sergey Levine, Rahul Sukthankar, and Jitendra Malik · 2017
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick · 2017
Earlier work this paper cites.
A-fast-rcnn: Hard positive generation via adversary for object detection
Xiaolong Wang, Abhinav Shrivastava, and Abhinav Gupta · 2017
Earlier work this paper cites.
Embodied question answering
Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Speaker-follower models for vision-and-language navigation
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell · 2018
Earlier work this paper cites.
Iqa: Visual question answering in interactive environments
Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi · 2018
Cited alongside, same era.
Deep learning for computer vision: A brief review
Athanasios Voulodimos, Nikolaos Doulamis, Anastasios Doulamis, and Eftychios Protopapadakis · 2018
Cited alongside, same era.
Tactical rewind: Self-correction via backtracking in vision-and-language navigation
Liyiming Ke, Xiujun Li, Yonatan Bisk, Ari Holtzman, Zhe Gan, Jingjing Liu, Jianfeng Gao, Yejin Choi, and Siddhartha Srinivasa · 2019
Cited alongside, same era.
Ai2-thor: An interactive 3d environment for visual ai, 2019
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi · 2019
Cited alongside, same era.
The regretful agent: Heuristic-aided navigation through progress estimation
Chih-Yao Ma, Zuxuan Wu, Ghassan AlRegib, Caiming Xiong, and Zsolt Kira · 2019
Cited alongside, same era.
Moca: A modular object-centric approach for interactive instruction following
Kunal Pratap Singh, Suvaansh Bhambri, Byeonghwi Kim, Roozbeh Mottaghi, and Jonghyun Choi · 2020
Later among the works it cites.
Language-mediated, object-centric representation learning
Ruocheng Wang, Jiayuan Mao, Samuel J Gershman, and Jiajun Wu · 2020
Later among the works it cites.
Vision-language navigation with self-supervised auxiliary reasoning tasks
Fengda Zhu, Yi Zhu, Xiaojun Chang, and Xiaodan Liang · 2020
Later among the works it cites.
A persistent spatial semantic representation for high-level natural language instruction execution
Valts Blukis, Chris Paxton, Dieter Fox, Animesh Garg, and Yoav Artzi · 2021
Closest in time.
Vtnet: Visual transformer network for object goal navigation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Cited alongside, same era.
Experience grounds language
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, et al · 2020
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Cited alongside, same era.
Embodied language grounding with 3d visual feature representations
Mihir Prabhudesai, Hsiao-Yu Fish Tung, Syed Ashar Javed, Maximilian Sieb, Adam W Harley, and Katerina Fragkiadaki · 2020
Cited alongside, same era.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 2020
Cited alongside, same era.
On Evaluation of Embodied Navigation Agents
Peter Anderson, Angel Chang, Devendra Singh Chaplot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, et al
Cited in the paper.
Heming Du, Xin Yu, and Liang Zheng · 2021
Closest in time.
Agent with the big picture: Perceiving surroundings for interactive instruction following
Byeonghwi Kim, Suvaansh Bhambri, Kunal Pratap Singh, Roozbeh Mottaghi, and Jonghyun Choi · 2021
Closest in time.
Look wide and interpret twice: Improving performance on interactive instruction-following tasks
Van-Quang Nguyen, Masanori Suganuma, and Takayuki Okatani · 2021
Closest in time.
Lav, 2021
Kolby Nottingham, Litian Liang, Daehyun Shin, Charless C. Fowlkes, Roy Fox, and Sameer Singh · 2021
Closest in time.
Episodic transformer for vision-and-language navigation
Alexander Pashevich, Cordelia Schmid, and Chen Sun · 2021
Closest in time.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht · 2021
Closest in time.
Hierarchical task learning from language instructions with unified transformers and self-monitoring
Yichi Zhang and Joyce Chai · 2021
Closest in time.