Fetching the paper…
Reading the bibliography…
Despite recent progress, learning new tasks through language instructions remains an extremely challenging problem.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019 · 1910
Earlier work this paper cites.
Htn planning: Complexity and expressivity
Kutluhan Erol, James Hendler, and Dana S Nau. 1994 · 1994
Earlier work this paper cites.
Interactive robot task training through dialog and demonstration
P. E. Rybski, K. Yoon, J. Stolarz, and M. M. Veloso. 2007 · 2007
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning. 2009 · 2009
Earlier work this paper cites.
Learning about objects with human teachers
A. L. Thomaz and M. Cakmak. 2009 · 2009
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew J. Hausknecht. 2020 · 2010
Earlier work this paper cites.
MOCA: A modular object-centric approach for interactive instruction following
Kunal Pratap Singh, Suvaansh Bhambri, Byeonghwi Kim, Roozbeh Mottaghi, and Jonghyun Choi. 2020 · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Learning goal-oriented hierarchical tasks from situated interactive instruction
Shiwali Mohan and John E. Laird. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Jointly learning grounded task structures from language instruction and visual demonstration
Changsong Liu, Shaohua Yang, Sari Saba-Sadiya, Nishant Shukla, Yunzhong He, Song-Chun Zhu, and Joyce Chai. 2016 · 2016
Earlier work this paper cites.
Mask R-CNN
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross B. Girshick. 2017 · 2017
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi. 2017 · 2017
Earlier work this paper cites.
Mapping instructions and visual observations to actions with reinforcement learning
Dipendra Misra, John Langford, and Yoav Artzi. 2017 · 2017
Cited alongside, same era.
Spoken instruction-based one-shot object and action learning in a cognitive robotic architecture
M. Scheutz, E. Krause, B. Oosterveld, T. Frasca, and R. Platt. 2017 · 2017
Cited alongside, same era.
Interactive learning of grounded verb semantics towards human-robot communication
Lanbo She and Joyce Chai. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Visual semantic planning using deep successor representations
Yuke Zhu, Daniel Gordon, Eric Kolve, Dieter Fox, Li Fei-Fei, Abhinav Gupta, Roozbeh Mottaghi, and Ali Farhadi. 2017 · 2017
Cited alongside, same era.
TOUCHDOWN: natural language navigation and spatial reasoning in visual street environments
Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi. 2019 · 2019
Later among the works it cites.
Stay on the path: Instruction fidelity in vision-and-language navigation
Vihan Jain, Gabriel Magalhaes, Alexander Ku, Ashish Vaswani, Eugene Ie, and Jason Baldridge. 2019 · 2019
Later among the works it cites.
Tactical rewind: Self-correction via backtracking in vision-and-language navigation
Liyiming Ke, Xiujun Li, Yonatan Bisk, Ari Holtzman, Zhe Gan, Jingjing Liu, Jianfeng Gao, Yejin Choi, and Siddhartha Srinivasa. 2019 · 2019
Later among the works it cites.
Self-monitoring navigation agent via auxiliary progress estimation
Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan AlRegib, Zsolt Kira, Richard Socher, and Caiming Xiong. 2019a · 2019
Later among the works it cites.
The regretful agent: Heuristic-aided navigation through progress estimation
Chih-Yao Ma, Zuxuan Wu, Ghassan AlRegib, Caiming Xiong, and Zsolt Kira. 2019b · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian D. Reid, Stephen Gould, and Anton van den Hengel. 2018 · 2018
Cited alongside, same era.
Language to action: Towards interactive task learning with physical agents
Joyce Y. Chai, Qiaozi Gao, Lanbo She, Shaohua Yang, Sari Saba-Sadiya, and Guangyue Xu. 2018 · 2018
Cited alongside, same era.
Embodied question answering
Abhishek Das, Samyak Datta, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra. 2018a · 2018
Cited alongside, same era.
Speaker-follower models for vision-and-language navigation
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. 2018 · 2018
Cited alongside, same era.
Interactive task learning: Humans, robots, and agents acquiring new tasks through natural interactions
Kevin A Gluck and John E Laird. 2018 · 2018
Cited alongside, same era.
IQA: visual question answering in interactive environments
Daniel Gordon, Aniruddha Kembhavi, Mohammad Rastegari, Joseph Redmon, Dieter Fox, and Ali Farhadi. 2018 · 2018
Cited alongside, same era.
Mapping instructions to actions in 3d environments with visual goal prediction
Dipendra Kumar Misra, Andrew Bennett, Valts Blukis, Eyvind Niklasson, Max Shatkhin, and Yoav Artzi. 2018 · 2018
Cited alongside, same era.
Situational fusion of visual representation for visual navigation
William B. Shen, Danfei Xu, Yuke Zhu, Fei-Fei Li, Leonidas J. Guibas, and Silvio Savarese. 2019 · 2019
Later among the works it cites.
Learning to navigate unseen environments: Back translation with environmental dropout
Hao Tan, Licheng Yu, and Mohit Bansal. 2019 · 2019
Later among the works it cites.
A recurrent vision-and-language bert for navigation
Yicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez Opazo, and Stephen Gould. 2020 · 2020
Later among the works it cites.
Visually-grounded planning without vision: Language models infer detailed plans from high-level instructions
Peter Jansen. 2020 · 2020
Later among the works it cites.
Learning adaptive language interfaces through decomposition
Siddharth Karamcheti, Dorsa Sadigh, and Percy Liang. 2020 · 2020
Later among the works it cites.
A hierarchical attention model for action learning from realistic environments and directives
Van-Quang Nguyen and Takayuki Okatani. 2020 · 2020
Later among the works it cites.
ALFRED: A benchmark for interpreting grounded instructions for everyday tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. 2020 · 2020
Later among the works it cites.
Diagnosing the environment bias in vision-and-language navigation
Yubo Zhang, Hao Tan, and Mohit Bansal. 2020 · 2020
Later among the works it cites.
Are we there yet? learning to localize in embodied instruction following
Shane Storks, Qiaozi Gao, Govind Thattai, and Gokhan Tur. 2021 · 2021
Closest in time.