Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Original
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht · 2010
Earlier work this paper cites.
Understanding and executing instructions for everyday manipulation tasks from the world wide web
Moritz Tenorth, Daniel Nyga, and Michael Beetz · 2010
Earlier work this paper cites.
Weakly supervised learning of semantic parsers for mapping instructions to actions
Yoav Artzi and Luke Zettlemoyer · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Environment-driven lexicon induction for high-level instructions
Dipendra Misra, Kejia Tao, Percy Liang, and Ashutosh Saxena · 2015
Earlier work this paper cites.
Learning from stories: using crowdsourced narratives to train virtual agents
Brent Harrison and Mark O Riedl · 2016
Earlier work this paper cites.
Tell me dave: Context-sensitive grounding of natural language to manipulation instructions
Dipendra K Misra, Jaeyong Sung, Kevin Lee, and Ashutosh Saxena · 2016
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Original
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia · 2017
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
Original
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Speaker-follower models for vision-and-language navigation
Original
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell · 2018
Earlier work this paper cites.
Virtualhome: Simulating household activities via programs
Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba · 2018
Earlier work this paper cites.
Commonsense knowledge mining from pretrained models
Joe Davison, Joshua Feldman, and Alexander M Rush · 2019
Earlier work this paper cites.
From language to goals: Inverse reinforcement learning for vision-based instruction following
Original
Justin Fu, Anoop Korattikara, Sergey Levine, and Sergio Guadarrama · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Original
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2019
Earlier work this paper cites.