Fetching the paper…
Reading the bibliography…
Interaction and navigation defined by natural language instructions in dynamic environments pose significant challenges for neural agents.
The Temporal Logic of Reactive and Concurrent Systems
Zohar Manna and Amir Pnueli · 1992
Earlier work this paper cites.
Learning to parse database queries using inductive logic programming
John M Zelle and Raymond J Mooney · 1996
Earlier work this paper cites.
PDDL: The Planning Domain Definition Language, 1998
M. Ghallab, A. Howe, C. Knoblock, D. Mcdermott, A. Ram, M. Veloso, D. Weld, and D. Wilkins · 1998
Earlier work this paper cites.
Using verbal instructions for route learning: Instruction analysis
Guido Bugmann, Stanislao Lauria, Theocharis Kyriacou, Ewan Klein, Johan Bos, and Kenny Coventry · 2001
Earlier work this paper cites.
Walk the talk: Connecting language, knowledge, and action in route instructions
Matt MacMahon, Brian Stankiewicz, and Benjamin Kuipers · 2006
Earlier work this paper cites.
Online learning of relaxed ccg grammars for parsing to logical form
Luke Zettlemoyer and Michael Collins · 2007
Earlier work this paper cites.
Reinforcement learning for mapping instructions to actions
Satchuthananthavale RK Branavan, Harr Chen, Luke S Zettlemoyer, and Regina Barzilay · 2009
Earlier work this paper cites.
Understanding and executing instructions for everyday manipulation tasks from the world wide web
Moritz Tenorth, Daniel Nyga, and Michael Beetz · 2010
Earlier work this paper cites.
Learning to interpret natural language navigation instructions from observations
David Chen and Raymond Mooney · 2011
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
Stefanie Tellex, Thomas Kollar, Steven Dickerson, Matthew Walter, Ashis Banerjee, Seth Teller, and Nicholas Roy · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Luke S Zettlemoyer and Michael Collins · 2012
Earlier work this paper cites.
Weakly supervised learning of semantic parsers for mapping instructions to actions
Yoav Artzi and Luke Zettlemoyer · 2013
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang · 2013
Earlier work this paper cites.
Interpreting and executing recipes with a cooking robot
Mario Bollini, Stefanie Tellex, Tyler Thompson, Nicholas Roy, and Daniela Rus · 2013
Earlier work this paper cites.
Alignment-based compositional semantics for instruction following
Jacob Andreas and Dan Klein · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Grammar as a foreign language
Oriol Vinyals, Lukasz Kaiser, Terry Koo, Slav Petrov, Ilya Sutskever, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, Adrià Puigdomènech Badia, Karl Moritz Hermann, Yori Zwols, Georg Ostrovski, Adam Cain, Helen King, Christopher Summerfield, Phil Blunsom, Koray Kavukcuoglu, and Demis Hassabis · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Listen, attend, and walk: Neural mapping of navigational instructions to action sequences
Hongyuan Mei, Mohit Bansal, and Matthew Walter · 2016
Earlier work this paper cites.
Tell me dave: Context-sensitive grounding of natural language to manipulation instructions
Dipendra K Misra, Jaeyong Sung, Kevin Lee, and Ashutosh Saxena · 2016
Earlier work this paper cites.
Mask r-cnn
Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick · 2017
Earlier work this paper cites.
AI2-THOR: An Interactive 3D Environment for Visual AI
Eric Kolve, Roozbeh Mottaghi, Winson Han, Eli VanderBilt, Luca Weihs, Alvaro Herrasti, Daniel Gordon, Yuke Zhu, Abhinav Gupta, and Ali Farhadi · 2017
Earlier work this paper cites.
Mapping instructions and visual observations to actions with reinforcement learning
Dipendra Misra, John Langford, and Yoav Artzi · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
On evaluation of embodied navigation agents
Peter Anderson, Angel X. Chang, Devendra Singh Chaplot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, and Amir Roshan Zamir · 2018
Cited alongside, same era.
Vision-and-Language Navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel · 2018
Cited alongside, same era.
Textworld: A learning environment for text-based games
Marc-Alexandre Côté, Ákos Kádár, Xingdi Yuan, Ben Kybartas, Tavian Barnes, Emery Fine, James Moore, Matthew Hausknecht, Layla El Asri, Mahmoud Adada, et al · 2018
Cited alongside, same era.
Speaker-follower models for vision-and-language navigation
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell · 2018
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
Xin Wang, Qiuyuan Huang, Celikyilmaz Asli, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Wang, and Lei Zhang · 2019
Later among the works it cites.
Quantifying attention flow in transformers
Samira Abnar and Willem Zuidema · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Later among the works it cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Later among the works it cites.
Robothor: An open simulation-to-real embodied ai platform
Matt Deitke, Winson Han, Alvaro Herrasti, Aniruddha Kembhavi, Eric Kolve, Roozbeh Mottaghi, Jordi Salvador, Dustin Schwenk, Eli VanderBilt, Matthew Wallingford, Luca Weihs, Mark Yatskar, and Ali Farhadi · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sequence-to-sequence language grounding of non-markovian task specifications
Nakul Gopalan, Dilip Arumugam, Lawson Wong, and Stefanie Tellex · 2018
Cited alongside, same era.
Efficient grounding of abstract spatial concepts for natural language interaction with robot platforms
Rohan Paul, Jacob Arkin, Derya Aksaray, Nicholas Roy, and Thomas M Howard · 2018
Cited alongside, same era.
Virtualhome: Simulating household activities via programs
Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, and Antonio Torralba · 2018
Cited alongside, same era.
Non-local neural networks
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He · 2018
Cited alongside, same era.
Tao Yu, Rui Zhang, Kai Yang, Michihiro Yasunaga, Dongxu Wang, Zifan Li, James Ma, Irene Li, Qingning Yao, Shanelle Roman, et al · 2018
Cited alongside, same era.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
Howard Chen, Alane Suhr, Dipendra Misra, Noah Snavely, and Yoav Artzi · 2019
Cited alongside, same era.
Causal confusion in imitation learning
Pim de Haan, Dinesh Jayaraman, and Sergey Levine · 2019
Cited alongside, same era.
Later among the works it cites.
David Ding, Felix Hill, Adam Santoro, and Matt Botvinick · 2020
Later among the works it cites.
Multi-modal transformer for video retrieval
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Later among the works it cites.
Beyond the nav-graph: Vision-and-language navigation in continuous environments
Jacob Krantz, Erik Wijmans, Arjun Majumdar, Dhruv Batra, and Stefan Lee · 2020
Later among the works it cites.
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding
Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge · 2020
Later among the works it cites.
Unsupervised reinforcement learning of transferable meta-skills for embodied navigation
Juncheng Li, Xin Wang, Siliang Tang, Haizhou Shi, Fei Wu, Yueting Zhuang, and William Yang Wang · 2020
Later among the works it cites.
Improving vision-and-language navigation with image-text pairs from the web
Arjun Majumdar, Ayush Shrivastava, Stefan Lee, Peter Anderson, Devi Parikh, and Dhruv Batra · 2020
Later among the works it cites.
Harsh Mehta, Yoav Artzi, Jason Baldridge, Eugene Ie, and Piotr Mirowski · 2020
Later among the works it cites.
Stabilizing transformers for reinforcement learning
Emilio Parisotto, Francis Song, Jack Rae, Razvan Pascanu, Caglar Gulcehre, Siddhant Jayakumar, Max Jaderberg, Raphaël Lopez Kaufman, Aidan Clark, Seb Noury, Matthew Botvinick, Nicolas Heess, and Raia Hadsell · 2020
Later among the works it cites.
Grounding language to non-markovian tasks with no supervision of task specifications
Roma Patel, Ellie Pavlick, and Stefanie Tellex · 2020
Later among the works it cites.
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 2020
Later among the works it cites.
MOCA: A Modular Object-Centric Approach for Interactive Instruction Following
Kunal Pratap Singh, Suvaansh Bhambri, Byeonghwi Kim, Roozbeh Mottaghi, and Jonghyun Choi · 2020
Later among the works it cites.
A hierarchical attention model for action learning from realistic environments and directives
Takayuki Okatani Van-Quang Nguyen · 2020
Later among the works it cites.
Vision-language navigation with self-supervised auxiliary reasoning tasks
Fengda Zhu, Yi Zhu, Xiaojun Chang, and Xiaodan Liang · 2020
Later among the works it cites.
Babywalk: Going farther in vision-and-language navigation by taking baby steps
Wang Zhu, Hexiang Hu, Jiacheng Chen, Zhiwei Deng, Vihan Jain, Eugene Ie, and Fei Sha · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Closest in time.
A Recurrent Vision-and-Language BERT for Navigation
Yicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez-Opazo, and Stephen Gould · 2021
Closest in time.
Grounding language in play
Corey Lynch and Pierre Sermanet · 2021
Closest in time.
Episodic Transformer for Vision-and-Language Navigation – supplementary material
Alexander Pashevich, Cordelia Schmid, and Chen Sun · 2021
Closest in time.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht · 2021
Closest in time.