Fetching the paper…
Reading the bibliography…
Recent work has shown that deep reinforcement-learning agents can learn to follow language-like instructions from infrequent environment rewards.
Procedures as a representation for data in a computer program for understanding natural language
Terry Winograd · 1971
Earlier work this paper cites.
Understanding natural language
Terry Winograd · 1972
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning
Andrew Y. Ng and Stuart Russell · 2000
Earlier work this paper cites.
Apprenticeship Learning via Inverse Reinforcement Learning
Pieter Abbeel and Andrew Y. Ng · 2004
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning
Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey · 2008
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The TAMER framework
W Bradley Knox and Peter Stone · 2009
Earlier work this paper cites.
Learning to Follow Navigational Directions
Adam Vogel and Dan Jurafsky · 2010
Earlier work this paper cites.
Learning to Interpret Natural Language Navigation Instructions from Observations
David L. Chen and Raymond J. Mooney · 2011
Earlier work this paper cites.
Understanding Natural Language Commands for Robotic Navigation and Mobile Manipulation
Stefanie Tellex, Thomas Kollar, Steven Dickerson, Matthew R. Walter, Ashis Gopal Banerjee, Seth Teller, and Nicholas Roy · 2011
Earlier work this paper cites.
A Bayesian approach for policy learning from trajectory preference queries
Aaron Wilson, Alan Fern, and Prasad Tadepalli · 2012
Earlier work this paper cites.
Weakly supervised learning of semantic parsers for mapping instructions to actions
Yoav Artzi and Luke Zettlemoyer · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Alignment-Based Compositional Semantics for Instruction Following
Jacob Andreas and Dan Klein · 2015
Cited alongside, same era.
Grounding english commands to reward functions
James MacGlashan, Monica Babes-Vroman, Marie desJardins, Michael L. Littman, Smaranda Muresan, Shawn Squire, Stefanie Tellex, Dilip Arumugam, and Lei Yang · 2015
Cited alongside, same era.
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Reinforcement Learning with Unsupervised Auxiliary Tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z. Leibo, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Mapping Instructions and Visual Observations to Actions with Reinforcement Learning
Dipendra Misra, John Langford, and Yoav Artzi · 2017
Later among the works it cites.
Zero-Shot Task Generalization with Multi-Task Deep Reinforcement Learning
Junhyuk Oh, Satinder Singh, Honglak Lee, and Pushmeet Kohli · 2017
Later among the works it cites.
FiLM: Visual Reasoning with a General Conditioning Layer
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron Courville · 2017
Later among the works it cites.
Third-Person Imitation Learning
Bradly C. Stadie, Pieter Abbeel, and Ilya Sutskever · 2017
Later among the works it cites.
Deep TAMER: Interactive agent shaping in high-dimensional state spaces
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hongyuan Mei, Mohit Bansal, and Matthew R. Walter · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
HoME: a Household Multimodal Environment
Simon Brodeur, Ethan Perez, Ankesh Anand, Florian Golemo, Luca Celotti, Florian Strub, Jean Rouat, Hugo Larochelle, and Aaron Courville · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Misha Denil, Sergio Gómez Colmenarejo, Serkan Cabi, David Saxton, and Nando de Freitas · 2017
Cited alongside, same era.
Grounded Language Learning in a Simulated 3d World
Karl Moritz Hermann, Felix Hill, Simon Green, Fumin Wang, Ryan Faulkner, Hubert Soyer, David Szepesvari, Wojciech Marian Czarnecki, Max Jaderberg, Denis Teplyashin, Marcus Wainwright, Chris Apps, Demis Hassabis, and Phil Blunsom · 2017
Cited alongside, same era.
Representation Learning for Grounded Spatial Reasoning
Michael Janner, Karthik Narasimhan, and Regina Barzilay · 2017
Cited alongside, same era.
Garrett Warnell, Nicholas Waytowich, Vernon Lawhern, and Peter Stone · 2017
Later among the works it cites.
Gated-Attention Architectures for Task-Oriented Language Grounding
Devendra Singh Chaplot, Kanthashree Mysore Sathyendra, Rama Kumar Pasumarthi, Dheeraj Rajagopal, and Ruslan Salakhutdinov · 2018
Closest in time.
From Language to Goals: Inverse Reinforcement Learning for Vision-Based Instruction Following
Justin Fu, Anoop Korattikara, Sergey Levine, and Sergio Guadarrama · 2018
Closest in time.
Synthesizing Programs for Images using Reinforced Adversarial Learning
Yaroslav Ganin, Tejas Kulkarni, Igor Babuschkin, S. M. Ali Eslami, and Oriol Vinyals · 2018
Closest in time.
Zero-shot visual imitation
Deepak Pathak, Parsa Mahmoudieh, Guanghao Luo, Pulkit Agrawal, Dian Chen, Yide Shentu, Evan Shelhamer, Jitendra Malik, Alexei A. Efros, and Trevor Darrell · 2018
Closest in time.
Learning to parse natural language to grounded reward functions with weak supervision
Edward C Williams, Nakul Gopalan, Mine Rhee, and Stefanie Tellex · 2018
Closest in time.
Building Generalizable Agents with a Realistic and Rich 3d Environment
Yi Wu, Yuxin Wu, Georgia Gkioxari, and Yuandong Tian · 2018
Closest in time.
Interactive Grounded Language Acquisition and Generalization in 2d Environment
Haonan Yu, Haochao Zhang, and Wei Xu · 2018
Closest in time.