Fetching the paper…
Reading the bibliography…
Recent work has shown that augmenting environments with language descriptions improves policy learning.
An analysis of model-based interval estimation for markov decision processes
Alexander L. Strehl and Michael L. Littman · 2008
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Observations on expert play in nethack
Chris Street · 2013
Earlier work this paper cites.
Language understanding for text-based games using deep reinforcement learning
Karthik Narasimhan, Tejas Kulkarni, and Regina Barzilay · 2015
Earlier work this paper cites.
Grounded Action Transformation for Robot Learning in Simulation
Josiah P. Hanna and Peter Stone · 2017
Earlier work this paper cites.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z. Leibo, David Silver, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Learning human behaviors from motion capture by adversarial imitation
Josh Merel, Yuval Tassa, Dhruva TB, Sriram Srinivasan, Jay Lemmon, Ziyu Wang, Greg Wayne, and Nicolas Heess · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Third-person imitation learning
Bradly C. Stadie, Pieter Abbeel, and Ilya Sutskever · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel · 2018
Earlier work this paper cites.
Touchdown: Natural language navigation and spatial reasoning in visual street environments
Howard Chen, Alane Suhr, Dipendra Kumar Misra, Noah Snavely, and Yoav Artzi · 2018
Earlier work this paper cites.
IMPALA: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Cited alongside, same era.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Decoupling strategy and generation in negotiation dialogues
He He, Derek Chen, Anusha Balakrishnan, and Percy Liang · 2018
Cited alongside, same era.
Internal Model from Observations for Reward Shaping
Daiki Kimura, Subhajit Chaudhury, Ryuki Tachibana, and Sakyasingha Dasgupta · 2018
Cited alongside, same era.
Reinforcement learning on web interfaces using workflow-guided exploration
Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang · 2018
Cited alongside, same era.
Behavioral cloning from observation
Faraz Torabi, Garrett Warnell, and Peter Stone · 2018
Room-Across-Room: Multilingual vision-and-language navigation with dense spatiotemporal grounding
Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge · 2020
Later among the works it cites.
The NetHack learning environment
Heinrich Küttler, Nantas Nardelli, Alexander H Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rocktäschel · 2020
Later among the works it cites.
RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated Environments
Roberta Raileanu and Tim Rocktäschel · 2020
Later among the works it cites.
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox · 2020
Later among the works it cites.
RTFM: Generalising to new environment dynamics via reading
Victor Zhong, Tim Rocktäschel, and Edward Grefenstette · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov · 2019
Cited alongside, same era.
Imitating Latent Policies from Observation
Ashley D Edwards, Himanshu Sahni, Yannick Schroecker, and Charles L Isbell · 2019
Cited alongside, same era.
Hybrid reinforcement learning with expert state sequences
Xiaoxiao Guo, Shiyu Chang, Mo Yu, Gerald Tesauro, and Murray Campbell · 2019
Cited alongside, same era.
TorchBeast: A PyTorch Platform for Distributed RL
Heinrich Küttler, Nantas Nardelli, Thibaut Lavril, Marco Selvatici, Viswanath Sivakumar, Tim Rocktäschel, and Edward Grefenstette · 2019
Cited alongside, same era.
Imitation learning from observations by minimizing inverse dynamics disagreement
Chao Yang, Xiaojian Ma, Wenbing Huang, Fuchun Sun, Huaping Liu, Junzhou Huang, and Chuang Gan · 2019
Cited alongside, same era.
Grounding language to entities and dynamics for generalization in reinforcement learning
Austin W. Hanjie, Victor Zhong, and Karthik Narasimhan · 2021
Later among the works it cites.
ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht · 2021
Later among the works it cites.
SILG: The multi-environment symbolic interactive language grounding benchmark
Victor Zhong, Austin W Hanjie, Sida I Wang, Karthik Narasimhan, and Luke Zettlemoyer · 2021
Later among the works it cites.
moolib: A platform for distributed RL
Vegard Mella, Eric Hambro, Danielle Rothermel, and Heinrich Küttler · 2022
Closest in time.
Semantic exploration from language abstractions and pretrained representations, 2022
Allison C. Tam, Neil C. Rabinowitz, Andrew K. Lampinen, Nicholas A. Roy, Stephanie C. Y. Chan, DJ Strouse, Jane X. Wang, Andrea Banino, and Felix Hill · 2022
Closest in time.