Fetching the paper…
Reading the bibliography…
Solving complex, temporally-extended tasks is a long-standing problem in reinforcement learning (RL).
Visual hindsight experience replay
Himanshu Sahni, Toby Buckley, Pieter Abbeel, and Ilya Kuzovkin · 1901
Earlier work this paper cites.
Skew-fit: State-covering self-supervised reinforcement learning
Vitchyr H. Pong, Murtaza Dalal, Steven Lin, Ashvin Nair, Shikhar Bahl, and Sergey Levine · 1903
Earlier work this paper cites.
Logic and conversation
H Paul Grice · 1975
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E Hinton · 1993
Earlier work this paper cites.
Hierarchical learning in stochastic domains: Preliminary results
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Ronald Parr and Stuart J Russell · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
Temporal abstraction in reinforcement learning
Doina Precup · 2000
Earlier work this paper cites.
Learning options in reinforcement learning
Martin Stolle and Doina Precup · 2002
Earlier work this paper cites.
Dynamic abstraction in reinforcement learning via clustering
Shie Mannor, Ishai Menache, Amit Hoze, and Uri Klein · 2004
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Nuttapong Chentanez, Andrew G Barto, and Satinder P Singh · 2005
Earlier work this paper cites.
Language and thought
Lila Gleitman and Anna Papafragou · 2005
Earlier work this paper cites.
Re-evaluation the role of bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn · 2006
Earlier work this paper cites.
Building portable options: Skill transfer in reinforcement learning
George Konidaris and Andrew G Barto · 2007
Earlier work this paper cites.
Hierarchical relative entropy policy search
Christian Daniel, Gerhard Neumann, and Jan Peters · 2012
Earlier work this paper cites.
Bootstrapping in a language of thought: A formal model of numerical concept learning
Steven T Piantadosi, Joshua B Tenenbaum, and Noah D Goodman · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches, 2014
Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks, 2014
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
Deep reinforcement learning with a natural language action space
Ji He, Jianshu Chen, Xiaodong He, Jianfeng Gao, Lihong Li, Li Deng, and Mari Ostendorf · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Language understanding for text-based games using deep reinforcement learning
Karthik Narasimhan, Tejas Kulkarni, and Regina Barzilay · 2015
Cited alongside, same era.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Deep reinforcement learning with double q-learning, 2015
Hado van Hasselt, Arthur Guez, and David Silver · 2015
Cited alongside, same era.
Modular multitask reinforcement learning with policy sketches, 2016
Jacob Andreas, Dan Klein, and Sergey Levine · 2016
Cited alongside, same era.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton van den Hengel · 2018
Later among the works it cites.
Learning to understand goal specifications by modelling reward, 2018
Dzmitry Bahdanau, Felix Hill, Jan Leike, Edward Hughes, Pushmeet Kohli, and Edward Grefenstette · 2018
Later among the works it cites.
Relational inductive biases, deep learning, and graph networks
Peter W Battaglia, Jessica B Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, et al · 2018
Later among the works it cites.
Neural modular control for embodied question answering
Abhishek Das, Georgia Gkioxari, Stefan Lee, Devi Parikh, and Dhruv Batra · 2018
Later among the works it cites.
Meta learning shared hierarchies
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicolas Heess, Greg Wayne, Yuval Tassa, Timothy Lillicrap, Martin Riedmiller, and David Silver · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Learning with latent language, 2017
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Cited alongside, same era.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Cited alongside, same era.
Kevin Frans, Jonathan Ho, Xi Chen, Pieter Abbeel, and John Schulman · 2018
Later among the works it cites.
Speaker-follower models for vision-and-language navigation
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning, 2018
Ofir Nachum, Shixiang Gu, Honglak Lee, and Sergey Levine · 2018
Later among the works it cites.
Visual reinforcement learning with imagined goals
Ashvin Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine · 2018
Later among the works it cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville · 2018
Later among the works it cites.
Temporal difference models: Model-free deep rl for model-based control, 2018
Vitchyr Pong, Shixiang Gu, Murtaza Dalal, and Sergey Levine · 2018
Later among the works it cites.
A structured review of the validity of bleu
Ehud Reiter · 2018
Later among the works it cites.
Policy search in continuous action domains: an overview
Olivier Sigaud and Freek Stulp · 2018
Later among the works it cites.
Bleu is not suitable for the evaluation of text simplification
Elior Sulem, Omri Abend, and Ari Rappoport · 2018
Later among the works it cites.
Action branching architectures for deep reinforcement learning
Arash Tavakoli, Fabio Pardo, and Petar Kormushev · 2018
Later among the works it cites.
BabyAI: First steps towards grounded language learning with a human in the loop
Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien Huu Nguyen, and Yoshua Bengio · 2019
Closest in time.
Meta-learning language-guided policy learning
John D Co-Reyes, Abhishek Gupta, Suvansh Sanjeev, Nick Altieri, John DeNero, Pieter Abbeel, and Sergey Levine · 2019
Closest in time.
From language to goals: Inverse reinforcement learning for vision-based instruction following
Justin Fu, Anoop Korattikara, Sergey Levine, and Sergio Guadarrama · 2019
Closest in time.
Hierarchical reinforcement learning with hindsight
Andrew Levy, Robert Platt, and Kate Saenko · 2019
Closest in time.
A survey of reinforcement learning informed by natural language
Jelena Luketina, Nantas Nardelli, Gregory Farquhar, Jakob Foerster, Jacob Andreas, Edward Grefenstette, Shimon Whiteson, and Tim Rocktäschel · 2019
Closest in time.
Near-optimal representation learning for hierarchical reinforcement learning
Ofir Nachum, Shixiang Gu, Honglak Lee, and Sergey Levine · 2019
Closest in time.
ACTRCE: Augmenting experience via teacher’s advice, 2019
Yuhuai Wu, Harris Chan, Jamie Kiros, Sanja Fidler, and Jimmy Ba · 2019
Closest in time.