Fetching the paper…
Reading the bibliography…
The tasks that an agent will need to solve often are not known during training.
Strips: A new approach to the application of theorem proving to problem solving
Richard E. Fikes and Nils J. Nilsson · 1971
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E. Hinton · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Finding structure in reinforcement learning
Sebastian Thrun and Anton Schwartz · 1994
Earlier work this paper cites.
Exploiting structure in policy construction
Craig Boutilier, Richard Dearden, and Moises Goldszmidt · 1995
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Ronald Parr and Stuart Russell · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Stochastic dynamic programming with factored representations
Craig Boutilier, Richard Dearden, and Moisés Goldszmidt · 2000
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G. Dietterich · 2000
Earlier work this paper cites.
Envelope-based planning in relational mdps
Natalia Hernandez-Gardiol and Leslie Pack Kaelbling · 2003
Earlier work this paper cites.
A Survey of Reinforcement Learning in Relational Domains
M. van Otterlo · 2005
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
Carlos Diuk, Andre Cohen, and Michael L. Littman · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
Horde: a scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S. Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M. Pilarski, Adam White, and Doina Precup · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Goal-based action priors
David Abel, D. Ellis Hershkowitz, Gabriel Barth-Maron, Stephen Brawner, Kevin O’Farrell, James MacGlashan, and Stefanie Tellex · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Universal Value Function Approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Cited alongside, same era.
Mazebase: A sandbox for learning from games
Sainbayar Sukhbaatar, Arthur Szlam, Gabriel Synnaeve, Soumith Chintala, and Rob Fergus · 2015
Cited alongside, same era.
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Later among the works it cites.
Misha Denil, Sergio Gomez Colmenarejo, Serkan Cabi, David Saxton, and Nando de Freitas · 2017
Later among the works it cites.
When Waiting is not an Option : Learning Options with a Deliberation Cost
J. Harb, P.-L. Bacon, M. Klissarov, and D. Precup · 2017
Later among the works it cites.
Gradient episodic memory for continuum learning
David Lopez-Paz and Marc’Aurelio Ranzato · 2017
Later among the works it cites.
A laplacian framework for option discovery in reinforcement learning
Marlos C. Machado, Marc G. Bellemare, and Michael H. Bowling · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to poke by poking: Experiential learning of intuitive physics
Pulkit Agrawal, Ashvin V Nair, Pieter Abbeel, Jitendra Malik, and Sergey Levine · 2016
Cited alongside, same era.
Charles Blundell, Benigno Uria, Alexander Pritzel, Yazhe Li, Avraham Ruderman, Joel Z. Leibo, Jack Rae, Daan Wierstra, and Demis Hassabis · 2016
Cited alongside, same era.
Learning to act by predicting the future
Alexey Dosovitskiy and Vladlen Koltun · 2016
Cited alongside, same era.
Using task features for zero-shot knowledge transfer in lifelong learning
David Isele, Mohammad Rostami, and Eric Eaton · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D. Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Cited alongside, same era.
Torchcraft: a library for machine learning research on real-time strategy games
Gabriel Synnaeve, Nantas Nardelli, Alex Auvolat, Soumith Chintala, Timothée Lacroix, Zeming Lin, Florian Richoux, and Nicolas Usunier · 2016
Cited alongside, same era.
Later among the works it cites.
Zero-shot task generalization with multi-task deep reinforcement learning
Junhyuk Oh, Satinder P. Singh, Honglak Lee, and Pushmeet Kohli · 2017
Later among the works it cites.
Neural episodic control
Alexander Pritzel, Benigno Uria, Sriram Srinivasan, Adrià Puigdomènech Badia, Oriol Vinyals, Demis Hassabis, Daan Wierstra, and Charles Blundell · 2017
Later among the works it cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sainbayar Sukhbaatar, Ilya Kostrikov, Arthur Szlam, and Rob Fergus · 2017
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
Chen Tessler, Shahar Givony, Tom Zahavy, Daniel J. Mankowitz, and Shie Mannor · 2017
Later among the works it cites.
Independently controllable factors
Valentin Thomas, Jules Pondard, Emmanuel Bengio, Marc Sarfati, Philippe Beaudoin, Marie-Jean Meurs, Joelle Pineau, Doina Precup, and Yoshua Bengio · 2017
Later among the works it cites.
Hybrid reward architecture for reinforcement learning
Harm van Seijen, Mehdi Fatemi, Joshua Romoff, Romain Laroche, Tavian Barnes, and Jeffrey Tsang · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu · 2017
Later among the works it cites.
From skills to symbols: Learning symbolic representations for abstract high-level planning
George Konidaris, Leslie Pack Kaelbling, and Tomás Lozano-Pérez · 2018
Closest in time.