Fetching the paper…
Reading the bibliography…
Learning robust value functions given raw observations and rewards is now possible with model-free and model-based deep reinforcement learning algorithms.
Improving generalization for temporal difference learning: The successor representation
Peter Dayan · 1993
Earlier work this paper cites.
Temporal difference learning of position evaluation in the game of go
Nicol N Schraudolph, Peter Dayan, and Terrence J Sejnowski · 1994
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Normalized cuts and image segmentation
Jianbo Shi and Jitendra Malik · 2000
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
Amy McGovern and Andrew G Barto · 2001
Earlier work this paper cites.
A distributed representation of temporal context
Marc W Howard and Michael J Kahana · 2002
Earlier work this paper cites.
Q-cut—dynamic discovery of sub-goals in reinforcement learning
Ishai Menache, Shie Mannor, and Nahum Shimkin · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
Andrew G Barto and Sridhar Mahadevan · 2003
Earlier work this paper cites.
Subgoal discovery for hierarchical reinforcement learning using learned policies
Sandeep Goel and Manfred Huber · 2003
Earlier work this paper cites.
Dynamic abstraction in reinforcement learning via clustering
Shie Mannor, Ishai Menache, Amit Hoze, and Uri Klein · 2004
Earlier work this paper cites.
Identifying useful subgoals in reinforcement learning by local graph partitioning
Özgür Şimşek, Alicia P Wolfe, and Andrew G Barto · 2005
Earlier work this paper cites.
Skill discovery in continuous reinforcement learning domains using skill chaining
George Konidaris and Andre S Barreto · 2009
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber · 2010
Earlier work this paper cites.
The successor representation and temporal context
Samuel J Gershman, Christopher D Moore, Michael T Todd, Kenneth A Norman, and Per B Sederberg · 2012
Earlier work this paper cites.
Model-based hierarchical reinforcement learning and human action control
Matthew Botvinick and Ari Weinstein · 2014
Cited alongside, same era.
The algorithmic anatomy of model-based evaluation
Nathaniel D Daw and Peter Dayan · 2014
Cited alongside, same era.
Design principles of the hippocampal cognitive map
Kimberly L Stachenfeld, Matthew Botvinick, and Samuel J Gershman · 2014
Cited alongside, same era.
Universal option models
Csaba Szepesvari, Richard S Sutton, Joseph Modayil, Shalabh Bhatnagar, et al · 2014
Cited alongside, same era.
Attractor network dynamics enable preplay and rapid path planning in maze–like environments
Dane S Corneil and Wulfram Gerstner · 2015
Cited alongside, same era.
Binding via reconstruction clustering
Klaus Greff, Rupesh Kumar Srivastava, and Jürgen Schmidhuber · 2015
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Later among the works it cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C Stadie, Sergey Levine, and Pieter Abbeel · 2015
Later among the works it cites.
Mazebase: A sandbox for learning from games
Sainbayar Sukhbaatar, Arthur Szlam, Gabriel Synnaeve, Soumith Chintala, and Rob Fergus · 2015
Later among the works it cites.
Attend, infer, repeat: Fast scene understanding with generative models
SM Eslami, Nicolas Heess, Theophane Weber, Yuval Tassa, Koray Kavukcuoglu, and Geoffrey E Hinton · 2016
Closest in time.
Vizdoom: A doom-based ai research platform for visual reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Draw: A recurrent neural network for image generation
Karol Gregor, Ivo Danihelka, Alex Graves, and Daan Wierstra · 2015
Cited alongside, same era.
Efficient inference in occlusion-aware generative models of images
Jonathan Huang and Kevin Murphy · 2015
Cited alongside, same era.
Deep convolutional inverse graphics network
Tejas D Kulkarni, William F Whitney, Pushmeet Kohli, and Josh Tenenbaum · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo Jimenez Rezende · 2015
Cited alongside, same era.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, et al · 2015
Cited alongside, same era.
Michał Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Jaśkowski · 2016
Closest in time.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik R Narasimhan, Ardavan Saeedi, and Joshua B Tenenbaum · 2016
Closest in time.
Learning purposeful behaviour in the absence of rewards
Marlos C Machado and Michael Bowling · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Closest in time.
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro · 2016
Closest in time.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Closest in time.
One-shot generalization in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, Ivo Danihelka, Karol Gregor, and Daan Wierstra · 2016
Closest in time.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Closest in time.
Understanding visual concepts with continuation learning
William F Whitney, Michael Chang, Tejas Kulkarni, and Joshua B Tenenbaum · 2016
Closest in time.