Fetching the paper…
Reading the bibliography…
Deep reinforcement learning has achieved many impressive results in recent years.
Information processing in dynamical systems: Foundations of harmony theory
Paul Smolensky · 1986
Earlier work this paper cites.
Learning stochastic feedforward networks
Radford M Neal · 1990
Earlier work this paper cites.
Curious model-building control systems
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Connectionist learning of belief networks
Radford M Neal · 1992
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Ronald Parr and Stuart Russell · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Geoffrey E Hinton · 2002
Earlier work this paper cites.
Learning options in reinforcement learning
Martin Stolle and Doina Precup · 2002
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Nuttapong Chentanez, Andrew G Barto, and Satinder P Singh · 2004
Earlier work this paper cites.
Dynamic abstraction in reinforcement learning via clustering
Shie Mannor, Ishai Menache, Amit Hoze, and Uri Klein · 2004
Earlier work this paper cites.
Learning movement primitives
Stefan Schaal, Jan Peters, Jun Nakanishi, and Auke Ijspeert · 2005
Earlier work this paper cites.
Identifying useful subgoals in reinforcement learning by local graph partitioning
Özgür Şimşek, Alicia P Wolfe, and Andrew G Barto · 2005
Earlier work this paper cites.
A generalized iterative lqg method for locally-optimal feedback control of constrained nonlinear stochastic systems
Emanuel Todorov and Weiwei Li · 2005
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh · 2006
Earlier work this paper cites.
Building portable options: Skill transfer in reinforcement learning
George Konidaris and Andrew G Barto · 2007
Earlier work this paper cites.
Multi-task reinforcement learning: a hierarchical bayesian approach
Aaron Wilson, Alan Fern, Soumya Ray, and Prasad Tadepalli · 2007
Earlier work this paper cites.
Representational power of restricted boltzmann machines and deep belief networks
Nicolas Le Roux and Yoshua Bengio · 2008
Cited alongside, same era.
Transfer learning for reinforcement learning domains: A survey
Matthew E Taylor and Peter Stone · 2009
Cited alongside, same era.
Bayesian multi-task reinforcement learning
Alessandro Lazaric and Mohammad Ghavamzadeh · 2010
Cited alongside, same era.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber · 2010
Cited alongside, same era.
Intrinsically motivated hierarchical skill learning in structured environments
Christopher M Vigorito and Andrew G Barto · 2010
Cited alongside, same era.
Autonomous skill acquisition on a mobile manipulator
George Konidaris, Scott Kuindersma, Roderic A Grupen, and Andrew G Barto · 2011
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz · 2015
Later among the works it cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Later among the works it cites.
The option-critic architecture
Pierre-Luc Bacon and Doina Precup · 2016
Later among the works it cites.
Unifying count-based exploration and intrinsic motivation
Marc G Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Later among the works it cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical relative entropy policy search
Christian Daniel, Gerhard Neumann, and Jan Peters · 2012
Cited alongside, same era.
Gaussian-bernoulli deep boltzmann machine
Kyung Hyun Cho, Tapani Raiko, and Alexander Ilin · 2013
Cited alongside, same era.
Autonomous reinforcement learning with hierarchical reps
Christian Daniel, Gerhard Neumann, and Jan Peters · 2013
Cited alongside, same era.
Learning stochastic feedforward neural networks
Yichuan Tang and Ruslan R Salakhutdinov · 2013
Cited alongside, same era.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Xiaoxiao Guo, Satinder Singh, Honglak Lee, Richard L Lewis, and Xiaoshi Wang · 2014
Cited alongside, same era.
Generalization and exploration via randomized value functions
Ian Osband, Benjamin Van Roy, and Zheng Wen · 2014
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Later among the works it cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Later among the works it cites.
Learning and transfer of modulated locomotor controllers
Nicolas Heess, Greg Wayne, Yuval Tassa, Timothy Lillicrap, Martin Riedmiller, and David Silver · 2016
Later among the works it cites.
Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Later among the works it cites.
Strategic attentive writer for learning macro-actions
Volodymyr Mnih, John Agapiou, Simon Osindero, Alex Graves, Oriol Vinyals, Koray Kavukcuoglu, et al · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
On multiplicative integration with recurrent neural networks
Yuhuai Wu, Saizheng Zhang, Ying Zhang, Yoshua Bengio, and Ruslan Salakhutdinov · 2016
Later among the works it cites.
Learning modular neural network policies for multi-task and multi-robot transfer
Coline Devin, Abhishek Gupta, Trevor Darrell, Pieter Abbeel, and Sergey Levine · 2017
Closest in time.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Closest in time.
The concrete distribution: A continuous relaxation of discrete random variables
Chris J Maddison, Andriy Mnih, and Yee Whye Teh · 2017
Closest in time.