Fetching the paper…
Reading the bibliography…
We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain.
On the theory of the brownian motion
Uhlenbeck, George E and Ornstein, Leonard S · 1930
Earlier work this paper cites.
Q-learning
Watkins, Christopher JCH and Dayan, Peter · 1992
Earlier work this paper cites.
Adaptive critic designs
Prokhorov, Danil V, Wunsch, Donald C, et al · 1997
Earlier work this paper cites.
A generalized iterative lqg method for locally-optimal feedback control of constrained nonlinear stochastic systems
Todorov, Emanuel and Li, Weiwei · 2005
Earlier work this paper cites.
Real-time reinforcement learning by sequential actor–critics and experience replay
Wawrzyński, Paweł · 2009
Earlier work this paper cites.
States versus rewards: dissociable neural prediction error signals underlying model-based and model-free reinforcement learning
Gläscher, Jan, Daw, Nathaniel, Dayan, Peter, and O’Doherty, John P · 2010
Earlier work this paper cites.
Double q-learning
Hasselt, Hado V · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, Marc and Rasmussen, Carl E · 2011
Earlier work this paper cites.
Deep sparse rectifier networks
Glorot, Xavier, Bordes, Antoine, and Bengio, Yoshua · 2011
Earlier work this paper cites.
Reinforcement learning in feedback control
Hafner, Roland and Riedmiller, Martin · 2011
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Krizhevsky, Alex, Sutskever, Ilya, and Hinton, Geoffrey E · 2012
Cited alongside, same era.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Tassa, Yuval, Erez, Tom, and Todorov, Emanuel · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, Emanuel, Erez, Tom, and Tassa, Yuval · 2012
Cited alongside, same era.
A survey on policy search for robotics
Deisenroth, Marc Peter, Neumann, Gerhard, Peters, Jan, et al · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Graves, Alex, Antonoglou, Ioannis, Wierstra, Daan, and Riedmiller, Martin · 2013
Autonomous reinforcement learning with experience replay
Wawrzyński, Paweł and Tanwani, Ajay Kumar · 2014
Later among the works it cites.
Compatible value gradients for reinforcement learning of continuous deep policies
Balduzzi, David and Ghifary, Muhammad · 2015
Closest in time.
Memory-based control with recurrent neural networks
Heess, N., Hunt, J. J, Lillicrap, T. P, and Silver, D · 2015
Closest in time.
Learning continuous control policies by stochastic value gradients
Heess, Nicolas, Wayne, Gregory, Silver, David, Lillicrap, Tim, Erez, Tom, and Tassa, Yuval · 2015
Closest in time.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, Sergey and Szegedy, Christian · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Cited alongside, same era.
Evolving deep unsupervised convolutional networks for vision-based reinforcement learning
Koutník, Jan, Schmidhuber, Jürgen, and Gomez, Faustino · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, David, Lever, Guy, Heess, Nicolas, Degris, Thomas, Wierstra, Daan, and Riedmiller, Martin · 2014
Cited alongside, same era.
Online evolution of deep convolutional network for vision-based reinforcement learning
Koutník, Jan, Schmidhuber, Jürgen, and Gomez, Faustino
Cited in the paper.
Gradient estimation using stochastic computation graphs
Schulman, John, Heess, Nicolas, Weber, Theophane, and Abbeel, Pieter
Cited in the paper.
Trust region policy optimization
Schulman, John, Levine, Sergey, Moritz, Philipp, Jordan, Michael I, and Abbeel, Pieter
Cited in the paper.
End-to-end training of deep visuomotor policies
Levine, Sergey, Finn, Chelsea, Darrell, Trevor, and Abbeel, Pieter · 2015
Closest in time.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Closest in time.
From pixels to torques: Policy learning with deep dynamical models
Wahlström, Niklas, Schön, Thomas B, and Deisenroth, Marc Peter · 2015
Closest in time.
Control policy with autocorrelated noise in reinforcement learning for robotics
Wawrzyński, Paweł · 2015
Closest in time.