Fetching the paper…
Reading the bibliography…
We propose a conceptually simple and lightweight framework for deep reinforcement learning that uses asynchronous gradient descent for optimization of deep neural network controllers.
Distributed dynamic programming
Bertsekas, Dimitri P · 1982
Earlier work this paper cites.
Learning from delayed rewards
Watkins, Christopher John Cornish Hellaby · 1989
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, Ronald J and Peng, Jing · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R.J · 1992
Earlier work this paper cites.
On-line q-learning using connectionist systems
Rummery, Gavin A and Niranjan, Mahesan · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and q-learning
Tsitsiklis, John N · 1994
Earlier work this paper cites.
Incremental multi-step q-learning
Peng, Jing and Williams, Ronald J · 1996
Earlier work this paper cites.
Reinforcement Learning: an Introduction
Sutton, R. and Barto, A · 1998
Earlier work this paper cites.
Parallel and distributed evolutionary algorithms: A review
Tomassini, Marco · 1999
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Riedmiller, Martin · 2005
Earlier work this paper cites.
Parallel reinforcement learning with linear function approximation
Grounds, Matthew and Kudenko, Daniel · 2008
Cited alongside, same era.
Mapreduce for parallel reinforcement learning
Li, Yuxi and Schuurmans, Dale · 2011
Cited alongside, same era.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Recht, Benjamin, Re, Christopher, Wright, Stephen, and Niu, Feng · 2011
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, Marc G, Naddaf, Yavar, Veness, Joel, and Bowling, Michael · 2012
Cited alongside, same era.
Model-free reinforcement learning with continuous action in practice
Degris, Thomas, Pilarski, Patrick M, and Sutton, Richard S · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, Tijmen and Hinton, Geoffrey · 2012
End-to-end training of deep visuomotor policies
Levine, Sergey, Finn, Chelsea, Darrell, Trevor, and Abbeel, Pieter · 2015
Later among the works it cites.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P, Hunt, Jonathan J, Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A., Veness, Joel, Bellemare, Marc G., Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K., Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Later among the works it cites.
Massively parallel methods for deep reinforcement learning
Nair, Arun, Srinivasan, Praveen, Blackwell, Sam, Alcicek, Cagdas, Fearon, Rory, Maria, Alessandro De, Panneershelvam, Vedavyas, Suleyman, Mustafa, Beattie, Charles, Petersen, Stig, Legg, Shane, Mnih, Volodymyr, Kavukcuoglu, Koray, and Silver, David · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Graves, Alex, Antonoglou, Ioannis, Wierstra, Daan, and Riedmiller, Martin · 2013
Cited alongside, same era.
Torcs: The open racing car simulator, v1.3.5, 2013
Wymann, B., Espié, E., Guionneau, C., Dimitrakakis, C., Coulom, R., and Sumner, A · 2013
Cited alongside, same era.
Evolving deep unsupervised convolutional networks for vision-based reinforcement learning
Koutník, Jan, Schmidhuber, Jürgen, and Gomez, Faustino · 2014
Cited alongside, same era.
Distributed deep q-learning
Chavez, Kevin, Ong, Hao Yi, and Hong, Augustus · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, John, Levine, Sergey, Moritz, Philipp, Jordan, Michael I, and Abbeel, Pieter
Cited in the paper.
High-dimensional continuous control using generalized advantage estimation
Schulman, John, Moritz, Philipp, Levine, Sergey, Jordan, Michael, and Abbeel, Pieter
Cited in the paper.
Schaul, Tom, Quan, John, Antonoglou, Ioannis, and Silver, David · 2015
Later among the works it cites.
MuJoCo: Modeling, Simulation and Visualization of Multi-Joint Dynamics with Contact (ed 1.0)
Todorov, E · 2015
Later among the works it cites.
Deep reinforcement learning with double q-learning
Van Hasselt, Hado, Guez, Arthur, and Silver, David · 2015
Later among the works it cites.
True Online Temporal-Difference Learning
van Seijen, H., Rupam Mahmood, A., Pilarski, P. M., Machado, M. C., and Sutton, R. S · 2015
Later among the works it cites.
Dueling Network Architectures for Deep Reinforcement Learning
Wang, Z., de Freitas, N., and Lanctot, M · 2015
Later among the works it cites.
Increasing the action gap: New operators for reinforcement learning
Bellemare, Marc G., Ostrovski, Georg, Guez, Arthur, Thomas, Philip S., and Munos, Rémi · 2016
Closest in time.