Fetching the paper…
Reading the bibliography…
This paper presents an actor-critic deep reinforcement learning agent with experience replay that is stable, sample efficient, and performs remarkably well on challenging environments, including the discrete 57-game Atari domain and several continuous control problems.
Self-improving reactive agents based on reinforcement learning, planning and teaching
L.J. Lin · 1992
Earlier work this paper cites.
Off-policy policy search
N. Meuleau, L. Peshkin, L. P. Kaelbling, and K. Kim · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup, R. S. Sutton, and S. Singh · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. Mcallester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Real-time reinforcement learning by sequential actor–critics and experience replay
P. Wawrzyński · 2009
Earlier work this paper cites.
On a connection between importance sampling and the likelihood ratio policy gradient
T. Jie and P. Abbeel · 2010
Earlier work this paper cites.
Off-policy actor-critic
T. Degris, M. White, and R. S. Sutton · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Guided policy search
S. Levine and V. Koltun · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa · 2015
Cited alongside, same era.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2015
Cited alongside, same era.
OpenAI Gym
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Closest in time.
Q ( λ \lambda ) with off-policy corrections
Anna Harutyunyan, Marc G Bellemare, Tom Stepleton, and Remi Munos · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. Puigdomènech Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Closest in time.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. G. Bellemare · 2016
Closest in time.
Control of memory, active perception, and action in Minecraft
J. Oh, V. Chockalingam, S. P. Singh, and H. Lee · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Lillicrap, J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Cited alongside, same era.
Language understanding for text-based games using deep reinforcement learning
K. Narasimhan, T. Kulkarni, and R. Barzilay · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz
Cited in the paper.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. I. Jordan, and P. Abbeel
Cited in the paper.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Closest in time.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C.J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Closest in time.
Dueling network architectures for deep reinforcement learning
Z. Wang, T. Schaul, M. Hessel, H. van Hasselt, M. Lanctot, and N. de Freitas · 2016
Closest in time.