Fetching the paper…
Reading the bibliography…
In reinforcement learning, policy gradient algorithms optimize the policy directly and rely on sampling efficiently an environment.
Neuronlike adaptive elements that can solve difficult learning control problems
A. Andrew, R. Sutton, and C. Anderson · 1983
Earlier work this paper cites.
Cautionary Note about R 2 R^{2}
T. Kvålseth · 1985
Earlier work this paper cites.
Dynamic reinforcement driven error propagation networks with application to game playing
A. Robinson and F. Fallside · 1989
Earlier work this paper cites.
Neural networks for control and system identification
P. Werbos · 1989
Earlier work this paper cites.
Learning to generate artificial fovea trajectories for target detection
J. Schmidhuber and R. Huber · 1991
Earlier work this paper cites.
Curious model-building control systems
J. Schmidhuber · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. Williams · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. Puterman · 1994
Earlier work this paper cites.
Natural gradient works efficiently in learning
S. Amari · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Temporal Abstraction in Reinforcement Learning
D. Precup · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. Sutton, D. McAllester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
A natural policy gradient
S. Kakade · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
S. Kakade · 2003
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
J. Peters and S. Schaal · 2008
Cited alongside, same era.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
M. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
S. Levine and P. Abbeel · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Roboschool, 2017
O. Klimov and J. Schulman · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
O. Nachum, M. Norouzi, K. Xu, and D. Schuurmans · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Sample efficient actor-critic with experience replay
Z. Wang, V. Bapst, N. Heess, V. Mnih, R. Munos, K. Kavukcuoglu, and N. de Freitas · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Y. Wu, E. Mansimov, R. Grosse, S. Liao, and J. Ba · 2017
Later among the works it cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa · 2015
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. Lillicrap, J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2016
Cited alongside, same era.
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
Later among the works it cites.
Recurrent world models facilitate policy evolution
D. Ha and J. Schmidhuber · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Later among the works it cites.
Stable baselines, 2018
A. Hill, A. Raffin, M. Ernestus, A. Gleave, A. Kanervisto, R. Traore, P. Dhariwal, C. Hesse, O. Klimov, A. Nichol, M. Plappert, A. Radford, J. Schulman, S. Sidor, and Y. Wu · 2018
Later among the works it cites.
Deep Reinforcement Learning Hands-On
M. Lapan · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. Sutton and A. Barto · 2018
Later among the works it cites.
MERL: Multi-Head Reinforcement Learning
Y. Flet-Berliac and P. Preux · 2019
Closest in time.
Learning to predict without looking ahead: World models without forward prediction
D. Freeman, L. Metz, and D. Ha · 2019
Closest in time.