Fetching the paper…
Reading the bibliography…
Very recently proximal policy optimization (PPO) algorithms have been proposed as first-order optimization methods for effective reinforcement learning.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Calculus of variations
I. M. Gelfand, R. A. Silverman, et al · 2000
Earlier work this paper cites.
Asymptopia: an exposition of statistical asymptotic theory. 2000
D. Pollard · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Learning tetris using the noisy cross-entropy method
I. Szita and A. Lörincz · 2006
Earlier work this paper cites.
Natural actor-critic algorithms
S. Bhatnagar, R. S. Sutton, M. Ghavamzadeh, and M. Lee · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
The Arcade Learning Environment - An Evaluation Platform for General Agents (Extended Abstract)
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2015
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Z. Wang, T. Schaul, M. Hessel, H. Van Hasselt, M. Lanctot, and N. De Freitas · 2015
Cited alongside, same era.
OpenAI Gym
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Benchmarking Deep Reinforcement Learning for Continuous Control
Y. Duan, X. Chen, J. Schulman, and P. Abbeel · 2016
Cited alongside, same era.
Successor features for transfer in reinforcement learning
A. Barreto, W. Dabney, R. Munos, J. Hunt, T. Schaul, D. Silver, and H. P. van Hasselt · 2017
Later among the works it cites.
S. Gu, T. Lillicrap, Z. Ghahramani, R. E. Turner, B. Schölkopf, and S. Levine · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
O. Nachum, M. Norouzi, K. Xu, and D. Schuurmans · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Hausknecht and P. Stone · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Sample efficient actor-critic with experience replay
Z. Wang, V. Bapst, N. Heess, V. Mnih, R. Munos, K. Kavukcuoglu, and N. de Freitas · 2016
Cited alongside, same era.
A brief survey of deep reinforcement learning
K. Arulkumaran, M. P. Deisenroth, M. Brundage, and A. A. Bharath · 2017
Cited alongside, same era.
Later among the works it cites.
# exploration: A study of count-based exploration for deep reinforcement learning
H. Tang, R. Houthooft, D. Foote, A. Stooke, X. Chen, Y. Duan, J. Schulman, F. DeTurck, and P. Abbeel · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Y. Wu, E. Mansimov, R. B. Grosse, S. Liao, and J. Ba · 2017
Later among the works it cites.
Pybullet, a python module for physics simulation for games, robotics and machine learning
E. Coumans and Y. Bai · 2018
Closest in time.