Fetching the paper…
Reading the bibliography…
Common policy gradient methods rely on the maximization of a sequence of surrogate functions.
Behaviour suite for reinforcement learning
Osband, I., Doron, Y., Hessel, M., Aslanides, J., Sezener, E., Saraiva, A., McKinney, K., Lattimore, T., Szepesvari, C., Singh, S., et al. (2019) · 1908
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J. (1991) · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, S. (2001) · 2001
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E. (2002) · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J. (2002) · 2002
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M. (2002) · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Beck, A. and Teboulle, M. (2003) · 2003
Earlier work this paper cites.
Mirror descent policy optimization
Tomar, M., Shani, L., Efroni, Y., and Ghavamzadeh, M. (2020) · 2005
Earlier work this paper cites.
Numerical Optimization
Nocedal, J. and Wright, S. J. (2006) · 2006
Earlier work this paper cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Cen, S., Cheng, C., Chen, Y., Wei, Y., and Chi, Y. (2020) · 2007
Earlier work this paper cites.
Revisiting design choices in proximal policy optimization
Hsu, C. C.-Y., Mendler-Dünner, C., and Hardt, M. (2020) · 2009
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Cited alongside, same era.
Local policy search in a convex space and conservative policy iteration as boosted policy search
Scherrer, B. and Geist, M. (2014) · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Bubeck, S. (2015) · 2015
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Heess, N., Wayne, G., Silver, D., Lillicrap, T., Erez, T., and Tassa, Y. (2015) · 2015
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Implementation matters in deep RL: A case study on PPO and TRPO
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A. (2019) · 2019
Later among the works it cites.
A theory of regularized Markov decision processes
Geist, M., Scherrer, B., and Pietquin, O. (2019) · 2019
Later among the works it cites.
On principled entropy exploration in policy optimization
Mei, J., Xiao, C., Huang, R., Schuurmans, D., and Müller, M. (2019) · 2019
Later among the works it cites.
Optimality and approximation with policy gradient methods in Markov decision processes
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. (2020) · 2020
Later among the works it cites.
An operator view of policy gradient methods
Ghosh, D., C Machado, M., and Le Roux, N. (2020) · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Openai baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., Wu, Y., and Zhokhov, P. (2017) · 2017
Cited alongside, same era.
A unified view of entropy-regularized Markov decision processes
Neu, G., Jonsson, A., and Gómez, V. (2017) · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M. A. (2018) · 2018
Cited alongside, same era.
Later among the works it cites.
Guided learning of nonconvex models through successive functional gradient optimization
Johnson, R. and Zhang, T. (2020) · 2020
Later among the works it cites.
Bandit algorithms
Lattimore, T. and Szepesvári, C. (2020) · 2020
Later among the works it cites.
On the global convergence rates of softmax policy gradient methods
Mei, J., Xiao, C., Szepesvari, C., and Schuurmans, D. (2020) · 2020
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Shani, L., Efroni, Y., and Mannor, S. (2020) · 2020
Later among the works it cites.
Old dog learns new tricks: Randomized ucb for bandit problems
Vaswani, S., Mehrabian, A., Durand, A., and Kveton, B. (2020) · 2020
Later among the works it cites.
What matters for on-policy deep actor-critic methods? a largescale study
Andrychowicz, M., Raichuk, A., Stanczyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., et al. (2021) · 2021
Closest in time.
Optimization issues in kl-constrained approximate policy iteration
Lazić, N., Hao, B., Abbasi-Yadkori, Y., Schuurmans, D., and Szepesvári, C. (2021) · 2021
Closest in time.