Fetching the paper…
Reading the bibliography…
Many reinforcement learning algorithms can be seen as versions of approximate policy iteration (API).
V-MPO: on-policy maximum a posteriori policy optimization for discrete and continuous control
Song, H. F., Abdolmaleki, A., Springenberg, J. T., Clark, A., Soyer, H., Rae, J. W., Noury, S., Ahuja, A., Liu, S., Tirumala, D., Heess, N., Belov, D., Riedmiller, M. A., and Botvinick, M. M · 1909
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Leverage the average: an analysis of regularization in rl
Vieillard, N., Kozuno, T., Scherrer, B., Pietquin, O., Munos, R., and Geist, M · 2003
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Bartlett, P. L. and Tewari, A · 2009
Earlier work this paper cites.
Approximate policy iteration: A survey and some new methods
Bertsekas, D. P · 2011
Earlier work this paper cites.
Value function approximation in reinforcement learning using the fourier basis
Konidaris, G., Osentoski, S., and Thomas, P · 2011
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Reward augmented maximum likelihood for neural structured prediction
Norouzi, M., Bengio, S., Jaitly, N., Schuster, M., Wu, Y., Schuurmans, D., et al · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Neu, G., Jonsson, A., and Gómez, V · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Understanding the impact of entropy on policy optimization
Ahmed, Z., Le Roux, N., Norouzi, M., and Schuurmans, D · 2019
Later among the works it cites.
Surrogate objectives for batch policy optimization in one-step decision making
Chen, M., Gummadi, R., Harris, C., and Schuurmans, D · 2019
Later among the works it cites.
Optimality and approximation with policy gradient methods in markov decision processes
Agarwal, A., Kakade, S., Lee, J., and Mahajan, G · 2020
Later among the works it cites.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., and Wang, Z · 2020
Later among the works it cites.
Provably efficient adaptive approximate policy iteration
Hao, B., Lazic, N., Abbasi-Yadkori, Y., Joulani, P., and Szepesvari, C · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M · 2018
Cited alongside, same era.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al · 2018
Cited alongside, same era.
Escaping the graviational pull of softmax
Mei, J., Xiao, C., Dai, B., Li, L., Szepesvári, C., and Schuurmans, D
Cited in the paper.
On the global convergence rates of softmax policy gradient methods
Mei, J., Xiao, C., Szepesvári, C., and Schuurmans, D
Cited in the paper.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Shani, L., Efroni, Y., and Mannor, S
Cited in the paper.
Optimistic policy optimization with bandit feedback
Shani, L., Efroni, Y., Rosenberg, A., and Mannor, S
Cited in the paper.
Momentum in reinforcement learning
Vieillard, N., Scherrer, B., Pietquin, O., and Geist, M
Cited in the paper.
Tomar, M., Shani, L., Efroni, Y., and Ghavamzadeh, M · 2020
Later among the works it cites.