Fetching the paper…
Reading the bibliography…
Modern deep reinforcement learning (RL) algorithms are motivated by either the generalised policy iteration (GPI) or trust-region learning (TRL) frameworks.
Sur diverses questions de calcul fonctionnel
Gateaux, R · 1922
Earlier work this paper cites.
A markov decision process. journal of mathematical mechanics
Bellman, R · 1957
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of visual pattern recognition
Fukushima, K. and Miyake, S · 1982
Earlier work this paper cites.
Global and local variational derivatives and integral representations of gâteaux differentials
Hamilton, E. and Nashed, M · 1982
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovskij, A. S. and Yudin, D. B · 1983
Earlier work this paper cites.
Q-learning
Watkins, C. J. C. H. and Dayan, P · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
A generalized theorem of the maximum
Ausubel, L. M. and Deneckere, R. J · 1993
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., Mcallester, D., Singh, S., and Mansour, Y · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J. and Bartlett, P. L · 2001
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Beck, A. and Teboulle, M · 2003
Earlier work this paper cites.
Analysis and improvement of policy gradient estimation
Zhao, T., Hachiya, H., Niu, G., and Sugiyama, M · 2011
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Duan, Y., Chen, X., Houthooft, R., Schulman, J., and Abbeel, P · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Probability functional descent: A unifying perspective on gans, variational inference, and reinforcement learning
Chu, C., Blanchet, J., and Glynn, P · 2019
Later among the works it cites.
Neural proximal/trust region policy optimization attains globally optimal policy
Liu, B., Cai, Q., Yang, Z., and Wang, Z · 2019
Later among the works it cites.
Implementation matters in deep policy gradients: A case study on ppo and trpo
Engstrom, L., Ilyas, A., Santurkar, S., Tsipras, D., Janoos, F., Rudolph, L., and Madry, A · 2020
Later among the works it cites.
Revisiting design choices in proximal policy optimization
Hsu, C. C.-Y., Mendler-Dünner, C., and Hardt, M · 2020
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Shani, L., Efroni, Y., and Mannor, S · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
A unified view of entropy-regularized markov decision processes
Neu, G., Jonsson, A., and Gómez, V · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Rudder: Return decomposition for delayed rewards
Arjona-Medina, J. A., Gillhofer, M., Widrich, M., Unterthiner, T., Brandstetter, J., and Hochreiter, S · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Later among the works it cites.
Mirror descent policy optimization
Tomar, M., Shani, L., Efroni, Y., and Ghavamzadeh, M · 2020
Later among the works it cites.
Truly proximal policy optimization
Wang, Y., He, H., and Tan, X · 2020
Later among the works it cites.
Trust region policy optimisation in multi-agent reinforcement learning
Kuba, J. G., Chen, R., Wen, M., Wen, Y., Sun, F., Wang, J., and Yang, Y · 2021
Later among the works it cites.
Lan, G · 2021
Later among the works it cites.
Generalized proximal policy optimization with sample reuse
Queeney, J., Paschalidis, I., and Cassandras, C · 2021
Later among the works it cites.
A functional mirror ascent view of policy gradient methods with function approximation
Vaswani, S., Bachem, O., Totaro, S., Mueller, R., Geist, M., Machado, M. C., Castro, P. S., and Roux, N. L · 2021
Later among the works it cites.
Zhan, W., Cen, S., Huang, B., Chen, Y., Lee, J. D., and Chi, Y · 2021
Later among the works it cites.