Fetching the paper…
Reading the bibliography…
Policy Mirror Descent (PMD) is a popular framework in reinforcement learning, serving as a unifying perspective that encompasses numerous algorithms.
On Information and Sufficiency
Kullback, S. and Leibler, R. A · 1951
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
Bregman, L. M · 1967
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
Nemirovski, A. and Yudin, D. B · 1983
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J · 1991
Earlier work this paper cites.
Parallel Optimization: Theory, Algorithms, and Applications
Censor, Y. and Zenios, S. A · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 1999
Earlier work this paper cites.
Evolutionary algorithms in noisy environments: theoretical issues and guidelines for practice
Beyer, H.-G · 2000
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. and Tsitsiklis, J · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Baxter, J. and Bartlett, P. L · 2001
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
Hansen, N. and Ostermeier, A · 2001
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2002
Earlier work this paper cites.
Natural actor-critic
Peters, J. and Schaal, S · 2008
Earlier work this paper cites.
A simple modification in cma-es achieving linear time and space complexity
Ros, R. and Hansen, N · 2008
Earlier work this paper cites.
Natural actor-critic algorithms
Bhatnagar, S., Sutton, R. S., Ghavamzadeh, M., and Lee, M · 2009
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Bubeck, S · 2015
Earlier work this paper cites.
Efficient bregman projections onto the simplex
Krichene, W., Krichene, S., and Bayen, A · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2016
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S., Shammah, S., and Shashua, A · 2016
Cited alongside, same era.
Evolution strategies as a scalable alternative to reinforcement learning
Salimans, T., Ho, J., Chen, X., Sidor, S., and Sutskever, I · 2017
Cited alongside, same era.
Actor-critic is implicitly biased towards high entropy optimal policies
Hu, Y., Ji, Z., and Telgarsky, M · 2022
Later among the works it cites.
Mirror learning: A unifying framework of policy optimisation
Kuba, J. G., De Witt, C. A. S., and Foerster, J · 2022
Later among the works it cites.
Discovered policy optimisation
Lu, C., Kuba, J., Letcher, A., Metz, L., Schroeder de Witt, C., and Foerster, J · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Mirror descent policy optimization
Tomar, M., Shani, L., Efroni, Y., and Ghavamzadeh, M · 2022
Later among the works it cites.
A general class of surrogate functions for stable and efficient reinforcement learning
Vaswani, S., Bachem, O., Totaro, S., Müller, R., Garg, S., Geist, M., Machado, M. C., Castro, P. S., and Le Roux, N · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Deep neuroevolution: Genetic algorithms are a competitive alternative for training deep neural networks for reinforcement learning, 2018
Such, F. P., Madhavan, V., Conti, E., Lehman, J., Stanley, K. O., and Clune, J · 2018
Cited alongside, same era.
Optuna: A next-generation hyperparameter optimization framework, 2019
Akiba, T., Sano, S., Yanase, T., Ohta, T., and Koyama, M · 2019
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Cited alongside, same era.
Behaviour suite for reinforcement learning
Osband, I., Doron, Y., Hessel, M., Aslanides, J., Sezener, E., Saraiva, A., McKinney, K., Lattimore, T., Szepesvari, C., Singh, S., et al · 2019
Cited alongside, same era.
Regularized evolution for image classifier architecture search
Real, E., Aggarwal, A., Huang, Y., and Le, Q. V · 2019
Cited alongside, same era.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Misra, D., Henaff, M., Krishnamurthy, A., and Langford, J · 2020
Cited alongside, same era.
Later among the works it cites.
On the convergence rates of policy gradient methods
Xiao, L · 2022
Later among the works it cites.
A novel framework for policy mirror descent with general parametrization and linear convergence
Alfano, C., Yuan, R., and Rebeschini, P · 2023
Later among the works it cites.
Policy mirror descent inherently explores action space
Li, Y. and Lan, G · 2023
Later among the works it cites.
Optimistic natural policy gradient: a simple efficient policy optimization framework for online rl
Liu, Q., Weisz, G., György, A., Jin, C., and Szepesvari, C · 2023
Later among the works it cites.
Adversarial cheap talk
Lu, C., Willi, T., Letcher, A., and Foerster, J. N · 2023
Later among the works it cites.
Decision-aware actor-critic with function approximation and theoretical guarantees
Vaswani, S., Kazemi, A., Babanezhad, R., and Roux, N. L · 2023
Later among the works it cites.
Linear convergence of natural policy gradient methods with log-linear policies
Yuan, R., Du, S. S., Gower, R. M., Lazaric, A., and Xiao, L · 2023
Later among the works it cites.
Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks
Chevalier-Boisvert, M., Dai, B., Towers, M., Perez-Vicente, R., Willems, L., Lahlou, S., Pal, S., Castro, P. S., and Terry, J · 2024
Closest in time.
Discovering temporally-aware reinforcement learning algorithms
Jackson, M. T., Lu, C., Kirsch, L., Lange, R. T., Whiteson, S., and Foerster, J. N · 2024
Closest in time.