Fetching the paper…
Reading the bibliography…
We cast policy gradient methods as the repeated application of two operators: a policy improvement operator $\mathcal{I}$, which maps any policy $\pi$ to a better one $\mathcal{I}\pi$, and a projection operator $\mathcal{P}$, which finds the best approximation of $\mathcal{I}\pi$ in the set of realizable policies.
Q-learning
Christopher Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
A view of the em algorithm that justifies incremental, sparse, and other variants
Radford M Neal and Geoffrey E Hinton · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder P. Singh · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Divergence measures and message passing
Tom Minka · 2005
Earlier work this paper cites.
Neural fitted Q-iteration – first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Cited alongside, same era.
Policy search for motor primitives in robotics
Jens Kober and Jan R Peters · 2009
Cited alongside, same era.
Efficient reductions for imitation learning
Stéphane Ross and Drew Bagnell · 2010
Cited alongside, same era.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, and Jan Peters · 2013
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Efficient iterative policy optimization
Nicolas Le Roux · 2016
Cited alongside, same era.
A unified view of entropy-regularized Markov decision processes
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Rémi Munos, Nicolas Heess, and Martin A. Riedmiller · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Later among the works it cites.
Imitation learning as f f -divergence minimization
Liyiming Ke, Matt Barnes, Wen Sun, Gilwoo Lee, Sanjiban Choudhury, and Siddhartha Srinivasa · 2019
Later among the works it cites.
Ray interference: a source of plateaus in deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Bridging the gap between value and policy based reinforcement learning
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans · 2017
Cited alongside, same era.
Equivalence between policy gradients and soft Q-learning
John Schulman, Pieter Abbeel, and Xi Chen
Cited in the paper.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov
Cited in the paper.
Tom Schaul, Diana Borsa, Joseph Modayil, and Razvan Pascanu · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Optimality and approximation with policy gradient methods in Markov decision processes
Alekh Agarwal, Sham M. Kakade, Jason D. Lee, and Gaurav Mahajan · 2020
Closest in time.
Taylor expansion policy optimization
Yunhao Tang, Michal Valko, and Rémi Munos · 2020
Closest in time.