Fetching the paper…
Reading the bibliography…
The policy gradient theorem describes the gradient of the expected discounted return with respect to an agent's policy parameters.
Asynchronous methods for deep reinforcement learning. In International conference on machine learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016 · 1937
Earlier work this paper cites.
Principles of Mathematical Analysis
Walter Rudin et al · 1964
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Average reward reinforcement learning: Foundations, algorithms, and empirical results
Sridhar Mahadevan. 1996 · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation. In Advances in Neural Information Processing Systems
John N Tsitsiklis and Benjamin Van Roy. 1997 · 1997
Earlier work this paper cites.
Direct gradient-based reinforcement learning
J. Baxter and P. L. Bartlett. 2000 · 2000
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
D. P. Bertsekas and J. N. Tsitsiklis. 2000 · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation. In Advances in neural information processing systems
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour. 2000 · 2000
Earlier work this paper cites.
Optimizing average reward using discounted rewards. In Proceedings of the 14th International Conference on Computational Learning Theory
Sham Kakade. 2001 · 2001
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation. In Proceedings of the 26th Annual International Conference on Machine Learning
Richard S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora. 2009 · 2009
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling. 2013 · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms. In Proceedings of the 31st International Conference on Machine Learning
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014 · 2014
Cited alongside, same era.
Bias in natural actor-critic algorithms. In Proceedings of the 31st International Conference on Machine Learning
Philip Thomas. 2014 · 2014
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. 2015b · 2015
Cited alongside, same era.
Introduction to reinforcement learning with function approximation. In Tutorial at the Conference on Neural Information Processing Systems
Richard S Sutton. 2015 · 2015
Yuhuai Wu, Elman Mansimov, Shun Liao, Roger Grosse, and Jimmy Ba. 2017 · 2017
Later among the works it cites.
Spinning Up in Deep Reinforcement Learning
Joshua Achiam. 2018 · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods. In Proceedings of the 35th International Conference on Machine Learning
Scott Fujimoto, Herke Hoof, and David Meger. 2018 · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Conference on Machine Learning
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Later among the works it cites.
Stable Baselines
Ashley Hill, Antonin Raffin, Maximilian Ernestus, Adam Gleave, Anssi Kanervisto, Rene Traore, Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, and Yuhuai Wu. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Continuous control with deep reinforcement learning. In Proceedings of the 4th International Conference on Learning Representations
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2016 · 2016
Cited alongside, same era.
OpenAI Baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Cited alongside, same era.
Sample efficient actor-critic with experience replay. In Proceedings of the 5th International Conference on Learning Representations
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas. 2017 · 2017
Cited alongside, same era.
Trust region policy optimization. In Proceedings of the 32nd International Conference on Machine Learning
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015a
Cited in the paper.
Communication
Hermann Amandus Schwarz. 1873
Cited in the paper.
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
Later among the works it cites.
Garage: A toolkit for reproducible reinforcement learning research
The garage contributors. 2019 · 2019
Closest in time.
TF-Agents: A library for reinforcement learning in TensorFlow
Sergio Guadarrama, Anoop Korattikara, Oscar Ramirez, Pablo Castro, Ethan Holly, Sam Fishman, Ke Wang, Ekaterina Gonina, Neal Wu, Efi Kokiopoulou, Luciano Sbaiz, Jamie Smith, Gábor Bartók, Jesse Berent, Chris Harris, Vincent Vanhoucke, Eugene Brevdo. 2018 · 2019
Closest in time.
The Autonomous Learning Library
Chris Nota. 2020 · 2020
Closest in time.