Fetching the paper…
Reading the bibliography…
Traditional off-policy actor-critic Reinforcement Learning (RL) algorithms learn value functions of a single target policy.
The monte carlo method
N. Metropolis and S. Ulam · 1949
Earlier work this paper cites.
On the experimental attainment of optimum conditions
G. E. P. Box and K. B. Wilson · 1951
Earlier work this paper cites.
Conditional Markov processes
RL Stratonovich · 1960
Earlier work this paper cites.
Temporal Credit Assignment in Reinforcement Learning
Richard S Sutton · 1984
Earlier work this paper cites.
Advances in importance sampling
Timothy Classen Hesterberg · 1988
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Networks adjusting networks
Jürgen Schmidhuber · 1990
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Paul J Werbos · 1990
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Memory-based stochastic optimization
Andrew W. Moore and Jeff G. Schneider · 1996
Earlier work this paper cites.
Optimization Using Surrogate Objectives on a Helicopter Test Example , pp. 49–58
Andrew J. Booker, J. E. Dennis, Paul D. Frank, David B. Serafini, and Virginia Torczon · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Actor-critic algorithms
Vijay Konda and John Tsitsiklis · 2001
Earlier work this paper cites.
Off-policy temporal difference learning with function approximation
Doina Precup, Richard S. Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
Natural actor-critic
Jan Peters and Stefan Schaal · 2007
Earlier work this paper cites.
Policy gradients with parameter-based exploration for control
Frank Sehnke, Christian Osendorfer, Thomas Rückstieß, Alex Graves, Jan Peters, and Jürgen Schmidhuber · 2008
Cited alongside, same era.
Stochastic approximation: a dynamical systems viewpoint , volume 48
Vivek S Borkar · 2009
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
Hamid R. Maei, Csaba Szepesvári, Shalabh Bhatnagar, Doina Precup, David Silver, and Richard S. Sutton · 2009
Cited alongside, same era.
Parameter-exploring policy gradients
Frank Sehnke, Christian Osendorfer, Thomas Rückstieß, Alex Graves, Jan Peters, and Jürgen Schmidhuber · 2009
Cited alongside, same era.
Learning bounds for importance weighting
Corinna Cortes, Yishay Mansour, and Mehryar Mohri · 2010
Cited alongside, same era.
Toward off-policy learning control with function approximation
Hamid Reza Maei, Csaba Szepesvári, Shalabh Bhatnagar, and Richard S. Sutton · 2010
Universal value function approximators
Tom Schaul, Dan Horgan, Karol Gregor, and David Silver · 2015
Later among the works it cites.
Scalable bayesian optimization using deep neural networks
Jasper Snoek, Oren Rippel, Kevin Swersky, Ryan Kiros, Nadathur Satish, Narayanan Sundaram, Md. Mostofa Ali Patwary, Prabhat Prabhat, and Ryan P. Adams · 2015
Later among the works it cites.
Simulation and the Monte Carlo Method
Reuven Y. Rubinstein and Dirk P. Kroese · 2016
Later among the works it cites.
An emphatic approach to the problem of off-policy temporal-difference learning
Richard S Sutton, A Rupam Mahmood, and Martha White · 2016
Later among the works it cites.
Sample efficient actor-critic with experience replay
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gradient temporal-difference learning algorithms
Hamid Reza Maei · 2011
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S. Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M. Pilarski, Adam White, and Doina Precup · 2011
Cited alongside, same era.
Off-policy actor-critic
Thomas Degris, Martha White, and Richard S. Sutton · 2012
Cited alongside, same era.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Max Jaderberg, Wojciech Marian Czarnecki, Simon Osindero, Oriol Vinyals, Alex Graves, David Silver, and Koray Kavukcuoglu · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Later among the works it cites.
Spinning Up in Deep Reinforcement Learning
Joshua Achiam · 2018
Later among the works it cites.
An off-policy policy gradient theorem using emphatic weightings
Ehsan Imani, Eric Graves, and Martha White · 2018
Later among the works it cites.
Simple random search of static linear policies is competitive for reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Later among the works it cites.
Policy optimization via importance sampling
Alberto Maria Metelli, Matteo Papini, Francesco Faccio, and Marcello Restelli · 2018
Later among the works it cites.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Later among the works it cites.
Algaedice: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 2019
Later among the works it cites.
Jean Harb, Tom Schaul, Doina Precup, and Pierre-Luc Bacon · 2020
Closest in time.
Predicting neural network accuracy from weights
Thomas Unterthiner, Daniel Keysers, Sylvain Gelly, Olivier Bousquet, and Ilya Tolstikhin · 2020
Closest in time.