Fetching the paper…
Reading the bibliography…
In reinforcement learning, the advantage function is critical for policy improvement, but is often extracted from a learned Q-function.
Learning from delayed rewards
Christopher J. C. H. Watkins · 1989
Earlier work this paper cites.
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Advantage updating
Leemon C. Baird III · 1993
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Tommi Jaakkola, Michael Jordan, and Satinder Singh · 1993
Earlier work this paper cites.
On the convergence of stochastic iterative dynamic programming algorithms
Tommi Jaakkola, Michael I. Jordan, and Satinder P. Singh · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and Q-learning
John N. Tsitsiklis · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon C. Baird · 1995
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S. Sutton and Andrew G. Barto · 1998
Cited alongside, same era.
Lecture 6.5 - RMSProp
Geoffrey Hinton and Tijman Tieleman · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado Van Hasselt, Marc Lanctot, and Nando De Freitas · 2015
Cited alongside, same era.
Increasing the action gap: New operators for reinforcement learning
Deep reinforcement learning with double Q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Later among the works it cites.
Sample efficient actor-critic with experience replay
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu, and Nando de Freitas · 2016
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Later among the works it cites.
Gap-increasing policy evaluation for efficient and noise-tolerant reinforcement learning
Tadashi Kozuno, Dongqi Han, and Kenji Doya · 2019
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville, and Marc G. Bellemare · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marc G Bellemare, Georg Ostrovski, Arthur Guez, Philip Thomas, and Rémi Munos · 2016
Cited alongside, same era.
DQN Zoo: Reference implementations of DQN-based agents
John Quan and Georg Ostrovski
Cited in the paper.
Hsiao-Ru Pan, Nico Gürtler, Alexander Neitz, and Bernhard Schölkopf · 2021
Later among the works it cites.