Fetching the paper…
Reading the bibliography…
We devise a distributional variant of gradient temporal-difference (TD) learning.
Practical issues in temporal difference learning
Gerald Tesauro · 1992
Earlier work this paper cites.
Consideration of risk in reinforcement learning
Matthias Heger · 1994
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
JN Tsitsiklis and B Van Roy · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
The ode method for convergence of stochastic approximation and reinforcement learning
Vivek S Borkar and Sean P Meyn · 2000
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Covariant policy search
J Andrew Bagnell and Jeff Schneider · 2003
Earlier work this paper cites.
E-statistics: The energy of statistical samples
GJ Székely · 2003
Cited alongside, same era.
Reinforcement learning with gaussian processes
Yaakov Engel, Shie Mannor, and Ron Meir · 2005
Cited alongside, same era.
Risk-aware decision making and dynamic programming
Boris Defourny, Damien Ernst, and Louis Wehenkel · 2008
Cited alongside, same era.
Achieving master level play in 9 x 9 computer go
Sylvain Gelly and David Silver · 2008
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
Shalabh Bhatnagar, Doina Precup, David Silver, Richard S Sutton, Hamid R Maei, and Csaba Szepesvári · 2009
Cited alongside, same era.
Toward off-policy learning control with function approximation
Hamid Reza Maei, Csaba Szepesvári, Shalabh Bhatnagar, and Richard S Sutton · 2010
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
PATTERN RECOGNITION AND MACHINE LEARNING
M Bishop Christopher · 2016
Later among the works it cites.
Vizdoom: A doom-based ai research platform for visual reinforcement learning
Michał Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Jaśkowski · 2016
Later among the works it cites.
Learning the variance of the reward-to-go
Aviv Tamar, Dotan Di Castro, and Shie Mannor · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Later among the works it cites.
Wasserstein gan
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Solving variational inequalities with stochastic mirror-prox algorithm
Anatoli Juditsky, Arkadi Nemirovski, and Claire Tauvel · 2011
Cited alongside, same era.
Policy evaluation with temporal differences: A survey and comparison
Christoph Dann, Gerhard Neumann, and Jan Peters · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos
Cited in the paper.
The cramer distance as a solution to biased wasserstein gradients
Marc G Bellemare, Ivo Danihelka, Will Dabney, Shakir Mohamed, Balaji Lakshminarayanan, Stephan Hoyer, and Rémi Munos
Cited in the paper.
A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation
Richard S Sutton, Hamid R Maei, and Csaba Szepesvári
Cited in the paper.
Later among the works it cites.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc G Bellemare, and Rémi Munos · 2017
Later among the works it cites.
The uncertainty bellman equation and exploration
Brendan O’Donoghue, Ian Osband, Remi Munos, and Volodymyr Mnih · 2017
Later among the works it cites.
An analysis of categorical distributional reinforcement learning
Mark Rowland, Marc Bellemare, Will Dabney, Remi Munos, and Yee Whye Teh · 2018
Closest in time.