Fetching the paper…
Reading the bibliography…
Distributional approaches to value-based reinforcement learning model the entire distribution of returns, rather than just their expected values, and have recently been shown to yield state-of-the-art empirical performance.
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
Probability and Measure
P. Billingsley · 1986
Earlier work this paper cites.
On the Convergence of Stochastic Iterative Dynamic Programming Algorithms
T. Jaakkola, M. I. Jordan, and S. P. Singh · 1994
Earlier work this paper cites.
On-line Q-learning using Connectionist Systems
G. A. Rummery and M. Niranjan · 1994
Earlier work this paper cites.
Stochastic Orders and their Applications
M. Shaked and J. G. Shanthikumar · 1994
Earlier work this paper cites.
Asynchronous Stochastic Approximation and Q-Learning
J. N. Tsitsiklis · 1994
Earlier work this paper cites.
Neuro-Dynamic Programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Cited alongside, same era.
An Analysis of Temporal-Difference Learning with Function Approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Cited alongside, same era.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 1998
Cited alongside, same era.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 1999
Cited alongside, same era.
E-statistics: The Energy of Statistical Samples
G. J. Székely · 2002
Cited alongside, same era.
A Distributional Perspective on Reinforcement Learning
M. G. Bellemare, W. Dabney, and R. Munos
Cited in the paper.
Stochastic Approximation and Recursive Algorithms and Applications
H. Kushner and G. Yin · 2003
Later among the works it cites.
Dynamic Programming and Optimal Control, Vol. II: Approximate Dynamic Programming
D. P. Bertsekas · 2012
Later among the works it cites.
Actor-Critic Algorithms for Risk-Sensitive MDPs
L. A. Prashanth and M. Ghavamzadeh · 2013
Later among the works it cites.
Learning the Variance of the Reward-to-go
A. Tamar, D. Di Castro, and S. Mannor · 2016
Later among the works it cites.
Distributional Reinforcement Learning with Quantile Regression
W. Dabney, M. Rowland, M. G. Bellemare, and R. Munos · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The Cramer Distance as a Solution to Biased Wasserstein Gradients
M. G. Bellemare, I. Danihelka, W. Dabney, S. Mohamed, B. Lakshminarayanan, S. Hoyer, and R. Munos
Cited in the paper.
Nonparametric Return Distribution Approximation for Reinforcement Learning
T. Morimura, M. Sugiyama, H. Kashima, H. Hachiya, and T. Tanaka
Cited in the paper.
Parametric Return Density Estimation for Reinforcement Learning
T. Morimura, M. Sugiyama, H. Kashima, H. Hachiya, and T. Tanaka
Cited in the paper.