Fetching the paper…
Reading the bibliography…
Despite many algorithmic advances, our theoretical understanding of practical distributional reinforcement learning methods remains limited.
Learning to predict by the methods of temporal differences
Sutton, Richard S · 1988
Earlier work this paper cites.
Learning from delayed rewards
Watkins, Christopher J. C. H · 1989
Earlier work this paper cites.
Markov Decision Processes: Discrete stochastic dynamic programming
Puterman, Martin L · 1994
Earlier work this paper cites.
Exponentially many local minima for single neurons
Auer, Peter, Herbster, Mark, and Warmuth, Manfred K · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, Dimitri P. and Tsitsiklis, John N · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, John N. and Van Roy, Benjamin · 1997
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, Richard S. and Barto, Andrew G · 1998
Cited alongside, same era.
E-statistics: The energy of statistical samples
Székely, Gabor J · 2002
Cited alongside, same era.
The Arcade Learning Environment: An evaluation platform for general agents
Bellemare, Marc G, Naddaf, Yavar, Veness, Joel, and Bowling, Michael · 2013
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, Marc G., Dabney, Will, and Munos, Rémi
Cited in the paper.
The Cramér distance as a solution to biased Wasserstein gradients
Bellemare, Marc G, Danihelka, Ivo, Dabney, Will, Mohamed, Shakir, Lakshminarayanan, Balaji, Hoyer, Stephan, and Munos, Rémi
Cited in the paper.
Implicit quantile networks for distributional reinforcement learning
Dabney, Will, Ostrovski, Georg, Silver, David, and Munos, Remi
Cited in the paper.
Distributional reinforcement learning with quantile regression
Dabney, Will, Rowland, Mark, Bellemare, Marc G., and Munos, Rémi
Cited in the paper.
Distributional policy gradients
Barth-Maron, Gabriel, Hoffman, Matthew W., Budden, David, Dabney, Will, Horgan, Dan, TB, Dhruva, Muldal, Alistair, Heess, Nicolas, and Lillicrap, Timothy · 2018
Later among the works it cites.
Dopamine: A research framework for deep reinforcement learning
Castro, Pablo S., Moitra, Subhodeep, Gelada, Carles, Kumar, Saurabh, and Bellemare, Marc G · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, Matteo, Modayil, Joseph, van Hasselt, Hado, Schaul, Tom, Ostrovski, Georg, Dabney, Will, Horgan, Dan, Piot, Bilal, Azar, Mohammad, and Silver, David · 2018
Later among the works it cites.
An analysis of categorical distributional reinforcement learning
Rowland, Mark, Bellemare, Marc G, Dabney, Will, Munos, Rémi, and Teh, Yee Whye · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…