Fetching the paper…
Reading the bibliography…
Distributional reinforcement learning (distributional RL) has seen empirical success in complex Markov Decision Processes (MDPs) in the setting of nonlinear function approximation.
The theory of dynamic programming
Richard Bellman · 1954
Earlier work this paper cites.
A class of wasserstein metrics for probability distributions
Clark R Givens, Rae Michael Shortt, et al · 1984
Earlier work this paper cites.
The wasserstein distance and approximation theorems
Ludger Rüschendorf · 1985
Earlier work this paper cites.
Matrix multiplication via arithmetic progressions
Don Coppersmith and Shmuel Winograd · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Gavin A Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Restricted boltzmann machines for collaborative filtering
Ruslan Salakhutdinov, Andriy Mnih, and Geoffrey Hinton · 2007
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller · 2013
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Deep exploration via bootstrapped DQN
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Later among the works it cites.
Connecting generative adversarial networks and actor-critic methods
David Pfau and Oriol Vinyals · 2016
Later among the works it cites.
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Later among the works it cites.
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Later among the works it cites.
Improved training of wasserstein gans
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Bayesian policy gradient and actor-critic algorithms
Mohammad Ghavamzadeh, Yaakov Engel, and Michal Valko · 2016
Cited alongside, same era.
Vincent Huang, Tobias Ley, Martha Vlachou-Konchylaki, and Wenfeng Hu · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
On the convergence properties of gan training
Lars Mescheder · 2018
Closest in time.