Fetching the paper…
Reading the bibliography…
In this paper we argue for the fundamental importance of the value distribution: the distribution of the random return received by a reinforcement learning agent.
Dynamic programming
Bellman, Richard E · 1957
Earlier work this paper cites.
Markov decision processes with a new optimality criterion: Discrete time
Jaquette, Stratton C · 1973
Earlier work this paper cites.
Some asymptotic theory for the bootstrap
Bickel, Peter J. and Freedman, David A · 1981
Earlier work this paper cites.
The variance of discounted markov decision processes
Sobel, Matthew J · 1982
Earlier work this paper cites.
Discounted mdp’s: Distribution functions and exponential utility maximization
Chung, Kun-Jen and Sobel, Matthew J · 1987
Earlier work this paper cites.
Mean, variance, and probabilistic criteria in finite markov decision processes: a review
White, D. J · 1988
Earlier work this paper cites.
A fixed point theorem for distributions
Rösler, Uwe · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete stochastic dynamic programming
Puterman, Martin L · 1994
Earlier work this paper cites.
Probability and measure
Billingsley, Patrick · 1995
Earlier work this paper cites.
Stable function approximation in dynamic programming
Gordon, Geoffrey · 1995
Earlier work this paper cites.
Reinforcement learning with selective perception and hidden state
McCallum, Andrew K · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, Dimitri P. and Tsitsiklis, John N · 1996
Earlier work this paper cites.
Multitask learning
Caruana, Rich · 1997
Earlier work this paper cites.
Bayesian Q-learning
Dearden, Richard, Friedman, Nir, and Russell, Stuart · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, Richard S. and Barto, Andrew G · 1998
Cited alongside, same era.
Approximately optimal approximate reinforcement learning
Kakade, Sham and Langford, John · 2002
Cited alongside, same era.
On the convergence of optimistic policy iteration
Tsitsiklis, John N · 2002
Cited alongside, same era.
Many-layered learning
Utgoff, Paul E. and Stracuzzi, David J · 2002
Cited alongside, same era.
Reinforcement learning with gaussian processes
Engel, Yaakov, Mannor, Shie, and Meir, Ron · 2005
Cited alongside, same era.
Probabilistic inference for solving discrete and continuous state markov decision processes
Toussaint, Marc and Storkey, Amos · 2006
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, Marc G, Naddaf, Yavar, Veness, Joel, and Bowling, Michael · 2013
Later among the works it cites.
Actor-critic algorithms for risk-sensitive mdps
Prashanth, LA and Ghavamzadeh, Mohammad · 2013
Later among the works it cites.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Later among the works it cites.
Massively parallel methods for deep reinforcement learning
Nair, Arun, Srinivasan, Praveen, Blackwell, Sam, Alcicek, Cagdas, Fearon, Rory, De Maria, Alessandro, Panneershelvam, Vedavyas, Suleyman, Mustafa, Beattie, Charles, and Petersen, Stig et al · 2015
Later among the works it cites.
Compress and control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, Tao, Lizotte, Daniel, Bowling, Michael, and Schuurmans, Dale · 2008
Cited alongside, same era.
An expectation maximization algorithm for continuous markov decision processes with arbitrary reward
Hoffman, Matthew D., de Freitas, Nando, Doucet, Arnaud, and Peters, Jan · 2009
Cited alongside, same era.
Kalman temporal differences
Geist, Matthieu and Pietquin, Olivier · 2010
Cited alongside, same era.
Mean-variance optimization in markov decision processes
Mannor, Shie and Tsitsiklis, John N · 2011
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R.S., Modayil, J., Delp, M., Degris, T., Pilarski, P.M., White, A., and Precup, D · 2011
Cited alongside, same era.
On the sample complexity of reinforcement learning with a generative model
Azar, Mohammad Gheshlaghi, Munos, Rémi, and Kappen, Hilbert · 2012
Cited alongside, same era.
Veness, Joel, Bellemare, Marc G., Hutter, Marcus, Chua, Alvin, and Desjardins, Guillaume · 2015
Later among the works it cites.
Prioritized experience replay
Schaul, Tom, Quan, John, Antonoglou, Ioannis, and Silver, David · 2016
Later among the works it cites.
Learning the variance of the reward-to-go
Tamar, Aviv, Di Castro, Dotan, and Mannor, Shie · 2016
Later among the works it cites.
Pixel recurrent neural networks
Van den Oord, Aaron, Kalchbrenner, Nal, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Deep reinforcement learning with double Q-learning
van Hasselt, Hado, Guez, Arthur, and Silver, David · 2016
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Wang, Ziyu, Schaul, Tom, Hessel, Matteo, Hasselt, Hado van, Lanctot, Marc, and de Freitas, Nando · 2016
Later among the works it cites.
The cramer distance as a solution to biased wasserstein gradients
Bellemare, Marc G., Danihelka, Ivo, Dabney, Will, Mohamed, Shakir, Lakshminarayanan, Balaji, Hoyer, Stephan, and Munos, Rémi · 2017
Closest in time.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, Max, Mnih, Volodymyr, Czarnecki, Wojciech Marian, Schaul, Tom, Leibo, Joel Z, Silver, David, and Kavukcuoglu, Koray · 2017
Closest in time.