Fetching the paper…
Reading the bibliography…
The distributional perspective on reinforcement learning (RL) has given rise to a series of successful Q-learning algorithms, resulting in state-of-the-art performance in arcade game environments.
Robust estimation of a location parameter
Peter J. Huber · 1964
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Premium calculation by transforming the layer premium density
Shaun Wang · 1996
Earlier work this paper cites.
Curvature of the probability weighting function
George Wu and Richard Gonzalez · 1996
Earlier work this paper cites.
Introduction to Reinforcement Learning
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
On the shape of the probability weighting function
Rodrigo González and George Wu · 1999
Earlier work this paper cites.
Optimization of conditional value-at-risk
R. Tyrrell Rockafellar and Stanislav Uryasev · 2000
Earlier work this paper cites.
A class of distortion operators for pricing financial and insurance risks
Shaun S. Wang · 2000
Earlier work this paper cites.
The cross-entropy method
Reuven Y. Rubinstein and Dirk P. Kroese · 2004
Earlier work this paper cites.
Double q-learning
Hado V. Hasselt · 2010
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Manfred Otto Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
How should a robot assess risk? Towards an axiomatic theory of risk in robotics
Anirudha Majumdar and Marco Pavone · 2017
Later among the works it cites.
Distributed distributional deterministic policy gradients
Gabriel Barth-Maron, Matthew W Hoffman, David Budden, Will Dabney, Dan Horgan, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap · 2018
Later among the works it cites.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Later among the works it cites.
Scalable deep reinforcement learning for vision-based robotic manipulation
Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, et al · 2018
Later among the works it cites.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
Sergey Levine, Peter Pastor, Alex Krizhevsky, Julian Ibarz, and Deirdre Quillen · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
Lerrel Pinto and Abhinav Gupta · 2016
Cited alongside, same era.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
Dex-net 2.0: Deep learning to plan robust grasps with synthetic point clouds and analytic grasp metrics
Jeffrey Mahler, Jacky Liang, Sherdil Niyaz, Michael Laskey, Richard Doan, Xinyu Liu, Juan Aparicio Ojea, and Ken Goldberg · 2017
Cited alongside, same era.
Implicit quantile networks for distributional reinforcement learning
Will Dabney, Georg Ostrovski, David Silver, and Remi Munos
Cited in the paper.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc G Bellemare, and Rémi Munos
Cited in the paper.
Later among the works it cites.
Deep reinforcement learning for vision-based robotic grasping: A simulated comparative evaluation of off-policy methods
Deirdre Quillen, Eric Jang, Ofir Nachum, Chelsea Finn, Julian Ibarz, and Sergey Levine · 2018
Later among the works it cites.
Striving for simplicity in off-policy deep reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2019
Closest in time.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Closest in time.