Fetching the paper…
Reading the bibliography…
In this work, we build on recent advances in distributional reinforcement learning to give a generally applicable, flexible, and state-of-the-art distributional variant of DQN.
Theory of Games and Economic Behavior
von Neumann, J. and Morgenstern, O · 1947
Earlier work this paper cites.
Dynamic Programming
Bellman, R. E · 1957
Earlier work this paper cites.
Robust estimation of a location parameter
Huber, P. J · 1964
Earlier work this paper cites.
Risk-sensitive markov decision processes
Howard, R. A. and Matheson, J. E · 1972
Earlier work this paper cites.
Markov decision processes with a new optimality criterion: discrete time
Jaquette, S. C · 1973
Earlier work this paper cites.
The variance of discounted markov decision processes
Sobel, M. J · 1982
Earlier work this paper cites.
The dual theory of choice under risk
Yaari, M. E · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Mean, variance, and probabilistic criteria in finite markov decision processes: a review
White, D. J · 1988
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Allais paradox
Allais, M · 1990
Earlier work this paper cites.
Advances in prospect theory: cumulative representation of uncertainty
Tversky, A. and Kahneman, D · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
Premium calculation by transforming the layer premium density
Wang, S · 1996
Earlier work this paper cites.
Curvature of the probability weighting function
Wu, G. and Gonzalez, R · 1996
Earlier work this paper cites.
Risk sensitive markov decision processes
Marcus, S. I., Fernández-Gaucherand, E., Hernández-Hernandez, D., Coraluppi, S., and Fard, P · 1997
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
Müller, A · 1997
Cited alongside, same era.
On the shape of the probability weighting function
Gonzalez, R. and Wu, G · 1999
Cited alongside, same era.
A class of distortion operators for pricing financial and insurance risks
Wang, S. S · 2000
Cited alongside, same era.
Quantile Regression
Koenker, R · 2005
Cited alongside, same era.
Kalman temporal differences
Geist, M. and Pietquin, O · 2010
Cited alongside, same era.
On the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, H. J · 2012
Cited alongside, same era.
Remarks on quantiles and distortion risk measures
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N · 2016
Later among the works it cites.
More than a million ways to be pushed. a high-fidelity experimental dataset of planar pushing
Yu, K.-T., Bauza, M., Fazeli, N., and Rodriguez, A · 2016
Later among the works it cites.
Wasserstein Generative Adversarial Networks
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Later among the works it cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Later among the works it cites.
From optimal transport to generative modeling: the vegan cookbook
Bousquet, O., Gelly, S., Tolstikhin, I., Simon-Gabriel, C.-J., and Schoelkopf, B · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dhaene, J., Kukush, A., Linders, D., and Tang, Q · 2012
Cited alongside, same era.
PAC bounds for discounted MDPs
Lattimore, T. and Hutter, M · 2012
Cited alongside, same era.
The Arcade Learning Environment: an evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B · 2013
Cited alongside, same era.
Algorithms for CVaR optimization in MDPs
Chow, Y. and Ghavamzadeh, M · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., et al · 2017
Later among the works it cites.
Maddison, C. J., Lawson, D., Tucker, G., Heess, N., Doucet, A., Mnih, A., and Teh, Y. W · 2017
Later among the works it cites.
How should a robot assess risk? Towards an axiomatic theory of risk in robotics
Majumdar, A. and Pavone, M · 2017
Later among the works it cites.
Efficient exploration with double uncertain value networks
Moerland, T. M., Broekens, J., and Jonker, C. M · 2017
Later among the works it cites.
Tolstikhin, I., Bousquet, O., Gelly, S., and Schoelkopf, B · 2017
Later among the works it cites.
Distributional policy gradients
Barth-Maron, G., Hoffman, M. W., Budden, D., Dabney, W., Horgan, D., TB, D., Muldal, A., Heess, N., and Lillicrap, T · 2018
Closest in time.
Distributional reinforcement learning with quantile regression
Dabney, W., Rowland, M., Bellemare, M. G., and Munos, R · 2018
Closest in time.
The Reactor: a fast and sample-efficient actor-critic agent for reinforcement learning
Gruslys, A., Dabney, W., Azar, M. G., Piot, B., Bellemare, M. G., and Munos, R · 2018
Closest in time.
Rainbow: combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2018
Closest in time.
An analysis of categorical distributional reinforcement learning
Rowland, M., Bellemare, M. G., Dabney, W., Munos, R., and Teh, Y. W · 2018
Closest in time.