Fetching the paper…
Reading the bibliography…
In contrast to classical reinforcement learning (RL), distributional RL algorithms aim to learn the distribution of returns rather than their expected value.
A Markovian decision process
R. Bellman · 1957
Earlier work this paper cites.
Mean, variance, and probabilistic criteria in finite Markov decision processes: A review
D. J. White · 1988
Earlier work this paper cites.
Aleatory and epistemic uncertainty in probability elicitation with an example from hazardous waste management
S. C. Hora · 1996
Earlier work this paper cites.
Ensemble methods in machine learning
T. G. Dietterich · 2000
Earlier work this paper cites.
Quantile regression
R. Koenker and K. F. Hallock · 2001
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
P. Auer · 2002
Earlier work this paper cites.
Aleatory or epistemic? Does it matter?
A. Der Kiureghian and O. Ditlevsen · 2009
Earlier work this paper cites.
Nonparametric return distribution approximation for reinforcement learning
T. Morimura, M. Sugiyama, H. Kashima, H. Hachiya, and T. Tanaka · 2010
Earlier work this paper cites.
Fundamentals of Applied Probability and Random Processes
O. Ibe · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Earlier work this paper cites.
ViZDoom: A Doom-based AI research platform for visual reinforcement learning
M. Kempka, M. Wydmuch, G. Runc, J. Toczek, and W. Jaśkowski · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped DQN
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Earlier work this paper cites.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Earlier work this paper cites.
UCB exploration via Q-ensembles
R. Y. Chen, S. Sidor, P. Abbeel, and J. Schulman · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
B. Lakshminarayanan, A. Pritzel, and C. Blundell · 2017
Cited alongside, same era.
Efficient exploration with double uncertain value networks
T. M. Moerland, J. Broekens, and C. M. Jonker · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Impala: Scalable distributed deep-RL with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
Cited alongside, same era.
Statistics and samples in distributional reinforcement learning
M. Rowland, R. Dadashi, S. Kumar, R. Munos, M. G. Bellemare, and W. Dabney · 2019
Later among the works it cites.
Fully parameterized quantile function for distributional reinforcement learning
D. Yang, L. Zhao, Z. Lin, T. Qin, J. Bian, and T.-Y. Liu · 2019
Later among the works it cites.
Temporal difference uncertainties as a signal for exploration
S. Flennerhag, J. X. Wang, P. Sprechmann, F. Visin, A. Galashov, S. Kapturowski, D. L. Borsa, N. Heess, A. Barreto, and R. Pascanu · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
C. Jin, Z. Yang, Z. Wang, and M. I. Jordan · 2020
Later among the works it cites.
Distributional reinforcement learning with ensembles
B. Lindenberg, J. Nordqvist, and K.-O. Lindahl · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Is Q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Cited alongside, same era.
Wasserstein and total variation distance between marginals of Lévy processes
E. Mariucci and M. Reiß · 2018
Cited alongside, same era.
The uncertainty Bellman equation and exploration
B. O’Donoghue, I. Osband, R. Munos, and V. Mnih · 2018
Cited alongside, same era.
An analysis of categorical distributional reinforcement learning
M. Rowland, M. Bellemare, W. Dabney, R. Munos, and Y. W. Teh · 2018
Cited alongside, same era.
Optuna: A next-generation hyperparameter optimization framework
T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama · 2019
Cited alongside, same era.
Estimating risk and uncertainty in deep reinforcement learning
W. R. Clements, B. Van Delft, B.-M. Robaglia, R. B. Slaoui, and S. Toth · 2019
Cited alongside, same era.
Successor uncertainties: Exploration and uncertainty in temporal difference learning
D. Janz, J. Hron, P. Mazur, K. Hofmann, J. M. Hernández-Lobato, and S. Tschiatschek · 2019
Cited alongside, same era.
Behaviour suite for reinforcement learning
I. Osband, Y. Doron, M. Hessel, J. Aslanides, E. Sezener, A. Saraiva, K. McKinney, T. Lattimore, C. Szepesvári, S. Singh, B. V. Roy, R. S. Sutton, D. Silver, and H. van Hasselt · 2020
Later among the works it cites.
Bayesian Bellman operators
M. Fellows, K. Hartikainen, and S. Whiteson · 2021
Later among the works it cites.
Distributional reinforcement learning via moment matching
T. Nguyen-Tang, S. Gupta, and S. Venkatesh · 2021
Later among the works it cites.
Fast and data-efficient training of rainbow: An experimental study on Atari
D. Schmidt and T. Schmied · 2021
Later among the works it cites.
DelftBlue Supercomputer (Phase 1), 2022
Delft High Performance Computing Centre (DHPC) · 2022
Later among the works it cites.
Sentinel: Taming uncertainty with ensemble based distributional reinforcement learning
H. Eriksson, D. Basu, M. Alibeigi, and C. Dimitrakakis · 2022
Later among the works it cites.
Distributional reinforcement learning
M. G. Bellemare, W. Dabney, and M. Rowland · 2023
Closest in time.
Ensemble quantile networks: Uncertainty-aware reinforcement learning with applications in autonomous driving
C.-J. Hoel, K. Wolff, and L. Laine · 2023
Closest in time.
Delft Artificial Intelligence Cluster (DAIC), 2024
2024
Closest in time.
On the importance of exploration for generalization in reinforcement learning
Y. Jiang, J. Z. Kolter, and R. Raileanu · 2024
Closest in time.