Fetching the paper…
Reading the bibliography…
Reinforcement learning agents are faced with two types of uncertainty.
Risk-sensitive markov decision processes
Howard, R. A. and Matheson, J. E · 1972
Earlier work this paper cites.
A possibility for implementing curiosity and boredom in model-building neural controllers
Schmidhuber, J · 1991
Earlier work this paper cites.
Bayesian learning for neural networks , volume 118
Neal, R. M · 1995
Earlier work this paper cites.
Quantile regression
Koenker, R. and Hallock, K. F · 2001
Earlier work this paper cites.
Information theory, inference and learning algorithms
MacKay, D. J · 2003
Earlier work this paper cites.
Planning to be surprised: Optimal bayesian exploration in dynamic environments
Sun, Y., Gomez, F., and Schmidhuber, J · 2011
Earlier work this paper cites.
The arcade learning environment: an evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Learning to optimize via information-directed sampling
Russo, D. and Van Roy, B · 2014
Earlier work this paper cites.
Weight uncertainty in neural networks
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Stadie, B. C., Levine, S., and Abbeel, P · 2015
Cited alongside, same era.
Dropout as a Bayesian approximation: representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Cited alongside, same era.
Risk versus uncertainty in deep learning: Bayes, bootstrap and the dangers of dropout
Osband, I · 2016
Cited alongside, same era.
Deep exploration via bootstrapped DQN
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B · 2016
Cited alongside, same era.
Learning the variance of the reward-to-go
Tamar, A., Di Castro, D., and Mannor, S · 2016
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K., Calandra, R., McAllister, R., and Levine, S · 2018
Later among the works it cites.
Decomposition of uncertainty in bayesian deep learning for efficient and risk-sensitive learning
Depeweg, S., Hernandez-Lobato, J.-M., Doshi-Velez, F., and Udluft, S · 2018
Later among the works it cites.
Exploration by distributional reinforcement learning
Tang, Y. and Agrawal, S · 2018
Later among the works it cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2019
Closest in time.
Model-predictive policy learning with uncertainty regularization for driving in dense traffic
Henaff, M., Canziani, A., and LeCun, Y · 2019
Closest in time.
Information-directed exploration for deep reinforcement learning
Nikolov, N., Kirschner, J., Berkenkamp, F., and Krause, A · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Cited alongside, same era.
Leike, J., Martic, M., Krakovna, V., Ortega, P. A., Everitt, T., Lefrancq, A., Orseau, L., and Legg, S · 2017
Cited alongside, same era.
Efficient exploration with double uncertain value networks
Moerland, T. M., Broekens, J., and Jonker, C. M · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Cited alongside, same era.
Efficient exploration through Bayesian deep Q-networks
Azizzadenesheli, K., Brunskill, E., and Anandkumar, A · 2018
Cited alongside, same era.
Implicit quantile networks for distributional reinforcement learning
Dabney, W., Ostrovski, G., Silver, D., and Munos, R
Cited in the paper.
Closest in time.
Randomized value functions via multiplicative normalizing flows
Touati, A., Satija, H., Romoff, J., Pineau, J., and Vincent, P · 2019
Closest in time.
Minatar: An atari-inspired testbed for more efficient reinforcement learning experiments
Young, K. and Tian, T · 2019
Closest in time.
Bayesian quantile regression
Yu, K. and Moyeed, R. A · 2019
Closest in time.
Uncertainty in neural networks: Bayesian ensembling
Pearce, T., Zaki, M., Brintrup, A., Anastassacos, N., and Neely, A · 2020
Closest in time.