Fetching the paper…
Reading the bibliography…
Characterizing aleatoric and epistemic uncertainty on the predicted rewards can help in building reliable reinforcement learning (RL) systems.
“On the likelihood that one unknown probability exceeds another in view of the evidence of two samples”
William Thompson · 1933
Earlier work this paper cites.
“Asynchronous methods for deep reinforcement learning”
Volodymyr Mnih, Adria Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver and Koray Kavukcuoglu · 1937
Earlier work this paper cites.
“Dynamic programming”
Richard Bellman · 1966
Earlier work this paper cites.
“Risk-Sensitive Markov Decision Processes”
Ronald. Howard and James. Matheson · 1972
Earlier work this paper cites.
“Neuronlike adaptive elements that can solve difficult learning control problems”
Andrew. Barto, Richard. Sutton and Charles. Anderson · 1983
Earlier work this paper cites.
“Generalization in Reinforcement Learning: Successful Examples Using Sparse Coarse Coding”
Richard Sutton · 1995
Earlier work this paper cites.
“No free lunch theorems for optimization”
D.H. Wolpert and W.G. Macready · 1997
Earlier work this paper cites.
“Reinforcement Learning through Active Inference”, 2020
Alexander Tschantz, Beren Millidge, Anil. Seth and Christopher. Buckley · 2002
Earlier work this paper cites.
“Gaussian Processes in Reinforcement Learning”
Malte Kuss and Carl Rasmussen · 2003
Earlier work this paper cites.
“Reinforcement Learning with Gaussian Processes”
Yaakov Engel, Shie Mannor and Ron Meir · 2005
Earlier work this paper cites.
“Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)”
Carl Rasmussen and Christopher K.. Williams · 2005
Earlier work this paper cites.
“The Challenge of Safe Lunar Landing”
Tye Brady and Stephen Paschall · 2010
Earlier work this paper cites.
“The Challenge of Safe Lunar Landing”
Tye Brady and Stephen Paschall · 2010
Earlier work this paper cites.
“Nonparametric Return Distribution Approximation for Reinforcement Learning”
Tetsuro Morimura, Masashi Sugiyama, Hisashi Kashima, Hirotaka Hachiya and Toshiyuki Tanaka · 2010
Earlier work this paper cites.
“Analysis of thompson sampling for the multi-armed bandit problem”
Shipra Agrawal and Navin Goyal · 2012
Earlier work this paper cites.
“Playing atari with deep reinforcement learning”
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra and Martin Riedmiller · 2013
Earlier work this paper cites.
“Automatic Model Construction with Gaussian Processes”, 2014
David Duvenaud · 2014
Earlier work this paper cites.
“Human-level control through deep reinforcement learning”
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei. Rusu, Joel Veness, Marc. Bellemare, Alex Graves, Martin Riedmiller, Andreas. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg and Demis Hassabis · 2015
Earlier work this paper cites.
“RLPy: A Value-Function-Based Reinforcement Learning Framework for Education and Research”
Alborz Geramifard, Christoph Dann, Robert. Klein, William Dabney and Jonathan. How · 2015
Earlier work this paper cites.
“A Comprehensive Survey on Safe Reinforcement Learning”
Javier Garc\’ia, Fern and o Fern\’andez · 2015
Earlier work this paper cites.
“Bayesian reinforcement learning: A survey”
Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau and Aviv Tamar · 2015
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“Sample efficient actor-critic with experience replay”
Ziyu Wang, Victor Bapst, Nicolas Heess, Volodymyr Mnih, Remi Munos, Koray Kavukcuoglu and Nando de Freitas · 2016
Earlier work this paper cites.
“Uncertainty in Deep Learning”, 2016
Yarin Gal · 2016
Earlier work this paper cites.
“Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning”
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
“An information-theoretic analysis of thompson sampling”
Daniel Russo and Benjamin Van · 2016
Earlier work this paper cites.
“Deep Exploration via Bootstrapped DQN”
Ian Osband, Charles Blundell, Alexander Pritzel and Benjamin Van · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang and Wojciech Zaremba · 2016
Earlier work this paper cites.
“Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles”
Balaji Lakshminarayanan, Alexander Pritzel and Charles Blundell · 2017
Earlier work this paper cites.
“A Distributional Perspective on Reinforcement Learning”
Marc. Bellemare, Will Dabney and R\’emi Munos · 2017
Cited alongside, same era.
“Proximal policy optimization algorithms”
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford and Oleg Klimov · 2017
Cited alongside, same era.
“Uncertainty-Aware Reinforcement Learning for Collision Avoidance”
Gregory Kahn, Adam Villaflor, Vitchyr Pong, Pieter Abbeel and Sergey Levine · 2017
Cited alongside, same era.
“Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic”
Shixiang Gu, Timothy Lillicrap, Zoubin Ghahramani, Richard. Turner and Sergey Levine · 2017
Cited alongside, same era.
“Predictive Uncertainty Estimation via Prior Networks”
Andrey Malinin and Mark Gales · 2018
Cited alongside, same era.
“Implicit Quantile Networks for Distributional Reinforcement Learning”
“Uncertainty Estimation Using a Single Deep Deterministic Neural Network”
Joost van Amersfoort, Lewis Smith, Yee Teh and Yarin Gal · 2020
Later among the works it cites.
“Posterior Network: Uncertainty Estimation without OOD Samples via Density-Based Pseudo-Counts”
Bertrand Charpentier, Daniel Zügner and Stephan Günnemann · 2020
Later among the works it cites.
“Deep Evidential Regression”
Alexander Amini, Wilko Schwarting, Ava Soleimany and Daniela Rus · 2020
Later among the works it cites.
“Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in Regret”
Yingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang and Qiaomin Xie · 2020
Later among the works it cites.
“Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in Regret”
Yingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang and Qiaomin Xie · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Will Dabney, Georg Ostrovski, David Silver and Remi Munos · 2018
Cited alongside, same era.
“Randomized Prior Functions for Deep Reinforcement Learning”
Ian Osband, John Aslanides and Albin Cassirer · 2018
Cited alongside, same era.
“Assessing Generalization in Deep Reinforcement Learning”
Charles Packer, Katelyn Gao, Jernej Kos, Philipp Krähenbühl, Vladlen Koltun and Dawn Song · 2018
Cited alongside, same era.
“Risk-Constrained Reinforcement Learning with Percentile Risk Criteria”
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson and Marco Pavone · 2018
Cited alongside, same era.
“Safe Reinforcement Learning with Model Uncertainty Estimates”
Bj\"orn L\"utjens, Michael Everett and Jonathan. How · 2018
Cited alongside, same era.
“Decomposition of Uncertainty in Bayesian Deep Learning for Efficient and Risk-sensitive Learning”
Stefan Depeweg, Jose-Miguel Hernandez-Lobato, Finale Doshi-Velez and Steffen Udluft · 2018
Cited alongside, same era.
“Bayesian Deep Reinforcement Learning via Deep Kernel Learning”
Junyu Xuan, Jie Lu, Zheng Yan and Guangquan Zhang · 2018
Cited alongside, same era.
Hannes Eriksson and Christos Dimitrakakis · 2020
Later among the works it cites.
“Towards Neural Networks that Provably Know When They Don’t Know”
Alexander Meinke and Matthias Hein · 2020
Later among the works it cites.
“Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU Networks”
Agustinus Kristiadi, Matthias Hein and Philipp Hennig · 2020
Later among the works it cites.
Andrey Malinin, Sergey Chervontsev, Ivan Provilkov and Mark Gales · 2020
Later among the works it cites.
“Behaviour Suite for Reinforcement Learning”
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvari, Satinder Singh, Benjamin Roy, Richard Sutton, David Silver and Hado Hasselt · 2020
Later among the works it cites.
In European Commission , 2020
“The assessment list for trustworthy artificial intelligence (ALTAI) for self assessment” · 2020
Later among the works it cites.
“Experiment Tracking with Weights and Biases” Software available from wandb.com, 2020
Lukas Biewald · 2020
Later among the works it cites.
“Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability”
Dibya Ghosh, Jad Rahme, Aviral Kumar, Amy Zhang, Ryan Adams and Sergey Levine · 2021
Later among the works it cites.
“Generalized out-of-distribution detection: A survey”
Jingkang Yang, Kaiyang Zhou, Yixuan Li and Ziwei Liu · 2021
Later among the works it cites.
“Out-of-Distribution Detection for Automotive Perception”
Julia Nitsch, Masha Itkina, Ransalu Senanayake, Juan Nieto, Max Schmidt, Roland Siegwart, Mykel. Kochenderfer and Cesar Cadena · 2021
Later among the works it cites.
“A review of uncertainty quantification in deep learning: Techniques, applications and challenges”
Moloud Abdar, Farhad Pourpanah, Sadiq Hussain, Dana Rezazadegan, Li Liu, Mohammad Ghavamzadeh, Paul Fieguth, Abbas Khosravi, U Acharya and Vladimir Makarenkov · 2021
Later among the works it cites.
“On Feature Collapse and Deep Kernel Learning for Single Forward Pass Uncertainty”
Joost van Amersfoort, Lewis Smith, Andrew Jesson, Oscar Key and Yarin Gal · 2021
Later among the works it cites.
“Shifts: A Dataset of Real Distributional Shift Across Multiple Large-Scale Tasks”
Andrey Malinin, Neil Band, Yarin Gal, Mark Gales, Alexander Ganshin, German Chesnokov, Alexey Noskov, Andrey Ploskonosov, Liudmila Prokhorenkova, Ivan Provilkov, Vatsal Raina, Vyas Raina, Denis Roginskiy, Mariya Shmatova, Panagiotis Tigas and Boris Yangel · 2021
Later among the works it cites.
“Domain Shifts in Reinforcement Learning: Identifying Disturbances in Environments”
Tom Haider, Felippe Roza, Dirk Eilers, Karsten Roscher and Stephan G\"unnemann · 2021
Later among the works it cites.
“Graph Posterior Network: Bayesian Predictive Uncertainty for Node Classification”
Maximilian Stadler, Bertrand Charpentier, Simon Geisler, Daniel Z\"ugner and Stephan G\"unnemann · 2021
Later among the works it cites.
“The Neural Testbed: Evaluating Predictive Distributions”
Ian Osband, Zheng Wen, Seyed Asghari, Vikranth Dwaracherla, Botao Hao, Morteza Ibrahimi, Dieterich Lawson, Xiuyuan Lu, Brendan O’Donoghue and Benjamin Van · 2021
Later among the works it cites.
“A Survey of Generalisation in Deep Reinforcement Learning”
Robert Kirk, Amy Zhang, Edward Grefenstette and Tim Rocktäschel · 2021
Later among the works it cites.
“Out-of-Distribution Dynamics Detection: RL-Relevant Benchmarks and Results”
Mohamad Danesh and Alan Fern · 2021
Later among the works it cites.
“Benchmark for Out-of-Distribution Detection in Deep Reinforcement Learning”
Aaqib Mohammed and Matias Valdenegro-Toro · 2021
Later among the works it cites.
“Towards Out-Of-Distribution Generalization: A Survey”
Zheyan Shen, Jiashuo Liu, Yue He, Xingxuan Zhang, Renzhe Xu, Han Yu and Peng Cui · 2021
Later among the works it cites.
“Assuring the Machine Learning Lifecycle: Desiderata, Methods, and Challenges”
Rob Ashmore, Radu Calinescu and Colin Paterson · 2021
Later among the works it cites.
“Natural Posterior Network: Deep Bayesian Predictive Uncertainty for Exponential Family Distributions”
Bertrand Charpentier, Oliver Borchert, Daniel Z\"ugner, Simon Geisler and Stephan G\"unnemann · 2022
Closest in time.
“Distributional Reinforcement Learning”
Marc. Bellemare, Will Dabney and Mark Rowland · 2022
Closest in time.
“Leveraging Procedural Generation to Benchmark Reinforcement Learning”
Karl Cobbe, Chris Hesse, Jacob Hilton and John Schulman · 2056
Closest in time.