Fetching the paper…
Reading the bibliography…
Learning a predictive model of the mean return, or value function, plays a critical role in many reinforcement learning algorithms.
Estimating risk and uncertainty in deep reinforcement learning, 2019
William R. Clements, Bastien Van Delft, Benoît-Marie Robaglia, Reda Bahi Slaoui, and Sébastien Toth · 1905
Earlier work this paper cites.
Robust Estimation of a Location Parameter
Peter J. Huber · 1964
Earlier work this paper cites.
Ohio supercomputer center, 1987
Ohio Supercomputer Center · 1987
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Estimating the mean and variance of the target probability distribution
D.A. Nix and A.S. Weigend · 1994
Earlier work this paper cites.
Aleatory and epistemic uncertainty in probability elicitation with an example from hazardous waste management
Stephen C. Hora · 1996
Earlier work this paper cites.
Sample-based distributional policy gradient
Rahul Singh, Keuntaek Lee, and Yongxin Chen · 2001
Earlier work this paper cites.
On the Markov chain central limit theorem
Galin L. Jones · 2004
Earlier work this paper cites.
A theoretical analysis of model-based interval estimation
Alexander L. Strehl and Michael L. Littman · 2005
Earlier work this paper cites.
SUNRISE: A simple unified framework for ensemble learning in deep reinforcement learning
Kimin Lee, Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2007
Earlier work this paper cites.
Implicit distributional reinforcement learning
Yuguang Yue, Zhendong Wang, and Mingyuan Zhou · 2007
Earlier work this paper cites.
Aleatory or epistemic? does it matter?
Armen Der Kiureghian and Ove Ditlevsen · 2008
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Sparse quantile huber regression for efficient and robust estimation
Aleksandr Y. Aravkin, Anju Kambadur, Aurélie C. Lozano, and Ronny Luss · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
What uncertainties do we need in bayesian deep learning for computer vision?
Alex Kendall and Yarin Gal · 2017
Later among the works it cites.
Efficient exploration with double uncertain value networks
Thomas M. Moerland, Joost Broekens, and Catholijn M. Jonker · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Spinning Up in Deep Reinforcement Learning, 2018
Joshua Achiam · 2018
Later among the works it cites.
Distributed distributional deterministic policy gradients
Gabriel Barth-Maron, Matthew W. Hoffman, David Budden, Will Dabney, Dan Horgan, Dhruva TB, Alistair Muldal, Nicolas Heess, and Timothy P. Lillicrap · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Philip Thomas · 2014
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Shixiang Gu, Timothy Lillicrap, Zoubin Ghahramani, Richard E Turner, and Sergey Levine · 2016
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles, 2016
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc G. Bellemare, and Rémi Munos · 2017
Cited alongside, same era.
Later among the works it cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos J. Storkey, and Oleg Klimov · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Sample efficient deep reinforcement learning via uncertainty estimation
Vincent Mai, Kaustubh Mani, and Liam Paull · 2022
Closest in time.
The phenomenon of policy churn, 2022
Tom Schaul, André Barreto, John Quan, and Georg Ostrovski · 2022
Closest in time.
Small batch deep reinforcement learning, 2023
Johan Obando-Ceron, Marc G. Bellemare, and Pablo Samuel Castro · 2023
Closest in time.