Fetching the paper…
Reading the bibliography…
In model-free deep reinforcement learning (RL) algorithms, using noisy value estimates to supervise policy evaluation and optimization is detrimental to the sample efficiency.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
A Simplex Method for Function Minimization
J. A. Nelder and R. Mead · 1965
Earlier work this paper cites.
Estimating the mean and variance of the target probability distribution
D.A. Nix and A.S. Weigend · 1994
Earlier work this paper cites.
Bayesian Q-learning
Richard Dearden, Nir Friedman, and Stuart J. Russel · 1998
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Malcolm Strens · 2001
Earlier work this paper cites.
Machine learning: a probabilistic perspective
Kevin P. Murphy · 2012
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Handling stochastic reward delays in machine reinforcement learning
Jeffrey S. Campbell, Sidney N. Givigi, and Howard M. Schwartz · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, and et al · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G. Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
UCB exploration via Q-ensembles
Richard Y. Chen, Szymon Sidor, Pieter Abbeel, and John Schulman · 2017
Earlier work this paper cites.
What uncertainties do we need in bayesian deep learning for computer vision?
Alex Kendall and Yarin Gal · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Cited alongside, same era.
Accurate uncertainties for deep learning using calibrated regression
Volodymyr Kuleshov, Nathan Fenner, and Stefano Ermon · 2018
Cited alongside, same era.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D Sculley, Sebastian Nowozin, Joshua Dillon, Balaji Lakshminarayanan, and Jasper Snoek · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba · 2019
Later among the works it cites.
Estimating risk and uncertainty in deep reinforcement learning
William R. Clements, Bastien Van Delft, Benoît-Marie Robaglia, Reda Bahi Slaoui, and Sébastien Toth · 2020
Later among the works it cites.
Temporal difference uncertainties as a signal for exploration
Sebastian Flennerhag, Jane X. Wang, Pablo Sprechmann, Francesco Visin, Alexandre Galashov, Steven Kapturowski, Diana L. Borsa, Nicolas Heess, Andre Barreto, and Razvan Pascanu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Cited alongside, same era.
Reward estimation for variance reduction in deep reinforcement learning
Joshua Romoff, Peter Henderson, Alexandre Piché, Vincent François-Lavet, and Joelle Pineau · 2018
Cited alongside, same era.
Reinforcement learning: an introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
The limits and potentials of deep learning for robotics
Niko Sünderhauf, Oliver Brock, Walter Scheirer, Raia Hadsell, Dieter Fox, Jürgen Leitner, Ben Upcroft, Pieter Abbeel, Wolfram Burgard, Michael Milford, and et al · 2018
Cited alongside, same era.
Montreal declaration for a responsible development of artificial intelligence, 2018
Université de Montréal · 2018
Cited alongside, same era.
Challenges of real-world reinforcement learning
Gabriel Dulac-Arnold, Daniel Mankowitz, and Todd Hester · 2019
Cited alongside, same era.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, and et al · 2019
Cited alongside, same era.
DisCor: Corrective feedback in reinforcement learning via distribution correction
Aviral Kumar, Abhishek Gupta, and Sergey Levine · 2020
Later among the works it cites.
Evaluating and calibrating uncertainty prediction in regression tasks
Dan Levi, Liran Gispan, Niv Giladi, and Ethan Fetaya · 2020
Later among the works it cites.
Behaviour suite for reinforcement learning
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvari, Satinder Singh, and et al · 2020
Later among the works it cites.
State-aware variational Thompson sampling for deep Q-networks
Siddharth Aravindan and Wee Sun Lee · 2021
Later among the works it cites.
f f -Cal: Calibrated aleatoric uncertainty estimation from neural networks for robot perception
Dhaivat Bhatt, Kaustubh Mani, Dishank Bansal, Krishna Murthy, Hanju Lee, and Liam Paull · 2021
Later among the works it cites.
Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods
Eyke Hüllermeier and Willem Waegeman · 2021
Later among the works it cites.
DEUP: Direct epistemic uncertainty prediction
Moksh Jain, Salem Lahlou, Hadi Nekoei, Victor Butoi, Paul Bertin, Jarrid Rector-Brooks, Maksym Korablyov, and Yoshua Bengio · 2021
Later among the works it cites.
Drivergym: Democratising reinforcement learning for autonomous driving
Parth Kothari, Christian Perone, Luca Bergamini, Alexandre Alahi, and Peter Ondruska · 2021
Later among the works it cites.
SUNRISE: A simple unified framework for ensemble learning in deep reinforcement learning
Kimin Lee, Michael Laskin, Aravind Srinivas, and Pieter Abbeel · 2021
Later among the works it cites.
Deep reinforcement learning for the control of robotic manipulation: A focussed mini-review
Rongrong Liu, Florent Nageotte, Philippe Zanne, Michel de Mathelin, and Birgitta Dresp-Langley · 2021
Later among the works it cites.
Batch inverse-variance weighting: Deep heteroscedastic regression
Vincent Mai, Waleed Khamies, and Liam Paull · 2021
Later among the works it cites.
Uncertainty weighted actor-critic for offline reinforcement learning
Yue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua M Susskind, Jian Zhang, Ruslan Salakhutdinov, and Hanlin Goh · 2021
Later among the works it cites.