Understanding deep learning requires rethinking generalization
Original
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Later among the works it cites.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Later among the works it cites.
A Distributional Perspective on Reinforcement Learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Later among the works it cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Later among the works it cites.
Concrete dropout
Yarin Gal, Jiri Hron, and Alex Kendall · 2017
Later among the works it cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Later among the works it cites.
Ensemble sampling
Xiuyuan Lu and Benjamin Van Roy · 2017
Later among the works it cites.
The uncertainty bellman equation and exploration
Original
Brendan O’Donoghue, Ian Osband, Remi Munos, and Volodymyr Mnih · 2017
Later among the works it cites.
Deep exploration via randomized value functions
Original
Ian Osband, Daniel Russo, Zheng Wen, and Benjamin Van Roy · 2017
Later among the works it cites.
Why is posterior sampling better than optimism for reinforcement learning?
Ian Osband and Benjamin Van Roy · 2017
Later among the works it cites.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aaron van den Oord, and Rémi Munos · 2017
Later among the works it cites.
Parameter space noise for exploration
Original
Matthias Plappert, Rein Houthooft, Prafulla Dhariwal, Szymon Sidor, Richard Y Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz · 2017
Later among the works it cites.
A tutorial on Thompson sampling
Original
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, and Ian Osband · 2017
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard Sutton and Andrew Barto · 2017
Later among the works it cites.
Variational deep q network
Original
Yunhao Tang and Alp Kucukelbir · 2017
Later among the works it cites.
Efficient exploration through bayesian deep q-networks
Original
Kamyar Azizzadenesheli, Emma Brunskill, and Animashree Anandkumar · 2018
Closest in time.
Distributed distributional deterministic policy gradients
Original
Gabriel Barth-Maron, Matthew W Hoffman, David Budden, Will Dabney, Dan Horgan, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap · 2018
Closest in time.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc G Bellemare, and Rémi Munos · 2018
Closest in time.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2018
Closest in time.
Distributed prioritized experience replay
Daniel Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado Van Hasselt, and David Silver · 2018
Closest in time.
Gradient descent quantizes ReLU network features
Original
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Closest in time.
Simple random search provides a competitive approach to reinforcement learning
Original
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Closest in time.
Deepmind control suite
Original
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Closest in time.
Randomized value functions via multiplicative normalizing flows
Original
Ahmed Touati, Harsh Satija, Joshua Romoff, Joelle Pineau, and Pascal Vincent · 2018
Closest in time.