Fetching the paper…
Reading the bibliography…
We introduce NoisyNet, a deep reinforcement learning agent with parametric noise added to its weights, and show that the induced stochasticity of the agent's policy can be used to aid efficient exploration.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
Dynamic programming and modern control theory
Richard Bellman and Robert Kalaba · 1965
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E Hinton and Drew Van Camp · 1993
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin Puterman · 1994
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Dimitri Bertsekas · 1995
Earlier work this paper cites.
Training with noise is equivalent to Tikhonov regularization
Chris M Bishop · 1995
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Evolutionary algorithms for reinforcement learning
David E Moriarty, Alan C Schultz, and John J Grefenstette · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David A. McAllester, Satinder P. Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Satinder P Singh, Andrew G Barto, and Nuttapong Chentanez · 2004
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Peter Auer and Ronald Ortner · 2007
Earlier work this paper cites.
What is intrinsic motivation? A typology of computational approaches
Pierre-Yves Oudeyer and Frederic Kaplan · 2007
Earlier work this paper cites.
Managing uncertainty within value function approximation in reinforcement learning
Matthieu Geist and Olivier Pietquin · 2010
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber · 2010
Cited alongside, same era.
Practical variational inference for neural networks
Alex Graves · 2011
Cited alongside, same era.
Monte-Carlo swarm policy search
Jeremy Fix and Matthieu Geist · 2012
Cited alongside, same era.
The sample-complexity of general reinforcement learning
Tor Lattimore, Marcus Hutter, and Peter Sunehag · 2013
Cited alongside, same era.
Generalization and exploration via randomized value functions
Ian Osband, Benjamin Van Roy, and Zheng Wen · 2014
Cited alongside, same era.
Efficient exploration for dialogue policy learning with BBQ networks & replay buffer spiking
Zachary C Lipton, Jianfeng Gao, Lihong Li, Xiujun Li, Faisal Ahmed, and Li Deng · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Training recurrent neural networks by diffusion
Hossein Mobahi · 2016
Later among the works it cites.
Deep exploration via bootstrapped DQN
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The arcade learning environment: An evaluation platform for general agents
Marc Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2015
Cited alongside, same era.
Weight uncertainty in neural networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas · 2016
Later among the works it cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 2016
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Closest in time.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Closest in time.
Bayesian recurrent neural networks
Meire Fortunato, Charles Blundell, and Oriol Vinyals · 2017
Closest in time.
Deep exploration via randomized value functions
Ian Osband, Daniel Russo, Zheng Wen, and Benjamin Van Roy · 2017
Closest in time.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aaron van den Oord, and Remi Munos · 2017
Closest in time.
Parameter space noise for exploration
Matthias Plappert, Rein Houthooft, Prafulla Dhariwal, Szymon Sidor, Richard Y Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz · 2017
Closest in time.
Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Tim Salimans, J. Ho, X. Chen, and I. Sutskever · 2017
Closest in time.