Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (RL) methods generally engage in exploratory behavior through noise injection in the action space.
On the theory of the brownian motion
George E Uhlenbeck and Leonard S Ornstein · 1930
Earlier work this paper cites.
Evolutionsstrategie: Optimierung technischer Systeme nach Prinzipien der biologishen Evolution
Ingo Rechenberg and Manfred Eigen · 1973
Earlier work this paper cites.
Numerische Optimierung von Computermodellen mittels der Evolutionsstrategie , volume 1
Hans-Paul Schwefel · 1977
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
A natural policy gradient
Sham Kakade · 2001
Earlier work this paper cites.
R-MAX - A general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael J. Kearns and Satinder P. Singh · 2002
Earlier work this paper cites.
The Levenberg-Marquardt algorithm
Ananth Ranganathan · 2004
Earlier work this paper cites.
Natural actor-critic
Jan Peters and Stefan Schaal · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2008
Earlier work this paper cites.
Policy search for motor primitives in robotics
Jens Kober and Jan Peters · 2008
Earlier work this paper cites.
State-dependent exploration for policy gradient methods
Thomas Rückstieß, Martin Felder, and Jürgen Schmidhuber · 2008
Earlier work this paper cites.
State-dependent exploration for policy gradient methods
Thomas Rückstieß, Martin Felder, and Jürgen Schmidhuber · 2008
Earlier work this paper cites.
Parameter-exploring policy gradients
Frank Sehnke, Christian Osendorfer, Thomas Rückstieß, Alex Graves, Jan Peters, and Jürgen Schmidhuber · 2009
Earlier work this paper cites.
Stochastic search using the natural gradient
Yi Sun, Daan Wierstra, Tom Schaul, and Jürgen Schmidhuber · 2009
Earlier work this paper cites.
Efficient natural evolution strategies
Yi Sun, Daan Wierstra, Tom Schaul, and Jürgen Schmidhuber · 2009
Cited alongside, same era.
A natural evolution strategy for multi-objective optimization
Tobias Glasmachers, Tom Schaul, and Jürgen Schmidhuber · 2010
Cited alongside, same era.
Exponential natural evolution strategies
Tobias Glasmachers, Tom Schaul, Yi Sun, Daan Wierstra, and Jürgen Schmidhuber · 2010
Cited alongside, same era.
Double Q-learning
Hado V Hasselt · 2010
Cited alongside, same era.
High dimensions and heavy tails for natural evolution strategies
Tom Schaul, Tobias Glasmachers, and Jürgen Schmidhuber · 2011
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas · 2015
Later among the works it cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 2015
Later among the works it cites.
Lei Jimmy Ba, Ryan Kiros, and Geoffrey E. Hinton · 2016
Later among the works it cites.
Unifying count-based exploration and intrinsic motivation
Marc G Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Later among the works it cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Natural evolution strategies
Daan Wierstra, Tom Schaul, Tobias Glasmachers, Yi Sun, Jan Peters, and Jürgen Schmidhuber · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Cited alongside, same era.
Later among the works it cites.
Benchmarking deep reinforcement learning for continous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Later among the works it cites.
VIME: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Later among the works it cites.
#Exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Later among the works it cites.
Surprise-based intrinsic motivation for deep reinforcement learning
Joshua Achiam and Shankar Sastry · 2017
Closest in time.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2017
Closest in time.
Count-based exploration with neural density models
Georg Ostrovski, Marc G. Bellemare, Aäron van den Oord, and Rémi Munos · 2017
Closest in time.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Closest in time.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, and Ilya Sutskever · 2017
Closest in time.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sainbayar Sukhbaatar, Ilya Kostrikov, Arthur Szlam, and Rob Fergus · 2017
Closest in time.