Fetching the paper…
Reading the bibliography…
Black-box optimizers that explore in parameter space have often been shown to outperform more sophisticated action space exploration methods developed specifically for the reinforcement learning problem.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Autonomous helicopter control using reinforcement learning policy search methods
J Andrew Bagnell and Jeff G Schneider · 2001
Earlier work this paper cites.
A natural policy gradient
Sham Kakade · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
The cross entropy method for fast policy search
Shie Mannor, Reuven Y Rubinstein, and Yohai Gat · 2003
Earlier work this paper cites.
Policy search by dynamic programming
J Andrew Bagnell, Sham M Kakade, Jeff G Schneider, and Andrew Y Ng · 2004
Earlier work this paper cites.
Online convex optimization in the bandit setting: gradient descent without a gradient
Abraham D Flaxman, Adam Tauman Kalai, and H Brendan McMahan · 2005
Earlier work this paper cites.
Learning tetris using the noisy cross-entropy method
István Szita and András Lörincz · 2006
Earlier work this paper cites.
Evolution strategies for direct policy search
Verena Heidrich-Meisner and Christian Igel · 2008
Cited alongside, same era.
Reinforcement learning of motor skills with policy gradients
Jan Peters and Stefan Schaal · 2008
Cited alongside, same era.
Optimal algorithms for online convex optimization with multi-point bandit feedback
Alekh Agarwal, Ofer Dekel, and Lin Xiao · 2010
Cited alongside, same era.
Parameter-exploring policy gradients
Frank Sehnke, Christian Osendorfer, Thomas Rückstieß, Alex Graves, Jan Peters, and Jürgen Schmidhuber · 2010
Cited alongside, same era.
Using response surfaces and expected improvement to optimize snake robot gait parameters
Matthew Tesch, Jeff Schneider, and Howie Choset · 2011
Cited alongside, same era.
Analysis and improvement of policy gradient estimation
Tingting Zhao, Hirotaka Hachiya, Gang Niu, and Masashi Sugiyama · 2011
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Later among the works it cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Later among the works it cites.
Optimal rates for zero-order convex optimization: The power of two function evaluations
J. C. Duchi, M. I. Jordan, M. J. Wainwright, and A. Wibisono · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I Jordan, and Philipp Moritz · 2015
Later among the works it cites.
Random gradient-free minimization of convex functions
Yurii Nesterov and Vladimir Spokoiny · 2017
Later among the works it cites.
Towards generalization and simplicity in continuous control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Cited alongside, same era.
On the complexity of bandit and derivative-free stochastic convex optimization
Ohad Shamir · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba
Cited in the paper.
Openai gym, 2016b
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba
Cited in the paper.
Aravind Rajeswaran, Kendall Lowrey, Emanuel V Todorov, and Sham M Kakade · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Later among the works it cites.
An optimal algorithm for bandit and zero-order convex optimization with two-point feedback
Ohad Shamir · 2017
Later among the works it cites.
Simple random search provides a competitive approach to reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Later among the works it cites.
Stephen Tu and Benjamin Recht · 2018
Later among the works it cites.