Fetching the paper…
Reading the bibliography…
A common belief in model-free reinforcement learning is that methods based on random search in the parameter space of policies exhibit significantly worse sample complexity than those that explore the space of actions.
Random optimization
J. Matyas · 1965
Earlier work this paper cites.
A natural policy gradient
S. M. Kakade · 2002
Earlier work this paper cites.
Online convex optimization in the bandit setting: gradient descent without a gradient
A. D. Flaxman, A. T. Kalai, and H. B. McMahan · 2005
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Optimal algorithms for online convex optimization with multi-point bandit feedback
A. Agarwal, O. Dekel, and L. Xiao · 2010
Earlier work this paper cites.
T. Degris, M. White, and R. S. Sutton · 2012
Earlier work this paper cites.
Query complexity of derivative-free optimization
K. G. Jamieson, R. Nowak, and B. Recht · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Guided policy search
S. Levine and V. Koltun · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Highly-smooth zero-th order online optimization
F. Bach and V. Perchet · 2016
Cited alongside, same era.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
S. Gu, T. Lillicrap, Z. Ghahramani, R. E. Turner, and S. Levine · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Reproducibility of benchmarked deep reinforcement learning tasks for continuous control
R. Islam, P. Henderson, M. Gomrokchi, and D. Precup · 2017
Later among the works it cites.
Ray: A distributed framework for emerging ai applications
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, W. Paul, M. I. Jordan, and I. Stoica · 2017
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine · 2017
Later among the works it cites.
Random gradient-free minimization of convex functions
Y. Nesterov and V. Spokoiny · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
Sample efficient actor-critic with experience replay
Z. Wang, V. Bapst, N. Heess, V. Mnih, R. Munos, K. Kavukcuoglu, and N. de Freitas · 2016
Cited alongside, same era.
On the sample complexity of the linear quadratic regulator
S. Dean, H. Mania, N. Matni, B. Recht, and S. Tu · 2017
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
T. Haarnoja, H. Tang, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Emergence of locomotion behaviours in rich environments
N. Heess, S. Sriram, J. Lemmon, J. Merel, G. Wayne, Y. Tassa, T. Erez, Z. Wang, A. Eslami, M. Riedmiller, et al · 2017
Cited alongside, same era.
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2017
Cited alongside, same era.
M. Plappert, R. Houthooft, P. Dhariwal, S. Sidor, R. Y. Chen, X. Chen, T. Asfour, P. Abbeel, and M. Andrychowicz · 2017
Later among the works it cites.
Towards generalization and simplicity in continuous control
A. Rajeswaran, K. Lowrey, E. Todorov, and S. Kakade · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
T. Salimans, J. Ho, X. Chen, and I. Sutskever · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Least-squares temporal difference learning for the linear quadratic regulator
S. Tu and B. Recht · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Y. Wu, E. Mansimov, S. Liao, R. Grosse, and J. Ba · 2017
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Closest in time.