Fetching the paper…
Reading the bibliography…
We describe an iterative procedure for optimizing policies, with guaranteed monotonic improvement.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A., Sutton, R., and Anderson, C · 1983
Earlier work this paper cites.
Adapting arbitrary normal mutation distributions in evolution strategies: The covariance matrix adaptation
Hansen, Nikolaus and Ostermeier, Andreas · 1996
Earlier work this paper cites.
Numerical optimization , volume 2
Wright, Stephen J and Nocedal, Jorge · 1999
Earlier work this paper cites.
PEGASUS: A policy search method for large mdps and pomdps
Ng, A. Y. and Jordan, M · 2000
Earlier work this paper cites.
Asymptopia: an exposition of statistical asymptotic theory
Pollard, David · 2000
Earlier work this paper cites.
A natural policy gradient
Kakade, Sham · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, Sham and Langford, John · 2002
Earlier work this paper cites.
Covariant policy search
Bagnell, J. A. and Schneider, J · 2003
Earlier work this paper cites.
Reinforcement learning as classification: Leveraging modern classifiers
Lagoudakis, Michail G and Parr, Ronald · 2003
Earlier work this paper cites.
A tutorial on MM algorithms
Hunter, David R and Lange, Kenneth · 2004
Earlier work this paper cites.
Stochastic policy gradient reinforcement learning on a simple 3d biped
Tedrake, R., Zhang, T., and Seung, H · 2004
Cited alongside, same era.
Dynamic programming and optimal control , volume 1
Bertsekas, D · 2005
Cited alongside, same era.
Efficient methods in convex programming
Nemirovski, Arkadi · 2005
Cited alongside, same era.
Fast biped walking with a reflexive controller and realtime policy searching
Geng, T., Porr, B., and Wörgötter, F · 2006
Cited alongside, same era.
Learning tetris using the noisy cross-entropy method
Szita, István and Lörincz, András · 2006
Cited alongside, same era.
Markov chains and mixing times
Levin, D. A., Peres, Y., and Wilmer, E. L · 2009
Cited alongside, same era.
MuJoCo: A physics engine for model-based control
Todorov, Emanuel, Erez, Tom, and Tassa, Yuval · 2012
Later among the works it cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Later among the works it cites.
A survey on policy search for robotics
Deisenroth, M., Neumann, G., and Peters, J · 2013
Later among the works it cites.
Approximate dynamic programming finally performs well in the game of Tetris
Gabillon, Victor, Ghavamzadeh, Mohammad, and Scherrer, Bruno · 2013
Later among the works it cites.
Playing Atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Later among the works it cites.
Monte Carlo theory, methods and examples
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimal gait and form for animal locomotion
Wampler, Kevin and Popović, Zoran · 2009
Cited alongside, same era.
Relative entropy policy search
Peters, J., Mülling, K., and Altün, Y · 2010
Cited alongside, same era.
Infinite-horizon policy-gradient estimation
Bartlett, P. L. and Baxter, J · 2011
Cited alongside, same era.
Training deep and recurrent networks with hessian-free optimization
Martens, J. and Sutskever, I · 2012
Cited alongside, same era.
Reinforcement learning of motor skills with policy gradients
Peters, J. and Schaal, S
Cited in the paper.
Natural actor-critic
Peters, Jan and Schaal, Stefan
Cited in the paper.
Owen, Art B · 2013
Later among the works it cites.
Revisiting natural gradient for deep networks
Pascanu, Razvan and Bengio, Yoshua · 2013
Later among the works it cites.
Safe policy iteration
Pirotta, Matteo, Restelli, Marcello, Pecorino, Alessio, and Calandriello, Daniele · 2013
Later among the works it cites.
Deep learning for real-time atari game play using offline Monte-Carlo tree search planning
Guo, X., Singh, S., Lee, H., Lewis, R. L., and Wang, X · 2014
Later among the works it cites.
Learning neural network policies with guided policy search under unknown dynamics
Levine, Sergey and Abbeel, Pieter · 2014
Later among the works it cites.