Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Natural exponential families with quadratic variance functions
Carl N Morris · 1982
Earlier work this paper cites.
Temporal Credit Assignment in Reinforcement Learning
Richard S. Sutton · 1984
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1985
Earlier work this paper cites.
Probability: Theory and Examples (Wadsworth and Brooks/Cole Statistics/Probability Series)
Richard Durrett · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Discrete Event Systems: Sensitivity Analysis and Stochastic Optimization by the Score Function Method , volume 1
Reuven Y Rubinstein and Alexander Shapiro · 1993
Earlier work this paper cites.
On the convergence of stochastic iterative dynamic programming algorithms
Tommi Jaakkola, Michael I Jordan, and Satinder P Singh · 1994
Earlier work this paper cites.
Stochastic approximation with two time scales
Vivek S Borkar · 1997
Earlier work this paper cites.
The cross-entropy method for combinatorial and continuous optimization
Reuven Rubinstein · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David A. McAllester, Satinder P. Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (cma-es)
Nikolaus Hansen, Sibylle D. Müller, and Petros Koumoutsakos · 2003
Earlier work this paper cites.
Reinforcement learning as classification: Leveraging modern classifiers
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
The cross entropy method for fast policy search
Shie Mannor, Reuven Rubinstein, and Yohai Gat · 2003
Earlier work this paper cites.
The cross-entropy method: a unified approach to combinatorial optimization, Monte-Carlo simulation, and machine learning , volume 133
Reuven Y Rubinstein and Dirk P Kroese · 2004
Earlier work this paper cites.
Learning tetris using the noisy cross-entropy method
István Szita and András Lörincz · 2006
Earlier work this paper cites.
A study on the cross-entropy method for rare-event probability estimation
Tito Homem-de Mello · 2007
Earlier work this paper cites.
A model reference adaptive search method for global optimization
Jiaqiao Hu, Michael C Fu, and Steven I Marcus · 2007
Earlier work this paper cites.
Reinforcement learning by reward-weighted regression for operational space control
Jan Peters and Stefan Schaal · 2007
Earlier work this paper cites.
Stochastic Approximation: A Dynamical Systems Viewpoint
Vivek S. Borkar · 2008
Earlier work this paper cites.
Policy search for motor primitives in robotics
Jens Kober and Jan R Peters · 2008
Earlier work this paper cites.