Constrained Markov decision processes , volume 7
Altman, E · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Beck, A. and Teboulle, M · 2003
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R · 2003
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
Weissman, T., Ordentlich, E., Seroussi, G., Verdu, S., and Weinberger, M. J · 2003
Earlier work this paper cites.
Policy gradient methods for robotics
Peters, J. and Schaal, S · 2006
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Peters, J. and Schaal, S · 2008
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
Maurer, A. and Pontil, M · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
The online loop-free stochastic shortest-path problem
Neu, G., György, A., and Szepesvári, C · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Deisenroth, M. and Rasmussen, C. E · 2011
Earlier work this paper cites.