Least-squares policy iteration
Lagoudakis, M. and R. Parr (2004) · 2004
Cited alongside, same era.
Tree-based batch mode reinforcement learning
Ernst, D., P. Geurts, and L. Wehenkel (2005) · 2005
Cited alongside, same era.
Finite time bounds for sampling based fitted value iteration
Szepesvari, C. and R. Munos (2005) · 2005
Cited alongside, same era.
Towards a unified theory of state abstraction for MDPs
Li, L., T. J. Walsh, and M. L. Littman (2006) · 2006
Cited alongside, same era.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., C. Szepesvári, and R. Munos (2008) · 2008
Cited alongside, same era.
Projected equation methods for approximate solution of large linear systems
Bertsekas, D. P. and H. Yu (2009) · 2009
Cited alongside, same era.
Markov Chains and Stochastic Stability
Meyn, S. and R. L. Tweedie (2009) · 2009
Cited alongside, same era.
Rademacher complexity bounds for non-i.i.d. processes
Mohri, M. and A. Rostamizadeh (2009) · 2009
Cited alongside, same era.
Hilbert space embeddings and metrics on probability measures
Sriperumbudur, B. K., A. Gretton, K. Fukumizu, B. Schölkopf, and G. R. Lanckriet (2010) · 2010
Cited alongside, same era.
Dynamic programming and optimal control
Bertsekas, D. P. (2012) · 2012
Cited alongside, same era.
A kernel two-sample test
Gretton, A., K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola (2012) · 2012
Cited alongside, same era.
Foundations of machine learning
Mohri, M., A. Rostamizadeh, and A. Talwalkar (2012) · 2012
Cited alongside, same era.