Rates of convergence of empirical processes for mixing sequences
B. Yu · 1994
Earlier work this paper cites.
Residual algorithms: reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Sphere packing numbers for subsets of the Boolean n n -cube with bounded Vapnik-Chervonenkis dimension
D. Haussler · 1995
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
A. Müller · 1997
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N. Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R.S. Sutton and A.G. Barto · 1998
Earlier work this paper cites.
The pagerank citation ranking: Bringing order to the web
Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, R. S. Sutton, and S. Singh · 2000
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Susan A Murphy, Mark J van der Laan, James M Robins, and Conduct Problems Prevention Research Group · 2001
Earlier work this paper cites.
Off-policy temporal-difference learning with funtion approximation
Doina Precup, Richard S. Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
An introduction to MCMC for machine learning
Christophe Andrieu, Nando de Freitas, Arnaud Doucet, and Michael I. Jordan · 2002
Earlier work this paper cites.
Real analysis and probability
R. M. Dudley · 2002
Earlier work this paper cites.
Deeper inside PageRank
Amy N. Langville and Carl D. Meyer · 2004
Earlier work this paper cites.
Maximum mean discrepancy
A. J. Smola, A. Gretton, and K. Borgwardt · 2006
Earlier work this paper cites.
Discriminative learning for differing training and test distributions
Steffen Bickel, Michael Brückner, and Tobias Scheffer · 2007
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Dataset shift in machine learning
Arthur Gretton, Alexander J. Smola, Jiayuan Huang, Marcel Schmittfull, Karsten Borgwardt, and Bernhard Schölkopf · 2008
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by penalized convex risk minimization
X.L. Nguyen, M. Wainwright, and M. Jordan · 2008
Earlier work this paper cites.
Direct importance estimation with model selection and its application to covariate shift adaptation
M. Sugiyama, S. Nakajima, H. Kashima, P. von Bünau, and M. Kawanabe · 2008
Earlier work this paper cites.
A least-squares approach to direct importance estimation
Takafumi Kanamori, Shohei Hido, and Masashi Sugiyama · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.