Optimal Information Processing and Bayes’s Theorem
Arnold Zellner · 1988
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L Puterman · 1994
Earlier work this paper cites.
The maxq method for hierarchical reinforcement learning
Thomas G Dietterich · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S. Sutton, and Satinder P. Singh · 2000
Earlier work this paper cites.
Marginal mean models for dynamic regimes
S. Murphy, M. van der Laan, and J. Robins · 2001
Earlier work this paper cites.
The linear programming approach to approximate dynamic programming
Daniela Pucci De Farias and Benjamin Van Roy · 2003
Earlier work this paper cites.
Bayes meets bellman: The gaussian process approach to temporal difference learning
Yaakov Engel, Shie Mannor, and Ron Meir · 2003
Earlier work this paper cites.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Convex approximations of chance constrained programs
Arkadi Nemirovski and Alexander Shapiro · 2007
Earlier work this paper cites.
A hilbert space embedding for distributions
Alex Smola, Arthur Gretton, Le Song, and Bernhard Schölkopf · 2007
Earlier work this paper cites.
Robust optimization
Aharon Ben-Tal, Laurent El Ghaoui, and Arkadi Nemirovski · 2009
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
J Zico Kolter and Andrew Y Ng · 2009
Earlier work this paper cites.
Pac-bayesian model selection for reinforcement learning
Mahdi M Fard and Joelle Pineau · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Original
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
Model selection in reinforcement learning
Amir-massoud Farahmand and Csaba Szepesvári · 2011
Earlier work this paper cites.
Universality, characteristic kernels and rkhs embedding of measures
Bharath K Sriperumbudur, Kenji Fukumizu, and Gert RG Lanckriet · 2011
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X. Charles, D. Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Earlier work this paper cites.