Reinforcement learning in finite mdps: Pac analysis
Alexander L Strehl, Lihong Li, and Michael L Littman · 2009
Cited alongside, same era.
Real time targeted exploration in large domains
Todd Hester and Peter Stone · 2010
Cited alongside, same era.
Interval estimation for reinforcement-learning algorithms in continuous-state domains
Original
Martha White and Adam White · 2010
Cited alongside, same era.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Cited alongside, same era.
Agnostic system identification for model-based reinforcement learning
Stephane Ross and J. Andrew Bagnell · 2012
Cited alongside, same era.
Counterfactual reasoning and learning systems: the example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quinonero Candela, Denis Xavier Charles, Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Y Simard, and Ed Snelson · 2013
Cited alongside, same era.
Off-policy Evaluation in Markov Decision Processes
Cosmin Paduraru · 2013
Cited alongside, same era.