Fetching the paper…
Reading the bibliography…
In environments with uncertain dynamics exploration is necessary to learn how to perform well.
Neuro-Dynamic Programming
Bertsekas, Dimitri P. and Tsitsiklis, John N · 1996
Earlier work this paper cites.
Reinforcement learning: an introduction
Sutton, Richard S. and Barto, Andrew G · 1998
Earlier work this paper cites.
Constrained Markov Decision Processes
Altman, Eitan · 1999
Earlier work this paper cites.
A survey of computational complexity results in systems and control
Blondel, Vincent D. and Tsitsiklis, John N · 2000
Earlier work this paper cites.
R-MAX - A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning
Brafman, Ronen I. and Tennenholtz, Moshe · 2001
Earlier work this paper cites.
Near-Optimal Reinforcement Learning in Polynomial Time
Kearns, Michael and Singh, Satinder · 2002
Earlier work this paper cites.
Robust Control of Markov Decision Processes with Uncertain Transition Matrices
Nilim, Arnab and El Ghaoui, Laurent · 2005
Cited alongside, same era.
A theoretical analysis of Model-Based Interval Estimation
Strehl, Alexander L. and Littman, Michael L · 2005
Cited alongside, same era.
Introduction: Mars Science Laboratory: The Next Generation of Mars Landers
Lockwood, Mary Kae · 2006
Cited alongside, same era.
Percentile optimization in uncertain Markov decision processes with application to efficient exploration
Delage, Erick and Mannor, Shie · 2007
Cited alongside, same era.
Safe exploration for reinforcement learning
Hans, A, Schneegaß, D, Schäfer, AM, and Udluft, S · 2008
Cited alongside, same era.
Knows what it knows: a framework for self-aware learning
Li, Lihong, Littman, Michael L., and Walsh, Thomas J · 2008
Later among the works it cites.
Near-Bayesian exploration in polynomial time
Kolter, J. Zico and Ng, Andrew Y · 2009
Later among the works it cites.
UAV Cooperative Control with Stochastic Risk Models
Geramifard, A, Redding, J, Roy, N, and How, J P · 2011
Later among the works it cites.
Guaranteed safe online learning of a bounded system
Gillula, Jeremy H. and Tomlin, Claire J · 2011
Later among the works it cites.
Extensions of Learning-Based Model Predictive Control for Real-Time Application to a Quadrotor Helicopter
Aswani, Anil and Bouffard, Patrick · 2012
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…