Fetching the paper…
Reading the bibliography…
We investigate the problem of best-policy identification in discounted Markov Decision Processes (MDPs) when the learner has access to a generative model.
Sequential design of experiments
Chernoff, H. (1959) · 1959
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. and Robbins, H. (1985) · 1985
Earlier work this paper cites.
Approximate Distributions of Order Statistics: With Applications to Nonparametric Statistics
Reiss, R.-D. (1989) · 1989
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
Kearns, M. and Singh, S. (1999) · 1999
Earlier work this paper cites.
Almost horizon-free structure-aware best policy identification with a generative model
Zanette, A., Kochenderfer, M. J., and Brunskill, E. (2019) · 1999
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M. (2003) · 2003
Cited alongside, same era.
Minimax PAC
Gheshlaghi Azar, M., Munos, R., and Kappen, H. J. (2013) · 2013
Cited alongside, same era.
Optimal best arm identification with fixed confidence
Garivier, A. and Kaufmann, E. (2016) · 2016
Cited alongside, same era.
On the complexity of best-arm identification in multi-armed bandit models
Kaufmann, E., Cappé, O., and Garivier, A. (2016) · 2016
Cited alongside, same era.
Mixture martingales revisited with applications to sequential tests and confidence intervals
Kaufmann, E. and Koolen, W. M. (2018) · 2018
Cited alongside, same era.
Near-optimal time and sample complexities for solving discounted markov decision process with a generative model
Sidford, A., Wang, M., Wu, X., Yang, L. F., and Ye, Y. (2018) · 2018
Later among the works it cites.
Model-based reinforcement learning with a generative model is minimax optimal
Agarwal, A., Kakade, S., and Yang, L. F. (2019) · 2019
Later among the works it cites.
Non-asymptotic sequential tests for overlapping hypotheses and application to near optimal arm identification in bandit models
Garivier, A. and Kaufmann, E. (2019) · 2019
Later among the works it cites.
Planning in markov decision processes with gap-dependent sample complexity
Jonsson, A., Kaufmann, E., Ménard, P., Domingues, O. D., Leurent, E., and Valko, M. (2020) · 2020
Closest in time.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Li, G., Wei, Y., Chi, Y., Gu, Y., and Chen, Y. (2020) · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…