Fetching the paper…
Reading the bibliography…
The theory of reinforcement learning has focused on two fundamental problems: achieving low regret, and identifying $\epsilon$-optimal policies.
On tail probabilities for martingales
Freedman, D. A · 1975
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M · 2003
Earlier work this paper cites.
Action elimination and stopping conditions for the multi-armed bandit and reinforcement learning problems
Even-Dar, E., Mannor, S., Mansour, Y., and Mahadevan, S · 2006
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
Maurer, A. and Pontil, M · 2009
Earlier work this paper cites.
Introduction to nonparametric estimation., 2009
Tsybakov, A. B · 2009
Earlier work this paper cites.
Zhang, Z., Ji, X., and Du, S. S · 2009
Earlier work this paper cites.
Nearly minimax optimal reward-free reinforcement learning
Zhang, Z., Du, S. S., and Ji, X · 2010
Earlier work this paper cites.
Pac bounds for discounted mdps
Lattimore, T. and Hutter, M · 2012
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, H. J · 2013
Earlier work this paper cites.
Online learning in episodic markovian decision processes by relative entropy policy search
Zimin, A. and Neu, G · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C. and Brunskill, E · 2015
Cited alongside, same era.
Optimal best arm identification with fixed confidence
Garivier, A. and Kaufmann, E · 2016
Cited alongside, same era.
On the complexity of best-arm identification in multi-armed bandit models
Kaufmann, E., Cappé, O., and Garivier, A · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Dann, C., Lattimore, T., and Brunskill, E · 2017
Cited alongside, same era.
Is q-learning provably efficient?
Model-based reinforcement learning with a generative model is minimax optimal
Agarwal, A., Kakade, S., and Yang, L. F · 2020
Later among the works it cites.
Planning in markov decision processes with gap-dependent sample complexity
Jonsson, A., Kaufmann, E., Ménard, P., Domingues, O. D., Leurent, E., and Valko, M · 2020
Later among the works it cites.
Is temporal difference learning optimal? an instance-dependent analysis
Khamaru, K., Pananjady, A., Ruan, F., Wainwright, M. J., and Jordan, M. I · 2020
Later among the works it cites.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Li, G., Wei, Y., Chi, Y., Gu, Y., and Chen, Y · 2020
Later among the works it cites.
Best policy identification in discounted mdps: Problem-specific sample complexity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Cited alongside, same era.
Exploration in structured reinforcement learning
Ok, J., Proutiere, A., and Tranos, D · 2018
Cited alongside, same era.
Sidford, A., Wang, M., Wu, X., Yang, L. F., and Ye, Y · 2018
Cited alongside, same era.
Policy certificates: Towards accountable reinforcement learning
Dann, C., Li, L., Wei, W., and Brunskill, E · 2019
Cited alongside, same era.
Pure exploration with multiple correct answers
Degenne, R. and Koolen, W. M · 2019
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Simchowitz, M. and Jamieson, K · 2019
Cited alongside, same era.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Zanette, A. and Brunskill, E · 2019
Cited alongside, same era.
Marjani, A. A. and Proutiere, A · 2020
Later among the works it cites.
Fast active learning for pure exploration in reinforcement learning
Ménard, P., Domingues, O. D., Jonsson, A., Kaufmann, E., Leurent, E., and Valko, M · 2020
Later among the works it cites.
Is long horizon reinforcement learning more difficult than short horizon reinforcement learning?
Wang, R., Du, S. S., Yang, L. F., and Kakade, S. M · 2020
Later among the works it cites.
Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning
Dann, C., Marinov, T. V., Mohri, M., and Zimmert, J · 2021
Closest in time.
Instance-optimality in optimal value estimation: Adaptivity via variance-reduced q-learning
Khamaru, K., Xia, E., Wainwright, M. J., and Jordan, M. I · 2021
Closest in time.
Navigating to the best policy in markov decision processes
Marjani, A. A., Garivier, A., and Proutiere, A · 2021
Closest in time.
Task-optimal exploration in linear dynamical systems
Wagenmaker, A., Simchowitz, M., and Jamieson, K · 2021
Closest in time.
Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap
Xu, H., Ma, T., and Du, S. S · 2021
Closest in time.