Fetching the paper…
Reading the bibliography…
A central issue lying at the heart of online reinforcement learning (RL) is data efficiency.
Q-learning with UCB exploration is sample efficient for infinite-horizon MDP
Dong, K., Wang, Y., Chen, X., and Wang, L. (2019) · 1901
Earlier work this paper cites.
Wainwright, M. J. (2019a) · 1905
Earlier work this paper cites.
Variance-reduced Q-learning is minimax optimal
Wainwright, M. J. (2019b) · 1906
Earlier work this paper cites.
Provably efficient exploration in policy optimization
Cai, Q., Yang, Z., Jin, C., and Wang, Z. (2019) · 1912
Earlier work this paper cites.
On tail probabilities for martingales
Freedman, D. A. (1975) · 1975
Earlier work this paper cites.
R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M. (2003) · 2003
Earlier work this paper cites.
Learning rates for Q-learning
Even-Dar, E. and Mansour, Y. (2003) · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, S. M. (2003) · 2003
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J. (2020) · 2005
Earlier work this paper cites.
On optimism in model-based reinforcement learning
Pacchiano, A., Ball, P., Parker-Holder, J., Choromanski, K., and Roberts, S. (2020) · 2006
Earlier work this paper cites.
PAC model-free reinforcement learning
Strehl, A. L., Li, L., Wiewiora, E., Langford, J., and Littman, M. L. (2006) · 2006
Earlier work this paper cites.
A unifying view of optimism in episodic reinforcement learning
Neu, G. and Pike-Burke, C. (2020) · 2007
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Strehl, A. L. and Littman, M. L. (2008) · 2008
Earlier work this paper cites.
Regal: a regularization based algorithm for reinforcement learning in weakly communicating mdps
Bartlett, P. L. and Tewari, A. (2009) · 2009
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
Kolter, J. Z. and Ng, A. Y. (2009) · 2009
Earlier work this paper cites.
Empirical Bernstein bounds and sample variance penalization
Maurer, A. and Pontil, M. (2009) · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Earlier work this paper cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Szita, I. and Szepesvári, C. (2010) · 2010
Earlier work this paper cites.
Error bounds for constant step-size Q-learning
Beck, C. L. and Srikant, R. (2012) · 2012
Earlier work this paper cites.
PAC bounds for discounted MDPs
Lattimore, T. and Hutter, M. (2012) · 2012
Earlier work this paper cites.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, H. J. (2013) · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B. (2013) · 2013
Earlier work this paper cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C. and Brunskill, E. (2015) · 2015
Earlier work this paper cites.
Open problem: First-order regret bounds for contextual bandits
Agarwal, A., Krishnamurthy, A., Langford, J., Luo, H., et al. (2017) · 2017
Earlier work this paper cites.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Agrawal, S. and Jia, R. (2017) · 2017
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Cited alongside, same era.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Dann, C., Lattimore, T., and Brunskill, E. (2017) · 2017
Cited alongside, same era.
Make the minority great again: First-order regret bound for contextual bandits
Allen-Zhu, Z., Bubeck, S., and Li, Y. (2018) · 2018
Cited alongside, same era.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Fruit, R., Pirotta, M., Lazaric, A., and Ortner, R. (2018) · 2018
Cited alongside, same era.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Jiang, N. and Agarwal, A. (2018) · 2018
Cited alongside, same era.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I. (2018) · 2018
Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning
Dann, C., Marinov, T. V., Mohri, M., and Zimmert, J. (2021) · 2021
Later among the works it cites.
Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited
Domingues, O. D., Ménard, P., Kaufmann, E., and Valko, M. (2021) · 2021
Later among the works it cites.
Is pessimism provably efficient for offline RL?
Jin, Y., Yang, Z., and Wang, Z. (2021) · 2021
Later among the works it cites.
UCB momentum Q-learning: Correcting the bias without forgetting
Ménard, P., Domingues, O. D., Shang, X., and Valko, M. (2021) · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S. (2021) · 2021
Later among the works it cites.
Nearly horizon-free offline reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Variance-aware regret bounds for undiscounted reinforcement learning in mdps
Talebi, M. S. and Maillard, O.-A. (2018) · 2018
Cited alongside, same era.
Provably efficient Q-learning with low switching cost
Bai, Y., Xie, T., Jiang, N., and Wang, Y.-X. (2019) · 2019
Cited alongside, same era.
Reinforcement learning and optimal control
Bertsekas, D. (2019) · 2019
Cited alongside, same era.
Policy certificates: Towards accountable reinforcement learning
Dann, C., Li, L., Wei, W., and Brunskill, E. (2019) · 2019
Cited alongside, same era.
Tight regret bounds for model-based reinforcement learning with greedy policies
Efroni, Y., Merlis, N., Ghavamzadeh, M., and Mannor, S. (2019) · 2019
Cited alongside, same era.
Worst-case regret bounds for exploration via randomized value functions
Russo, D. (2019) · 2019
Cited alongside, same era.
Ren, T., Li, J., Dai, B., Du, S. S., and Sanghavi, S. (2021) · 2021
Later among the works it cites.
Stochastic shortest path: Minimax, parameter-free and towards horizon-free regret
Tarbouriech, J., Zhou, R., Du, S. S., Pirotta, M., Valko, M., and Lazaric, A. (2021) · 2021
Later among the works it cites.
A fully problem-dependent regret lower bound for finite-horizon MDPs
Tirinzoni, A., Pirotta, M., and Lazaric, A. (2021) · 2021
Later among the works it cites.
Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Xie, T., Jiang, N., Wang, H., Xiong, C., and Bai, Y. (2021) · 2021
Later among the works it cites.
Fine-grained gap-dependent bounds for tabular MDPs via adaptive multi-step bootstrap
Xu, H., Ma, T., and Du, S. (2021) · 2021
Later among the works it cites.
Q Q -learning with logarithmic regret
Yang, K., Yang, L., and Du, S. (2021) · 2021
Later among the works it cites.
Minimax-optimal multi-agent RL in Markov games with a generative model
Li, G., Chi, Y., Wei, Y., and Chen, Y. (2022) · 2022
Later among the works it cites.
Pessimistic Q-learning for offline reinforcement learning: Towards optimal sample complexity
Shi, L., Li, G., Wei, Y., Chen, Y., and Chi, Y. (2022) · 2022
Later among the works it cites.
First-order regret in reinforcement learning with linear function approximation: A robust estimation approach
Wagenmaker, A. J., Chen, Y., Simchowitz, M., Du, S., and Jamieson, K. (2022) · 2022
Later among the works it cites.
On gap-dependent bounds for offline reinforcement learning
Wang, X., Cui, Q., and Du, S. S. (2022) · 2022
Later among the works it cites.
Near-optimal randomized exploration for tabular markov decision processes
Xiong, Z., Shen, R., Cui, Q., Fazel, M., and Du, S. S. (2022) · 2022
Later among the works it cites.
Yin, M., Duan, Y., Wang, M., and Wang, Y.-X. (2022) · 2022
Later among the works it cites.
Horizon-free reinforcement learning in polynomial time: the power of stationary policies
Zhang, Z., Ji, X., and Du, S. (2022) · 2022
Later among the works it cites.
Regret-optimal model-free reinforcement learning for discounted MDPs with short burn-in time
Ji, X. and Li, G. (2023) · 2023
Closest in time.
The curious price of distributional robustness in reinforcement learning with a generative model
Shi, L., Li, G., Wei, Y., Chen, Y., Geist, M., and Chi, Y. (2023) · 2023
Closest in time.
The benefits of being distributional: Small-loss bounds for reinforcement learning
Wang, K., Zhou, K., Wu, R., Kallus, N., and Sun, W. (2023) · 2023
Closest in time.
The efficacy of pessimism in asynchronous Q-learning
Yan, Y., Li, G., Chen, Y., and Fan, J. (2023) · 2023
Closest in time.
Zhao, H., He, J., Zhou, D., Zhang, T., and Gu, Q. (2023) · 2023
Closest in time.
Sharp variance-dependent bounds in reinforcement learning: Best of both worlds in stochastic and deterministic environments
Zhou, R., Zihan, Z., and Du, S. S. (2023) · 2023
Closest in time.