Fetching the paper…
Reading the bibliography…
Q-learning, which seeks to learn the optimal Q-function of a Markov decision process (MDP) in a model-free fashion, lies at the heart of reinforcement learning.
A theoretical analysis of deep Q-learning
Fan, J., Wang, Z., Xie, Y., and Yang, Z. (2019) · 1901
Earlier work this paper cites.
Performance of Q-learning with linear function approximation: Stability and finite-time analysis
Chen, Z., Zhang, S., Doan, T. T., Maguluri, S. T., and Clarke, J.-P. (2019) · 1905
Earlier work this paper cites.
Wainwright, M. J. (2019b) · 1905
Earlier work this paper cites.
Variance-reduced Q-learning is minimax optimal
Wainwright, M. J. (2019c) · 1906
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S. (1951) · 1951
Earlier work this paper cites.
On the theory of dynamic programming
Bellman, R. (1952) · 1952
Earlier work this paper cites.
On tail probabilities for martingales
Freedman, D. A. (1975) · 1975
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H. (1989) · 1989
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B. (1992) · 1992
Earlier work this paper cites.
Q-learning
Watkins, C. J. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Jaakkola, T., Jordan, M. I., and Singh, S. P. (1994) · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and Q-learning
Tsitsiklis, J. N. (1994) · 1994
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Tsitsiklis, J. and Van Roy, B. (1997) · 1997
Earlier work this paper cites.
The asymptotic convergence-rate of Q-learning
Szepesvári, C. (1998) · 1998
Earlier work this paper cites.
Finite-sample convergence rates for Q-learning and indirect algorithms
Kearns, M. J. and Singh, S. P. (1999) · 1999
Earlier work this paper cites.
The ODE method for convergence of stochastic approximation and reinforcement learning
Borkar, V. S. and Meyn, S. P. (2000) · 2000
Earlier work this paper cites.
Finite-sample analysis of stochastic approximation using smooth convex envelopes
Chen, Z., Maguluri, S. T., Shakkottai, S., and Shanmugam, K. (2020) · 2002
Earlier work this paper cites.
Q-learning with uniformly bounded variance: Large discounting is not a barrier to fast learning
Devraj, A. M. and Meyn, S. P. (2020) · 2002
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large Markov decision processes
Kearns, M., Mansour, Y., and Ng, A. Y. (2002) · 2002
Earlier work this paper cites.
Learning rates for Q-learning
Even-Dar, E. and Mansour, Y. (2003) · 2003
Cited alongside, same era.
Nash Q-learning for general-sum stochastic games
Hu, J. and Wellman, M. P. (2003) · 2003
Cited alongside, same era.
On the sample complexity of reinforcement learning
Kakade, S. (2003) · 2003
Cited alongside, same era.
On linear stochastic approximation: Fine-grained Polyak-Ruppert and non-asymptotic concentration
Mou, W., Li, C. J., Wainwright, M. J., Bartlett, P. L., and Jordan, M. I. (2020) · 2004
Cited alongside, same era.
A generalization error for Q-learning
Murphy, S. (2005) · 2005
Cited alongside, same era.
A finite time analysis of two time-scale actor critic methods
Wu, Y., Zhang, W., Xu, P., and Gu, Q. (2020) · 2005
Near-optimal time and sample complexities for solving Markov decision processes with a generative model
Sidford, A., Wang, M., Wu, X., Yang, L., and Ye, Y. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Provably efficient Q-learning with low switching cost
Bai, Y., Xie, T., Jiang, N., and Wang, Y.-X. (2019) · 2019
Later among the works it cites.
Neural temporal-difference and Q-learning converges to global optima
Cai, Q., Yang, Z., Lee, J. D., and Wang, Z. (2019) · 2019
Later among the works it cites.
Finite-time analysis of distributed TD(0) with linear function approximation on multi-agent reinforcement learning
Doan, T., Maguluri, S., and Romberg, J. (2019) · 2019
Later among the works it cites.
Finite-time performance bounds and adaptive learning rate selection for two time-scale reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Momentum Q-learning with finite-sample convergence guarantee
Weng, B., Xiong, H., Zhao, L., Liang, Y., and Zhang, W. (2020a) · 2007
Cited alongside, same era.
Introduction to nonparametric estimation
Tsybakov, A. B. and Zaiats, V. (2009) · 2009
Cited alongside, same era.
Double Q-learning
Hasselt, H. (2010) · 2010
Cited alongside, same era.
Reinforcement learning with a near optimal rate of convergence
Azar, M. G., Munos, R., Ghavamzadeh, M., and Kappen, H. (2011) · 2011
Cited alongside, same era.
Freedman’s inequality for matrix martingales
Tropp, J. (2011) · 2011
Cited alongside, same era.
Error bounds for constant step-size Q-learning
Beck, C. L. and Srikant, R. (2012) · 2012
Cited alongside, same era.
Gupta, H., Srikant, R., and Ying, L. (2019) · 2019
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation and TD learning
Srikant, R. and Ying, L. (2019) · 2019
Later among the works it cites.
Variance reduced policy evaluation with smooth function approximation
Wai, H.-T., Hong, M., Yang, Z., Wang, Z., and Tang, K. (2019) · 2019
Later among the works it cites.
Model-based reinforcement learning with a generative model is minimax optimal
Agarwal, A., Kakade, S., and Yang, L. F. (2020) · 2020
Later among the works it cites.
Instance-dependent ℓ ∞ \ell_{\infty} -bounds for policy evaluation in tabular reinforcement learning
Pananjady, A. and Wainwright, M. J. (2020) · 2020
Later among the works it cites.
Finite-time analysis of asynchronous stochastic approximation and Q-learning
Qu, G. and Wierman, A. (2020) · 2020
Later among the works it cites.
Finite-time analysis for double Q-learning
Xiong, H., Zhao, L., Liang, Y., and Zhang, W. (2020) · 2020
Later among the works it cites.
A finite-time analysis of Q-learning with neural network function approximation
Xu, P. and Gu, Q. (2020) · 2020
Later among the works it cites.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zhang, Z., Zhou, Y., and Ji, X. (2020) · 2020
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J., Russo, D., and Singal, R. (2021) · 2021
Closest in time.
A Lyapunov theory for finite-sample guarantees of asynchronous Q-learning and TD-learning variants
Chen, Z., Maguluri, S. T., Shakkottai, S., and Shanmugam, K. (2021) · 2021
Closest in time.
Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning
Li, G., Shi, L., Chen, Y., and Chi, Y. (2021) · 2021
Closest in time.
Pessimistic Q-learning for offline reinforcement learning: Towards optimal sample complexity
Shi, L., Li, G., Wei, Y., Chen, Y., and Chi, Y. (2022) · 2022
Closest in time.
The efficacy of pessimism in asynchronous Q-learning
Yan, Y., Li, G., Chen, Y., and Fan, J. (2022) · 2022
Closest in time.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Li, G., Wei, Y., Chi, Y., and Chen, Y. (2023) · 2023
Closest in time.