Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Later among the works it cites.
Concentration inequalities for Markov chains by Marton couplings and spectral methods
Paulin, D. (2015) · 2015
Later among the works it cites.
Dynamic programming and optimal control (4th edition)
Bertsekas, D. P. (2017) · 2017
Later among the works it cites.
Stochastic variance reduction methods for policy evaluation
Du, S. S., Chen, J., Li, L., Xiao, L., and Zhou, D. (2017) · 2017
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J., Russo, D., and Singal, R. (2018) · 2018
Later among the works it cites.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I. (2018) · 2018
Later among the works it cites.
Q-learning with nearest neighbors
Shah, D. and Xie, Q. (2018) · 2018
Later among the works it cites.
Provably efficient q q -learning with low switching cost
Bai, Y., Xie, T., Jiang, N., and Wang, Y.-X. (2019) · 2019
Later among the works it cites.
Neural temporal-difference and q-learning converges to global optima
Cai, Q., Yang, Z., Lee, J. D., and Wang, Z. (2019) · 2019
Later among the works it cites.
Finite-time analysis of distributed TD(0) with linear function approximation on multi-agent reinforcement learning
Doan, T., Maguluri, S., and Romberg, J. (2019) · 2019
Later among the works it cites.
Provably efficient Q-learning with function approximation via distribution shift error checking oracle
Du, S. S., Luo, Y., Wang, R., and Zhang, H. (2019) · 2019
Later among the works it cites.
Finite-time performance bounds and adaptive learning rate selection for two time-scale reinforcement learning
Gupta, H., Srikant, R., and Ying, L. (2019) · 2019
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation and TD learning
Srikant, R. and Ying, L. (2019) · 2019
Later among the works it cites.
Two time-scale off-policy TD learning: Non-asymptotic analysis over Markovian samples
Xu, T., Zou, S., and Liang, Y. (2019) · 2019
Later among the works it cites.
Sample-optimal parametric Q-learning using linearly additive features
Yang, L. and Wang, M. (2019) · 2019
Later among the works it cites.
Finite-sample analysis for SARSA with linear function approximation
Zou, S., Xu, T., and Liang, Y. (2019) · 2019
Later among the works it cites.
Finite-time analysis of asynchronous stochastic approximation and Q-learning
Qu, G. and Wierman, A. (2020) · 2020
Closest in time.
Markov chain block coordinate descent
Sun, T., Sun, Y., Xu, Y., and Yin, W. (2020) · 2020
Closest in time.
Q-learning with UCB exploration is sample efficient for infinite-horizon MDP
Wang, Y., Dong, K., Chen, X., and Wang, L. (2020) · 2020
Closest in time.
A finite-time analysis of Q-learning with neural network function approximation
Xu, P. and Gu, Q. (2020) · 2020
Closest in time.