Fetching the paper…
Reading the bibliography…
This paper develops an unified framework to study finite-sample convergence guarantees of a large class of value-based asynchronous reinforcement learning (RL) algorithms.
Chen, Z., Zhang, S., Doan, T. T., Clarke, J.-P., and Maguluri, S. T. (2019) · 1905
Earlier work this paper cites.
Wainwright, M. J. (2019) · 1905
Earlier work this paper cites.
Sur les opérations dans les ensembles abstraits et leur application aux équations intégrales
Banach, S. (1922) · 1922
Earlier work this paper cites.
Improved upper bounds on the expected error in constant step-size Q Q -learning
Beck, C. L. and Srikant, R. (2013) · 1931
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S. (1951) · 1951
Earlier work this paper cites.
Watkins, C. J. C. H. (1989) · 1989
Earlier work this paper cites.
Q Q -learning
Watkins, C. J. and Dayan, P. (1992) · 1992
Earlier work this paper cites.
TD( λ \lambda ) converges with probability 1 1
Dayan, P. and Sejnowski, T. J. (1994) · 1994
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Jaakkola, T., Jordan, M. I., and Singh, S. P. (1994) · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and Q Q -learning
Tsitsiklis, J. N. (1994) · 1994
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L. (1995) · 1995
Earlier work this paper cites.
Neuro-dynamic programming
Bertsekas, D. P. and Tsitsiklis, J. N. (1996) · 1996
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
Singh, S. P. and Sutton, R. S. (1996) · 1996
Earlier work this paper cites.
The asymptotic convergence-rate of Q Q -learning
Szepesvári, C. et al. (1997) · 1997
Earlier work this paper cites.
Analysis of temporal-difference learning with function approximation
Tsitsiklis, J. N. and Van Roy, B. (1997) · 1997
Earlier work this paper cites.
Open theoretical questions in reinforcement learning
Sutton, R. S. (1999) · 1999
Earlier work this paper cites.
Average cost temporal-difference learning
Tsitsiklis, J. N. and Van Roy, B. (1999) · 1999
Cited alongside, same era.
The ODE method for convergence of stochastic approximation and reinforcement learning
Borkar, V. S. and Meyn, S. P. (2000) · 2000
Cited alongside, same era.
Bias-Variance Error Bounds for Temporal Difference Updates
Kearns, M. J. and Singh, S. P. (2000) · 2000
Cited alongside, same era.
Eligibility Traces for Off-Policy Policy Evaluation
Precup, D., Sutton, R. S., and Singh, S. P. (2000) · 2000
Cited alongside, same era.
Learning rates for Q Q -learning
Even-Dar, E. and Mansour, Y. (2003) · 2003
Cited alongside, same era.
Boundedness of iterates in Q Q -learning
Gosavi, A. (2006) · 2006
Cited alongside, same era.
First-order methods in optimization
Beck, A. (2017) · 2017
Later among the works it cites.
Zap Q Q -learning
Devraj, A. M. and Meyn, S. (2017) · 2017
Later among the works it cites.
Markov chains and mixing times
Levin, D. A. and Peres, Y. (2017) · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. (2017) · 2017
Later among the works it cites.
A Finite Time Analysis of Temporal Difference Learning With Linear Function Approximation
Bhandari, J., Russo, D., and Singal, R. (2018) · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, G., Wei, Y., Chi, Y., Gu, Y., and Chen, Y. (2020) · 2006
Cited alongside, same era.
Stochastic approximation: a dynamical systems viewpoint
Borkar, V. S. (2009) · 2009
Cited alongside, same era.
Accelerating Optimization and Reinforcement Learning with Quasi-Stochastic Approximation
Chen, S., Devraj, A., Bernstein, A., and Meyn, S. (2020a) · 2009
Cited alongside, same era.
Stochastic approximation: a survey
Kushner, H. (2010) · 2010
Cited alongside, same era.
Adaptive algorithms and stochastic approximations
Benveniste, A., Métivier, M., and Priouret, P. (2012) · 2012
Cited alongside, same era.
Stochastic approximation methods for constrained and unconstrained systems
Kushner, H. J. and Clark, D. S. (2012) · 2012
Cited alongside, same era.
Devraj, A. M., Bušic, A., and Meyn, S. (2018) · 2018
Later among the works it cites.
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al. (2018) · 2018
Later among the works it cites.
Is Q Q -learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Policy certificates: Towards accountable reinforcement learning
Dann, C., Li, L., Wei, W., and Brunskill, E. (2019) · 2019
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation and TD learning
Srikant, R. and Ying, L. (2019) · 2019
Later among the works it cites.
First-order and Stochastic Optimization Methods for Machine Learning
Lan, G. (2020) · 2020
Later among the works it cites.
Finite-Time Analysis of Asynchronous Stochastic Approximation and Q Q -Learning
Qu, G. and Wierman, A. (2020) · 2020
Later among the works it cites.
Finite-Sample Analysis of Off-Policy Natural Actor-Critic Algorithm
Khodadadian, S., Chen, Z., and Maguluri, S. T. (2021) · 2021
Closest in time.
Is Q Q -learning minimax optimal? a tight sample complexity analysis
Li, G., Cai, C., Chen, Y., Wei, Y., and Chi, Y. (2023) · 2023
Closest in time.