Fetching the paper…
Reading the bibliography…
Various algorithms in reinforcement learning exhibit dramatic variability in their convergence rates and ultimate accuracy as a function of the problem structure.
1905
Earlier work this paper cites.
Martin L. Puterman and Shelby L. Brumelle, On the convergence of policy iteration in stationary dynamic programming , Mathematics of Operations Research 4
1979
Earlier work this paper cites.
Lucien Birgé, Approximation dans les espaces métriques et théorie de l’estimation , Zeitschrift für Wahrscheinlichkeitstheorie und Verwandte Gebiete 65
1983
Earlier work this paper cites.
Christopher JCH Watkins and Peter Dayan, Q Q -learning , Machine Learning 8
1992
Earlier work this paper cites.
Tommi Jaakkola, Michael I Jordan, and Satinder P Singh, On the convergence of stochastic iterative dynamic programming algorithms , Neural Computation 6
1994
Earlier work this paper cites.
John N Tsitsiklis, Asynchronous stochastic approximation and Q Q -learning , Machine Learning 16
1994
Earlier work this paper cites.
Csaba Szepesvári, The asymptotic convergence-rate of Q Q -learning , Advances in Neural Information Processing Systems, vol. 10, 1997, pp. 1064–1070
1997
Earlier work this paper cites.
Dimitri P. Bertsekas, Neuro-dynamic programming , Springer US, Boston, MA, 2009
2009
Earlier work this paper cites.
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen, Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model , Machine Learning 91
2013
Earlier work this paper cites.
Rie Johnson and Tong Zhang, Accelerating stochastic gradient descent using predictive variance reduction , Advances in Neural Information Processing Systems, vol. 26, 2013, pp. 315–323
2013
Earlier work this paper cites.
Odalric-Ambrym Maillard, Timothy A Mann, and Shie Mannor, How hard is my MDP? ”The distribution-norm to the rescue” , Advances in Neural Information Processing Systems, vol. 27, 2014, pp. 1835–1843
2014
Cited alongside, same era.
Martin L Puterman, Markov Decision Processes: Discrete stochastic dynamic programming , John Wiley & Sons, 2014
2014
Cited alongside, same era.
T Cai and Mark Low, A framework for estimation of convex functions , Statistica Sinica 25
2015
Cited alongside, same era.
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel, End-to-end training of deep visuomotor policies , Journal of Machine Learning Research 17
2016
Cited alongside, same era.
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, and Marc Lanctot, Mastering the game of Go with deep neural networks and tree search , Nature 529
Richard S. Sutton and Andrew G. Barto, Reinforcement learning: An introduction , second ed., The MIT Press, 2018
2018
Later among the works it cites.
Aaron Sidford, Mengdi Wang, Xian Wu, Lin F Yang, and Yinyu Ye, Near-optimal time and sample complexities for solving markov decision processes with a generative model , Advances in Neural Information Processing Systems, vol. 33, 2018, pp. 5192–5202
2018
Later among the works it cites.
Aaron Sidford, Mengdi Wang, Xian Wu, and Yinyu Ye, Variance reduced value iteration and faster algorithms for solving markov decision processes , ACM-SIAM Symposium on Discrete Algorithms, vol. 29, SIAM, 2018, pp. 770–787
2018
Later among the works it cites.
Max Simchowitz and Kevin Jamieson, Non-asymptotic gap-dependent regret bounds for tabular MDPs , Advances in Neural Information Processing Systems, vol. 33, 2019, pp. 1153–1162
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel, Domain randomization for transferring deep neural networks from simulation to the real world , International Conference on Intelligent Robots and Systems (IROS), IEEE, 2017, pp. 23–30
2017
Cited alongside, same era.
Jalaj Bhandari, Daniel Russo, and Raghav Singal, A finite time analysis of temporal difference learning with linear function approximation , Conference On Learning Theory, PMLR, 2018, pp. 1691–1692
2018
Cited alongside, same era.
Gal Dalal, Balázs Szörényi, Gugan Thoppe, and Shie Mannor, Finite sample analyses for TD (0) with function approximation , AAAI Conference on Artificial Intelligence, vol. 32, 2018, pp. 6144–6153
2018
Cited alongside, same era.
Chandrashekar Lakshminarayanan and Csaba Szepesvari, Linear stochastic approximation: How far does constant step-size and iterate averaging go? , AISTATS: Conference on AI and Statistics, vol. 21, PMLR, 2018, pp. 1347–1355
2018
Cited alongside, same era.
Martin J. Wainwright, High-dimensional statistics: A non-asymptotic viewpoint , Cambridge Series in Statistical and Probabilistic Mathematics, Cambridge University Press, 2019
2019
Later among the works it cites.
Andrea Zanette and Emma Brunskill, Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds , International Conference on Machine Learning, PMLR, 2019, pp. 7304–7312
2019
Later among the works it cites.
Andrea Zanette, Mykel J Kochenderfer, and Emma Brunskill, Almost horizon-free structure-aware best policy identification with a generative model , Advances in Neural Information Processing Systems, vol. 32, 2019, pp. 5625–5634
2019
Later among the works it cites.
2020
Later among the works it cites.
A. Pananjady and M. J. Wainwright, Instance-dependent ℓ ∞ \ell_{\infty} -bounds for policy evaluation in tabular reinforcement learning , IEEE Transactions on Information Theory 67
2020
Later among the works it cites.