Fetching the paper…
Reading the bibliography…
In many branches of engineering, Banach contraction mapping theorem is employed to establish the convergence of certain deterministic algorithms.
L. E. Dubins and D. A. Freedman, “Invariant probabilities for certain Markov processes,” The Annals of Mathematical Statistics
1966
Earlier work this paper cites.
A. Maitra, “Discounted dynamic programming on compact metric spaces,” Sankhyā: The Indian Journal of Statistics, Series A
1968
Earlier work this paper cites.
Springer Series in Statistics, Springer-Verlag New York, 1984
D. Pollard, Convergence of stochastic processes · 1984
Earlier work this paper cites.
M. F. Barnsley and S. Demko, “Iterated function systems and the global construction of fractals,” Proceedings of the Royal Society of London. A. Mathematical and Physical Sciences
1985
Earlier work this paper cites.
M. F. Barnsley, J. H. Elton, and D. P. Hardin, “Recurrent iterated function systems,” Constructive approximation
1989
Earlier work this paper cites.
Athena scientific Belmont, MA, 1995
D. P. Bertsekas, Dynamic programming and optimal control, Vol. I and II · 1995
Earlier work this paper cites.
Athena Scientific, 1996
D. P. Bertsekas and J. N. Tsitsiklis, Neuro-Dynamic Programming · 1996
Earlier work this paper cites.
Springer-Verlag New York, 1996
A. W. Van Der Vaart and J. A. Wellner, Weak Convergence and Empirical Processes With Applications to Statistics · 1996
Earlier work this paper cites.
J. N. Tsitsiklis and B. Van Roy, “Analysis of temporal-difference learning with function approximation,” in Advances in neural information processing systems
1997
Earlier work this paper cites.
John Wiley & Sons, Ltd., 1998
A.-A. Borovkov and V. Yurinsky, Ergodicity and Stability of Stochastic Processes · 1998
Earlier work this paper cites.
Springer Science & Business Media, 1998
S. T. Rachev and L. Rüschendorf, Mass Transportation Problems: Volume I: Theory · 1998
Earlier work this paper cites.
C. Kollman, K. Baggerly, D. Cox, and R. Picard, “Adaptive importance sampling on discrete Markov chains,” Annals of Applied Probability
1999
Earlier work this paper cites.
P. Diaconis and D. Freedman, “Iterated random functions,” SIAM review
1999
Earlier work this paper cites.
V. S. Borkar and S. P. Meyn, “The ODE method for convergence of stochastic approximation and reinforcement learning,” SIAM Journal on Control and Optimization
2000
Earlier work this paper cites.
Stanford University, 2001
P. Y. Desai, Adaptive Monte Carlo methods for solving eigenvalue problems · 2001
Earlier work this paper cites.
P. Y. Desai and P. W. Glynn, “A Markov chain perspective on adaptive Monte Carlo algorithms,” in Proceeding of the 2001 Winter Simulation Conference (Cat. No. 01CH37304)
2001
Earlier work this paper cites.
89, American Mathematical Soc., 2001
M. Ledoux, The concentration of measure phenomenon · 2001
Earlier work this paper cites.
J. Huang, I. Kontoyiannis, and S. P. Meyn, “The ODE method and spectral theory of Markov operators,” in Stochastic Theory and Control
2002
Earlier work this paper cites.
Springer, 2002
L. Györfi, Principles of nonparametric learning · 2002
Earlier work this paper cites.
Springer Science & Business Media, 2003
H. Kushner and G. G. Yin, Stochastic approximation and recursive algorithms and applications · 2003
Earlier work this paper cites.
S. Kakade, M. J. Kearns, and J. Langford, “Exploration in metric state spaces,” in Proceedings of the 20th International Conference on Machine Learning (ICML-03)
2003
Earlier work this paper cites.
K. Hinderer, “Lipschitz continuity of value functions in Markovian decision processes,” Mathematical Methods of Operations Research
2005
Earlier work this paper cites.
Springer-Verlag Berlin Heidelberg, 2006
C. D. Aliprantis and K. Border, Infinite Dimensional Analysis: A Hitchhiker’s Guide · 2006
Cited alongside, same era.
T. I. Ahamed, V. S. Borkar, and S. Juneja, “Adaptive importance sampling technique for markov chains using stochastic approximation,” Operations Research
2006
Cited alongside, same era.
Springer Science & Business Media, 2007
M. Shaked and J. G. Shanthikumar, Stochastic Orders · 2007
Cited alongside, same era.
Springer Science & Business Media, 2008
I. Steinwart and A. Christmann, Support vector machines · 2008
Cited alongside, same era.
R. Munos and C. Szepesvári, “Finite-time bounds for fitted value iteration,” Journal of Machine Learning Research
2008
Cited alongside, same era.
Cambridge University Press, 2009
S. P. Meyn and R. L. Tweedie, Markov Chains and Stochastic Stability · 2009
A. Gupta, R. Jain, and P. W. Glynn, “An empirical algorithm for relative value iteration for average-cost MDPs,” in Proc. of 54th IEEE Conference on Decision and Control (CDC)
2015
Later among the works it cites.
R. Harikandeh, M. O. Ahmed, A. Virani, M. Schmidt, J. Konečnỳ, and S. Sallinen, “Stopwasting my gradients: Practical SVRG,” in Advances in Neural Information Processing Systems
2015
Later among the works it cites.
W. B. Haskell, R. Jain, and D. Kalathil, “Empirical dynamic programming,” Mathematics of Operations Research
2016
Later among the works it cites.
A. Dieuleveut and F. Bach, “Nonparametric stochastic approximation with large step-sizes,” The Annals of Statistics
2016
Later among the works it cites.
W. B. Haskell, P. Yu, H. Sharma, and R. Jain, “Randomized function fitting-based empirical value iteration,” in 2017 IEEE 56th Annual Conference on Decision and Control (CDC)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Springer, 2009
V. S. Borkar, Stochastic approximation: A dynamical systems viewpoint · 2009
Cited alongside, same era.
A. A. Kulkarni and V. S. Borkar, “Finite dimensional approximation and newton-based algorithm for stochastic approximation in Hilbert space,” Automatica
2009
Cited alongside, same era.
Cambridge University Press, 2009
M. Anthony and P. L. Bartlett, Neural network learning: Theoretical foundations · 2009
Cited alongside, same era.
2009
Cited alongside, same era.
D. P. Bertsekas and H. Yu, “Q-learning and enhanced policy iteration in discounted dynamic programming,” Mathematics of Operations Research
2012
Cited alongside, same era.
Ö. Stenflo, “A survey of average contractive iterated function systems,” Journal of Difference Equations and Applications
2012
Cited alongside, same era.
2017
Later among the works it cites.
P. Yu, W. B. Haskell, and H. Xu, “Approximate value iteration for risk-aware markov decision processes,” IEEE Transactions on Automatic Control
2018
Closest in time.
Springer, 2018
R. Douc, E. Moulines, P. Priouret, and P. Soulier, Markov chains · 2018
Closest in time.
A. Sidford, M. Wang, X. Wu, L. Yang, and Y. Ye, “Near-optimal time and sample complexities for solving Markov decision processes with a generative model,” in Advances in Neural Information Processing Systems
2018
Closest in time.
D. Shah and Q. Xie, “Q-learning with nearest neighbors,” in Advances in Neural Information Processing Systems
2018
Closest in time.
E. A. Feinberg and J. Huang, “Reduction of total-cost and average-cost MDPs with weakly continuous transition probabilities to discounted MDPs,” Operations research letters
2018
Closest in time.
M. J. Wainwright, “Variance-reduced Q-learning is minimax optimal,” arXiv preprint arXiv:1906.04697
2019
Closest in time.
B. Hanin, “Universal function approximation by deep neural nets with bounded width and Relu activations,” Mathematics
2019
Closest in time.
W. B. Haskell, R. Jain, H. Sharma, and P. Yu, “A universal empirical dynamic programming algorithm for continuous state MDPs,” IEEE Transactions on Automatic Control
2019
Closest in time.
H. Sharma, M. Jafarnia-Jahromi, and R. Jain, “Approximate relative value learning for average-reward continuous state MDPs,” in Proceedings UAI
2019
Closest in time.
H. Sharma and R. Jain, “An approximately optimal relative value learning algorithm for averaged MDPs with continuous states and actions,” in 2019 57th Annual Allerton Conference on Communication, Control, and Computing (Allerton)
2019
Closest in time.
H. Sharma, R. Jain, and A. Gupta, “An empirical relative value learning algorithm for non-parametric MDPs with continuous state space,” in 2019 18th European Control Conference (ECC)
2019
Closest in time.
E. A. Feinberg and J. Huang, “On the reduction of total-cost and average-cost MDPs to discounted MDPs,” Naval Research Logistics (NRL)
2019
Closest in time.
J. Heng, A. N. Bishop, G. Deligiannidis, and A. Doucet, “Controlled sequential Monte Carlo,” The Annals of Statistics
2020
Closest in time.
A. Boustati, O. D. Akyildiz, T. Damoulas, and A. Johansen, “Generalised Bayesian filtering via sequential Monte Carlo,” Advances in neural information processing systems
2020
Closest in time.
A. Gupta and W. B. Haskell, “Convergence of recursive stochastic algorithms using wasserstein divergence,” SIAM Journal on Mathematics of Data Science
2021
Closest in time.
S. Shao, F. Harirchi, D. Dave, and A. Gupta, “Premptive scheduling of ev charging for providing demand response services,” submitted to IEEE Open Journal of Control Systems
2022
Closest in time.