Fetching the paper…
Reading the bibliography…
Reinforcement learning, mathematically described by Markov Decision Problems, may be approached either through dynamic programming or policy search.
Mnih V, Badia AP, Mirza M, Graves A, Lillicrap T, Harley T, Silver D, Kavukcuoglu K (2016) Asynchronous methods for deep reinforcement learning. In: International Conference on Machine Learning, pp 1928–1937
1937
Earlier work this paper cites.
Bellman R (1954) The theory of dynamic programming. Tech. rep., RAND Corp Santa Monica CA
1954
Earlier work this paper cites.
Bellman RE (1957) Dynamic Programming. Courier Dover Publications
1957
Earlier work this paper cites.
Nesterov YE (1983) A method for solving the convex programming problem with convergence rate o (1/kˆ 2). In: Dokl. akad. nauk Sssr, vol 269, pp 543–547
1983
Earlier work this paper cites.
Sutton RS (1988) Learning to predict by the methods of temporal differences. Machine learning 3(1):9–44
1988
Earlier work this paper cites.
Cybenko G (1989) Approximation by superpositions of a sigmoidal function. Mathematics of control, signals and systems 2(4):303–314
1989
Earlier work this paper cites.
Park J, Sandberg IW (1991) Universal approximation using radial-basis-function networks. Neural computation 3(2):246–257
1991
Earlier work this paper cites.
Watkins CJ, Dayan P (1992) Q-learning. Machine learning 8(3-4):279–292
1992
Earlier work this paper cites.
Tsitsiklis JN (1994) Asynchronous stochastic approximation and q-learning. Machine learning 16(3):185–202
1994
Earlier work this paper cites.
Baird L (1995) Residual algorithms: Reinforcement learning with function approximation. In: Machine Learning Proceedings 1995, Elsevier, pp 30–37
1995
Earlier work this paper cites.
Bertsekas DP, Bertsekas DP, Bertsekas DP, Bertsekas DP (1995) Dynamic programming and optimal control, vol 1. Athena scientific Belmont, MA
1995
Earlier work this paper cites.
Boyan JA, Moore AW (1995) Generalization in reinforcement learning: Safely approximating the value function. In: Advances in neural information processing systems, pp 369–376
1995
Earlier work this paper cites.
Tesauro G, et al. (1995) Temporal difference learning and td-gammon. Communications of the ACM 38(3):58–68
1995
Earlier work this paper cites.
Borkar VS (1997) Stochastic approximation with two time scales. Systems & Control Letters 29(5):291–294
1997
Earlier work this paper cites.
Tsitsiklis JN, Van Roy B (1997) Analysis of temporal-diffference learning with function approximation. In: Advances in Neural Information Processing Systems, pp 1075–1081
1997
Earlier work this paper cites.
Bottou L (1998) Online learning and stochastic approximations. On-line learning in neural networks 17(9):142
1998
Earlier work this paper cites.
Konda VR, Borkar VS (1999) Actor-critic–type learning algorithms for Markov decision processes. SIAM Journal on Control and Optimization 38(1):94–123
1999
Earlier work this paper cites.
Borkar VS, Meyn SP (2000) The ode method for convergence of stochastic approximation and reinforcement learning. SIAM Journal on Control and Optimization 38(2):447–469
2000
Earlier work this paper cites.
Doya K (2000) Reinforcement learning in continuous time and space. Neural Computation 12(1):219–245
2000
Earlier work this paper cites.
Konda VR, Tsitsiklis JN (2000) Actor-critic algorithms. In: Advances in Neural Information Processing Systems, pp 1008–1014
2000
Earlier work this paper cites.
Sutton RS, McAllester DA, Singh SP, Mansour Y (2000) Policy gradient methods for reinforcement learning with function approximation. In: Advances in Neural Information Processing Systems, pp 1057–1063
2000
Earlier work this paper cites.
Bousquet O, Elisseeff A (2002) Stability and generalization. Journal of machine learning research 2(Mar):499–526
2002
Earlier work this paper cites.
Giannoccaro I, Pontrandolfo P (2002) Inventory management in supply chains: a reinforcement learning approach. International Journal of Production Economics 78(2):153–161
2002
Earlier work this paper cites.
Kushner HJ, Yin GG (2003) Stochastic approximation and recursive algorithms and applications. Springer, New York, NY
2003
Earlier work this paper cites.
Bertsekas DP (2005) Dynamic Programming and Optimal Control, vol 1
2005
Cited alongside, same era.
Powell WB (2007) Approximate Dynamic Programming: Solving the curses of dimensionality, vol 703. John Wiley & Sons
2007
Cited alongside, same era.
Antos A, Szepesvári C, Munos R (2008) Fitted q-iteration in continuous action-space mdps. In: Advances in neural information processing systems, pp 9–16
2008
Cited alongside, same era.
Bhatnagar S, Ghavamzadeh M, Lee M, Sutton RS (2008) Incremental natural actor-critic algorithms. In: Advances in Neural Information Processing Systems, pp 105–112
2008
Cited alongside, same era.
Sutton RS, Szepesvári C, Maei HR (2008) A convergent o (n) algorithm for off-policy temporal-difference learning with linear function approximation. Advances in neural information processing systems 21(21):1609–1616
2008
Jin C, Allen-Zhu Z, Bubeck S, Jordan MI (2018) Is q-learning provably efficient? In: Advances in Neural Information Processing Systems 31, pp 4863–4873
2018
Later among the works it cites.
Lakshminarayanan C, Szepesvari C (2018) Linear stochastic approximation: How far does constant step-size and iterate averaging go? In: International Conference on Artificial Intelligence and Statistics, pp 1347–1355
2018
Later among the works it cites.
Papini M, Binaghi D, Canonaco G, Pirotta M, Restelli M (2018) Stochastic variance-reduced policy gradient. In: International Conference on Machine Learning, pp 4026–4035
2018
Later among the works it cites.
Paternain S (2018) Stochastic control foundations of autonomous behavior. PhD thesis, University of Pennsylvania
2018
Later among the works it cites.
Thoppe G, Borkar V (2019) A concentration bound for stochastic approximation via alekseev’s formula. Stochastic Systems 9(1):1–26, DOI 10.1287/stsy.2018.0019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bhatnagar S, Sutton R, Ghavamzadeh M, Lee M (2009) Natural actor-critic algorithms. Automatica 45(11):2471–2482
2009
Cited alongside, same era.
Borkar VS (2009) Stochastic approximation: a dynamical systems viewpoint, vol 48. Springer
2009
Cited alongside, same era.
Castro DD, Meir R (2010) A convergent online single-time-scale actor-critic algorithm. Journal of Machine Learning Research 11(Jan):367–410
2010
Cited alongside, same era.
Maei HR, Szepesvári C, Bhatnagar S, Sutton RS (2010) Toward off-policy learning control with function approximation. In: Proceedings of the 27th International Conference on Machine Learning (ICML-10), pp 719–726
2010
Cited alongside, same era.
Kober J, Peters J (2012) Reinforcement learning in robotics: A survey. In: Reinforcement Learning, Springer, pp 579–610
2012
Cited alongside, same era.
Meyn SP, Tweedie RL (2012) Markov chains and stochastic stability. Springer Science & Business Media
2012
Cited alongside, same era.
Jiang DR, Pham TV, Powell WB, Salas DF, Scott WR (2014) A comparison of approximate dynamic programming techniques on benchmark energy storage problems: Does anything work? In: 2014 IEEE Symposium on Adaptive Dynamic Programming and Reinforcement Learning (ADPRL), IEEE, pp 1–8
2014
Cited alongside, same era.
2018
Later among the works it cites.
Tolstaya E, Koppel A, Stump E, Ribeiro A (2018) Nonparametric stochastic compositional gradient descent for q-learning in continuous markov decision problems. In: 2018 Annual American Control Conference (ACC), IEEE, pp 6608–6615
2018
Later among the works it cites.
Yang Z, Zhang K, Hong M, Başar T (2018) A finite sample analysis of the actor-critic algorithm. In: 2018 IEEE Conference on Decision and Control (CDC), IEEE, pp 2759–2764
2018
Later among the works it cites.
Zhang K, Yang Z, Liu H, Zhang T, Başar T (2018) Fully decentralized multi-agent reinforcement learning with networked agents. In: International Conference on Machine Learning, pp 5872–5881
2018
Later among the works it cites.
Cai Q, Yang Z, Lee JD, Wang Z (2019) Neural temporal-difference learning converges to global optima. Advances in Neural Information Processing Systems 32
2019
Closest in time.
Parisi S, Tangkaratt V, Peters J, Khan ME (2019) Td-regularized actor-critic methods. Machine Learning 108(8-9):1467–1501
2019
Closest in time.
Srikant R, Ying L (2019) Finite-time error bounds for linear stochastic approximation andtd learning. In: Conference on Learning Theory, PMLR, pp 2803–2830
2019
Closest in time.
Wang L, Cai Q, Yang Z, Wang Z (2019) Neural policy gradient methods: Global optimality and rates of convergence. In: International Conference on Learning Representations
2019
Closest in time.
Zhang K, Koppel A, Zhu H, Başar T (2019) Convergence and iteration complexity of policy gradient method for infinite-horizon reinforcement learning. IEEE Conference on Decision and Control
2019
Closest in time.
Zou S, Xu T, Liang Y (2019) Finite-sample analysis for sarsa with linear function approximation. Advances in neural information processing systems 32
2019
Closest in time.
Dalal G, Szörényi B, Thoppe G (2020) A tale of two-timescale reinforcement learning with the tightest finite-time bound. AAAI Press, pp 3701–3708, URL https://aaai.org/ojs/index.php/AAAI/article/view/5779
2020
Closest in time.
Shen H, Zhang K, Hong M, Chen T (2020) Asynchronous advantage actor critic: Non-asymptotic analysis and linear speedup. arXiv preprint arXiv:201215511
2020
Closest in time.
Wu YF, Zhang W, Xu P, Gu Q (2020) A finite-time analysis of two time-scale actor-critic methods. Advances in Neural Information Processing Systems 33:17617–17628
2020
Closest in time.
Xu T, Wang Z, Liang Y (2020) Non-asymptotic convergence analysis of two time-scale (natural) actor-critic algorithms. arXiv preprint arXiv:200503557
2020
Closest in time.
Zhang K, Koppel A, Zhu H, Basar T (2020) Global convergence of policy gradient methods to (almost) locally optimal policies. SIAM Journal on Control and Optimization 58(6):3586–3612
2020
Closest in time.
Qiu S, Yang Z, Ye J, Wang Z (2021) On finite-time convergence of actor-critic algorithm. IEEE Journal on Selected Areas in Information Theory 2(2):652–664
2021
Closest in time.
Cayci S, He N, Srikant R (2022) Finite-time analysis of entropy-regularized neural natural actor-critic algorithm. arXiv preprint arXiv:220600833
2022
Closest in time.
Olshevsky A, Gharesifard B (2022) A small gain analysis of single timescale actor critic. arXiv preprint arXiv:220302591
2022
Closest in time.
Zeng S, Chen T, Garcia A, Hong M (2022) Learning to coordinate in multi-agent systems: A coordinated actor-critic algorithm and finite-time guarantees. In: Learning for Dynamics and Control Conference, PMLR, pp 278–290
2022
Closest in time.