Fetching the paper…
Reading the bibliography…
We study the policy evaluation problem in multi-agent reinforcement learning, modeled by a Markov decision process.
R. Horn and C. Johnson, Matrix Analysis . Cambridge, U.K.: Cambridge Univ. Press, 1985
1985
Earlier work this paper cites.
R. S. Sutton, “Learning to predict by the methods of temporal differences,” Machine Learning , vol. 3, no. 1, pp. 9–44, Aug 1988
1988
Earlier work this paper cites.
P. Dayan, “The convergence of TD( λ \lambda ) for general λ \lambda ,” pp. 341–362, 1992
1992
Earlier work this paper cites.
L. Gurvits, L. J. Lin, and S. J. Hanson, “Incremental learning of evaluation functions for absorbing Markov chains: New methods and theorems,” 1994
1994
Earlier work this paper cites.
G. Tesauro, “Temporal difference learning and td-gammon,” Commun. ACM , vol. 38, no. 3, pp. 58–68, 1995
1995
Earlier work this paper cites.
S. Bradtke and A. Barto, “Linear least-squares algorithms for temporal difference learning,” Machine Learning , vol. 22, no. 1, pp. 33–57, Mar 1996
1996
Earlier work this paper cites.
F. J. Pineda, “Mean-field theory for batched TD( λ \lambda ),” Neural Computation , vol. 9, pp. 1403–1419, 1997
1997
Earlier work this paper cites.
J. N. Tsitsiklis and B. V. Roy, “An analysis of temporal-difference learning with function approximation,” IEEE Transactions on Automatic Control , vol. 42, no. 5, pp. 674–690, 1997
1997
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction , 1st ed. MIT Press, 1998
1998
Earlier work this paper cites.
D. Bertsekas and J. Tsitsiklis, Neuro-Dynamic Programming , 2nd ed. Athena Scientific, Belmont, MA, 1999
1999
Earlier work this paper cites.
——, “Average cost temporal-difference learning,” Automatica , vol. 35, pp. 1799–1808, 1999
1999
Earlier work this paper cites.
V. Borkar and S. Meyn, “The o.d.e. method for convergence of stochastic approximation and reinforcement learning,” SIAM Journal on Control and Optimization , vol. 38, no. 2, pp. 447–469, 2000
2000
Earlier work this paper cites.
A. Nedić and D. P. Bertsekas, “Least squares policy evaluation algorithms with linear function approximation,” Discrete Event Dynamic Systems , vol. 13, no. 1, pp. 79–110, Jan 2003
2003
Earlier work this paper cites.
J. Cortes, S. Martinez, T. Karatas, and F. Bullo, “Coverage control for mobile sensing networks,” IEEE Transactions on Robotics and Automation , vol. 20, no. 2, pp. 243–255, 2004
2004
Earlier work this paper cites.
P. Ogren, E. Fiorelli, and N. E. Leonard, “Cooperative control of mobile sensor networks:adaptive gradient climbing in a distributed environment,” IEEE Transactions on Automatic Control , vol. 49, no. 8, pp. 1292–1302, 2004
2004
Earlier work this paper cites.
P. Abbeel, A. Coates, M. Quigley, and A. Ng, “An application of reinforcement learning to aerobatic helicopter flight,” in Advances in Neural Information Processing Systems 19 , 2007, pp. 1–8
2007
Cited alongside, same era.
V. Borkar, Stochastic Approximation: A Dynamical Systems Viewpoint . Cambridge University Press, 2008
2008
Cited alongside, same era.
R. S. Sutton, H. R. Maei, and C. Szepesvári, “A convergent o(n) temporal-difference algorithm for off-policy learning with linear function approximation,” in Advances in Neural Information Processing Systems 21 , 2009, pp. 1609–1616
2009
Cited alongside, same era.
R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora, “Fast gradient-descent methods for temporal-difference learning with linear function approximation,” in Proceedings of the 26th Annual International Conference on Machine Learning , ser. ICML ’09, 2009, pp. 993–1000
2009
Cited alongside, same era.
M. S. Stanković and S. S. Stanković, “Multi-agent temporal-difference learning with linear function approximation: Weak convergence under time-varying network topologies,” in 2016 American Control Conference (ACC) , 2016, pp. 167–172
2016
Later among the works it cites.
S. Gu, E. Holly, T. Lillicrap, and S. Levine, “Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates,” 2017 IEEE International Conference on Robotics and Automation (ICRA) , pp. 3389–3396, 2017
2017
Later among the works it cites.
A. Mathkar and V. S. Borkar, “Distributed reinforcement learning via gossip,” IEEE Transactions on Automatic Control , vol. 62, no. 3, pp. 1465–1470, 2017
2017
Later among the works it cites.
G. Dalal, B. Szörényi, G. Thoppe, and S. Mannor, “Finite sample analyses for TD(0) with function approximation,” in AAAI , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Yu and D. P. Bertsekas, “Convergence results for some temporal difference methods based on least squares,” IEEE Transactions on Automatic Control , vol. 54, no. 7, pp. 1515–1531, 2009
2009
Cited alongside, same era.
C. Szepesvari, Algorithms for Reinforcement Learning . Morgan and Claypool Publishers, 2010
2010
Cited alongside, same era.
H. Yu, “Convergence of least squares temporal difference methods under general conditions,” in Proceedings of the 27th International Conference on International Conference on Machine Learning , ser. ICML’10, 2010
2010
Cited alongside, same era.
S. Kar, J. M. F. Moura, and H. V. Poor, “Qd-learning: A collaborative distributed strategy for multi-agent reinforcement learning through consensus + innovations,” IEEE Trans. Signal Processing , vol. 61, pp. 1848–1862, 2013
2013
Cited alongside, same era.
M. Bennis, S. M. Perlaza, P. Blasco, Z. Han, and H. V. Poor, “Self-organization in small cell networks: A reinforcement learning approach,” IEEE Transactions on Wireless Communications , vol. 12, no. 7, pp. 3202–3212, 2013
2013
Cited alongside, same era.
2013
Cited alongside, same era.
C. Chen, A. Seff, A. Kornhauser, and J. Xiao, “Deepdriving: Learning affordance for direct perception in autonomous driving,” in Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV) , ser. ICCV ’15, Washington, DC, USA, 2015, pp. 2722–2730
2015
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nature , vol. 518, pp. 529–33, 02 2015
2015
Cited alongside, same era.
J. Bhandari, D. Russo, and R. Singal, “A finite time analysis of temporal difference learning with linear function approximation,” in COLT , 2018
2018
Later among the works it cites.
C. Lakshminarayanan and C. Szepesvari, “Linear stochastic approximation: How far does constant step-size and iterate averaging go?” in Proceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics , 2018, pp. 1347–1355
2018
Later among the works it cites.
K. Zhang, Z. Yang, H. Liu, T. Zhang, and T. Basar, “Fully decentralized multi-agent reinforcement learning with networked agents,” in Proceedings of the 35th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 80, 2018, pp. 5872–5881
2018
Later among the works it cites.
H.-T. Wai, Z. Yang, Z. Wang, and M. Hong, “Multi-agent reinforcement learning via double averaging primal-dual optimization,” in Annual Conference on Neural Information Processing Systems , 2018, pp. 9672–9683
2018
Later among the works it cites.
S. Tu and B. Recht, “Least-squares temporal difference learning for the linear quadratic regulator,” in Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018 , 2018, pp. 5012–5021
2018
Later among the works it cites.
2018
Later among the works it cites.
T. T. Doan, S. T. Maguluri, and J. Romberg, “Finite-time analysis of distributed TD(0) with linear function approximation on multi-agent reinforcement learning,” in Proceedings of the 36th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 97, Long Beach, California, USA, 2019, pp. 1626–1635
2019
Closest in time.
D. Lee and N. He, “Target-based temporal-difference learning,” in Proceedings of the 36th International Conference on Machine Learning , 2019, pp. 3713–3722
2019
Closest in time.
R. Srikant and L. Ying, “Finite-time error bounds for linear stochastic approximation and TD learning,” in COLT , 2019
2019
Closest in time.
B. Hu and U. Syed, “Characterizing the exact behaviors of temporal difference learning algorithms using markov jump linear system theory,” in Advances in Neural Information Processing Systems 32 , 2019
2019
Closest in time.