Fetching the paper…
Reading the bibliography…
The goal of this paper is to study a distributed version of the gradient temporal-difference (GTD) learning algorithm for multi-agent Markov decision processes (MDPs).
L. Baird, “Residual algorithms: Reinforcement learning with function approximation,” in Machine Learning Proceedings 1995 , 1995, pp. 30–37
1995
Earlier work this paper cites.
D. P. Bertsekas and J. N. Tsitsiklis, Neuro-dynamic programming . Athena Scientific Belmont, MA, 1996
1996
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT Press, 1998
1998
Earlier work this paper cites.
D. P. Bertsekas, Nonlinear programming . Athena scientific Belmont, 1999
1999
Earlier work this paper cites.
V. S. Borkar and S. P. Meyn, “The ODE method for convergence of stochastic approximation and reinforcement learning,” SIAM Journal on Control and Optimization , vol. 38, no. 2, pp. 447–469, 2000
2000
Earlier work this paper cites.
H. Kushner and G. G. Yin, Stochastic approximation and recursive algorithms and applications . Springer Science & Business Media, 2003, vol. 35
2003
Earlier work this paper cites.
S. Boyd and L. Vandenberghe, Convex Optimization . Cambridge University Press, 2004
2004
Earlier work this paper cites.
R. S. Sutton, H. R. Maei, and C. Szepesvári, “A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation,” in Advances in neural information processing systems , 2009, pp. 1609–1616
2009
Earlier work this paper cites.
R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora, “Fast gradient-descent methods for temporal-difference learning with linear function approximation,” in Proceedings of the 26th Annual International Conference on Machine Learning , 2009, pp. 993–1000
2009
Earlier work this paper cites.
P. Pennesi and I. C. Paschalidis, “A distributed actor-critic algorithm and applications to mobile sensor network coordination problems,” IEEE Transactions on Automatic Control , vol. 55, no. 2, pp. 492–497, 2010
2010
Cited alongside, same era.
S. S. Ram, A. Nedić, and V. V. Veeravalli, “Distributed stochastic subgradient projection algorithms for convex optimization,” Journal of optimization theory and applications , vol. 147, no. 3, pp. 516–545, 2010
2010
Cited alongside, same era.
J. Wang and N. Elia, “Control approach to distributed optimization,” in 48th Annual Allerton Conference on Communication, Control, and Computing (Allerton) , 2010, pp. 557–561
2010
Cited alongside, same era.
A. Nedic, A. Ozdaglar, and P. A. Parrilo, “Constrained consensus and optimization in multi-agent networks,” IEEE Transactions on Automatic Control , vol. 55, no. 4, pp. 922–938, 2010
2010
Cited alongside, same era.
B. Gharesifard and J. Cortés, “Distributed continuous-time convex optimization on weight-balanced digraphs,” IEEE Transactions on Automatic Control , vol. 59, no. 3, pp. 781–786, 2014
2014
Later among the works it cites.
2014
Later among the works it cites.
S. V. Macua, J. Chen, S. Zazo, and A. H. Sayed, “Distributed policy evaluation under multiple behavior strategies,” IEEE Transactions on Automatic Control , vol. 60, no. 5, pp. 1260–1274, 2015
2015
Later among the works it cites.
M. S. Stanković and S. S. Stanković, “Multi-agent temporal-difference learning with linear function approximation: weak convergence under time-varying network topologies,” in American Control Conference (ACC) , 2016, pp. 167–172
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
——, “A control perspective for centralized and distributed convex optimization,” in 50th IEEE Conference on Decision and Control and European Control Conference (CDC-ECC) , 2011, pp. 3800–3805
2011
Cited alongside, same era.
S. Bhatnagar, H. Prasad, and L. Prashanth, Stochastic recursive algorithms for optimization: simultaneous perturbation methods . Springer, 2012, vol. 434
2012
Cited alongside, same era.
S. Kar, J. M. Moura, and H. V. Poor, “QD-learning: a collaborative distributed strategy for multi-agent reinforcement learning through consensus + + innovations,” IEEE Transactions on Signal Processing , vol. 61, no. 7, pp. 1848–1862, 2013
2013
Cited alongside, same era.
P. Bianchi and J. Jakubowicz, “Convergence of a multi-agent projected stochastic gradient algorithm for non-convex optimization,” IEEE Transactions on Automatic Control , vol. 58, no. 2, pp. 391–405, 2013
2013
Cited alongside, same era.
R. Tutunov, H. B. Ammar, and A. Jadbabaie, “An exact distributed newton method for reinforcement learning,” in 2016 IEEE 55th Conference on Decision and Control (CDC) , 2016, pp. 1003–1008
2016
Later among the works it cites.
2016
Later among the works it cites.
A. Mathkar and V. S. Borkar, “Distributed reinforcement learning via gossip,” IEEE Transactions on Automatic Control , vol. 62, no. 3, pp. 1465–1470, 2017
2017
Later among the works it cites.