Fetching the paper…
Reading the bibliography…
This work develops a fully decentralized multi-agent algorithm for policy evaluation.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International conference on machine learning , 2016, pp. 1928–1937
1937
Earlier work this paper cites.
R. S. Sutton, “Learning to predict by the methods of temporal differences,” Machine learning , vol. 3, no. 1, pp. 9–44, 1988
1988
Earlier work this paper cites.
S. P. Singh and R. S. Sutton, “Reinforcement learning with replacing eligibility traces,” Machine learning , vol. 22, no. 1-3, pp. 123–158, 1996
1996
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Advances in Neural Information Processing Systems (NIPS) , Denver, USA, 2000, pp. 1057–1063
2000
Earlier work this paper cites.
2005
Earlier work this paper cites.
S.-Q. Shen, T.-Z. Huang, and G.-H. Cheng, “A condition for the nonsymmetric saddle point matrix being diagonalizable and having real and positive eigenvalues,” Journal of Computational and Applied Mathematics , vol. 220, no. 1-2, pp. 8–12, 2008
2008
Earlier work this paper cites.
R. S. Sutton, H. R. Maei, and C. Szepesvári, “A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation,” in Advances in Neural Information Processing Systems (NIPS) , Vancouver, Canada, 2009, pp. 1609–1616
2009
Earlier work this paper cites.
R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora, “Fast gradient-descent methods for temporal-difference learning with linear function approximation,” in Proceedings of the 26th Annual International Conference on Machine Learning , Montreal, Canada, 2009, pp. 993–1000
2009
Earlier work this paper cites.
A. Nedic and A. Ozdaglar, “Distributed subgradient methods for multi-agent optimization,” IEEE Transactions on Automatic Control , vol. 54, no. 1, p. 48, 2009
2009
Earlier work this paper cites.
H. R. Maei, “Gradient temporal-difference learning algorithms,” University of Alberta , 2011
2011
Earlier work this paper cites.
I. Grondman, L. Busoniu, G. A. Lopes, and R. Babuska, “A survey of actor-critic reinforcement learning: Standard and natural policy gradients,” IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews) , vol. 42, no. 6, pp. 1291–1307, 2012
2012
Earlier work this paper cites.
B. Kehoe, A. Matsukawa, S. Candido, J. Kuffner, and K. Goldberg, “Cloud-based robot grasping with the Google object recognition engine,” in IEEE International Conference on Robotics and Automation (ICRA) , Karlsruhe, Germany, May 2013, pp. 4263–4270
2013
Earlier work this paper cites.
R. Johnson and T. Zhang, “Accelerating stochastic gradient descent using predictive variance reduction,” in Advances in Neural Information Processing Systems (NIPS) , Lake Tahoe, USA, 2013, pp. 315–323
2013
Earlier work this paper cites.
J. F. Mota, J. M. Xavier, P. M. Aguiar, and M. Püschel, “D-admm: A communication-efficient distributed algorithm for separable optimization,” IEEE Transactions on Signal Processing , vol. 61, no. 10, pp. 2718–2723, 2013
2013
Earlier work this paper cites.
H. van Hasselt, A. R. Mahmood, and R. S. Sutton, “Off-policy TD( λ \lambda ) with a true online equivalence,” in Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence , Quebec City, Canada, 2014, pp. 330–339
2014
Cited alongside, same era.
A. Defazio, F. Bach, and S. Lacoste-Julien, “Saga: A fast incremental gradient method with support for non-strongly convex composite objectives,” in Advances in Neural Information Processing Systems (NIPS) , Montreal, Canada, 2014, pp. 1646–1654
2014
Cited alongside, same era.
A. H. Sayed, “Adaptive networks,” Proceedings of the IEEE , vol. 102, no. 4, pp. 460–497, 2014
2014
Cited alongside, same era.
A. H. Sayed, “Adaptation, learning, and optimization over networks,” Foundations and Trends in Machine Learning , vol. 7, 2014
2014
Cited alongside, same era.
S. Gu, E. Holly, T. Lillicrap, and S. Levine, “Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates,” in IEEE International Conference on Robotics and Automation (ICRA) , Singapore, Singapore, May 2017, pp. 3389–3396
2017
Later among the works it cites.
S. S. Du, J. Chen, L. Li, L. Xiao, and D. Zhou, “Stochastic variance reduction methods for policy evaluation,” in Proc. International Conference on Machine Learning , Sydney, Australia, 2017, pp. 1049–1058
2017
Later among the works it cites.
A. Nedic, A. Olshevsky, and W. Shi, “Achieving geometric convergence for distributed optimization over time-varying graphs,” SIAM Journal on Optimization , vol. 27, no. 4, pp. 2597–2633, 2017
2017
Later among the works it cites.
G. Qu and N. Li, “Harnessing smoothness to accelerate distributed optimization,” IEEE Transactions on Control of Network Systems , vol. 5, no. 3, pp. 1245–1260, 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Dann, G. Neumann, and J. Peters, “Policy evaluation with temporal differences: A survey and comparison,” The Journal of Machine Learning Research , vol. 15, no. 1, pp. 809–883, 2014
2014
Cited alongside, same era.
B. Kehoe, S. Patil, P. Abbeel, and K. Goldberg, “A survey of research on cloud robotics and automation,” IEEE Transactions on automation science and engineering , vol. 12, no. 2, pp. 398–409, 2015
2015
Cited alongside, same era.
B. Liu, J. Liu, M. Ghavamzadeh, S. Mahadevan, and M. Petrik, “Finite-sample analysis of proximal gradient td algorithms.” in Conference on Uncertainty in Artificial Intelligence (UAI) , Amsterdam, Holland, 2015, pp. 504–513
2015
Cited alongside, same era.
S. V. Macua, J. Chen, S. Zazo, and A. H. Sayed, “Distributed policy evaluation under multiple behavior strategies,” IEEE Transactions on Automatic Control , vol. 60, no. 5, pp. 1260–1274, 2015
2015
Cited alongside, same era.
W. Shi, Q. Ling, G. Wu, and W. Yin, “Extra: An exact first-order algorithm for decentralized consensus optimization,” SIAM Journal on Optimization , vol. 25, no. 2, pp. 944–966, 2015
2015
Cited alongside, same era.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al. , “Mastering the game of Go with deep neural networks and tree search,” Nature , vol. 529, no. 7587, p. 484, 2016
2016
Cited alongside, same era.
M. S. Stanković and S. S. Stanković, “Multi-agent temporal-difference learning with linear function approximation: Weak convergence under time-varying network topologies,” in Proc. American Control Conference , Boston, USA, July 2016, pp. 167–172
2016
Cited alongside, same era.
K. Yuan, Q. Ling, and W. Yin, “On the convergence of decentralized gradient descent,” SIAM Journal on Optimization , vol. 26, no. 3, pp. 1835–1854, 2016
2016
Cited alongside, same era.
C. Xi and U. A. Khan, “Dextra: A fast algorithm for optimization over directed graphs,” IEEE Transactions on Automatic Control , vol. 62, no. 10, pp. 4980–4993, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
B. Ying, K. Yuan, and A. H. Sayed, “Convergence of variance-reduced stochastic learning under random reshuffling,” in Proc. International Conference on Acoustics, Speech and Signal Processing (ICASSP) , Alberta, Canada, April 2018, pp. 2286–2290
2018
Closest in time.
2018
Closest in time.
B. Dai, A. Shaw, L. Li, L. Xiao, N. He, Z. Liu, J. Chen, and L. Song, “SBEED: Convergent reinforcement learning with nonlinear function approximation,” in International Conference on Machine Learning , 2018, pp. 1133–1142
2018
Closest in time.
K. Zhang, Z. Yang, H. Liu, T. Zhang, and T. Başar, “Fully decentralized multi-agent reinforcement learning with networked agents,” in Proc. International Conference on Machine Learning (ICML) , Stockholm, Sweden, July 2018, pp. 10–15
2018
Closest in time.
K. Yuan, B. Ying, X. Zhao, and A. H. Sayed, “Exact diffusion for distributed optimization and learning—part I: Algorithm development,” IEEE Transactions on Signal Processing , vol. 67, no. 3, pp. 708–723, 2018
2018
Closest in time.
——, “Exact diffusion for distributed optimization and learning—part II: Convergence analysis,” IEEE Transactions on Signal Processing , vol. 67, no. 3, pp. 724–739, 2018
2018
Closest in time.
L. Cassano, K. Yuan, and A. H. Sayed, “Distributed value-function learning with linear convergence rates,” in to appear in Proceedings of European Control Conference , Napoli, Italy, June 2019
2019
Closest in time.