Fetching the paper…
Reading the bibliography…
In this paper, we propose a distributed off-policy actor critic method to solve multi-agent reinforcement learning problems.
1903
Earlier work this paper cites.
D. Lee, H. Yoon, and N. Hovakimyan, “Primal-dual algorithm for distributed reinforcement learning: distributed gtd,” in 2018 IEEE Conference on Decision and Control (CDC) . IEEE, 2018, pp. 1967–1972
1972
Earlier work this paper cites.
D. P. Bertsekas, Dynamic Programming and Optimal Control . Athena Scientific, 1995
1995
Earlier work this paper cites.
L. Baird, “Residual algorithms: Reinforcement learning with function approximation,” in Machine Learning Proceedings 1995 . Elsevier, 1995, pp. 30–37
1995
Earlier work this paper cites.
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Advances in neural information processing systems , 2000, pp. 1057–1063
2000
Earlier work this paper cites.
V. R. Konda and J. N. Tsitsiklis, “Actor-critic algorithms,” in Advances in neural information processing systems , 2000, pp. 1008–1014
2000
Earlier work this paper cites.
R. S. Sutton, C. Szepesvári, and H. R. Maei, “A convergent o (n) algorithm for off-policy temporal-difference learning with linear function approximation,” Advances in neural information processing systems , vol. 21, no. 21, pp. 1609–1616, 2008
2008
Earlier work this paper cites.
V. S. Borkar, Stochastic approximation: a dynamical systems viewpoint . Springer, 2009, vol. 48
2009
Earlier work this paper cites.
H. R. Maei, C. Szepesvári, S. Bhatnagar, and R. S. Sutton, “Toward off-policy learning control with function approximation.” in ICML , 2010, pp. 719–726
2010
Earlier work this paper cites.
P. Pennesi and I. C. Paschalidis, “A distributed actor-critic algorithm and applications to mobile sensor network coordination problems,” IEEE Transactions on Automatic Control , vol. 55, no. 2, pp. 492–497, 2010
2010
Cited alongside, same era.
T. Degris, M. White, and R. Sutton, “Off-policy actor-critic,” in International Conference on Machine Learning , 2012
2012
Cited alongside, same era.
2013
Cited alongside, same era.
P. Bianchi and J. Jakubowicz, “Convergence of a multi-agent projected stochastic gradient algorithm for non-convex optimization,” IEEE Transactions on Automatic Control , vol. 58, no. 2, pp. 391–405, 2013
2013
Cited alongside, same era.
2017
Later among the works it cites.
K. Yuan, B. Ying, J. Liu, and A. H. Sayed, “Variance-reduced stochastic learning by networked agents under random reshuffling,” IEEE Transactions on Signal Processing , vol. 67, no. 2, pp. 351–366, 2017
2017
Later among the works it cites.
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative-competitive environments,” in Advances in Neural Information Processing Systems , 2017, pp. 6379–6390
2017
Later among the works it cites.
S. Srinivasan, M. Lanctot, V. Zambaldi, J. Pérolat, K. Tuyls, R. Munos, and M. Bowling, “Actor-critic policy optimization in partially observable multiagent environments,” in Advances in Neural Information Processing Systems , 2018, pp. 3426–3439
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in ICML , 2014
2014
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, p. 529, 2015
2015
Cited alongside, same era.
S. V. Macua, J. Chen, S. Zazo, and A. H. Sayed, “Distributed policy evaluation under multiple behavior strategies,” IEEE Transactions on Automatic Control , vol. 60, no. 5, pp. 1260–1274, 2015
2015
Cited alongside, same era.
M. S. Stanković and S. S. Stanković, “Multi-agent temporal-difference learning with linear function approximation: Weak convergence under time-varying network topologies,” in 2016 American Control Conference (ACC) . IEEE, 2016, pp. 167–172
2016
Cited alongside, same era.
H.-T. Wai, Z. Yang, P. Z. Wang, and M. Hong, “Multi-agent reinforcement learning via double averaging primal-dual optimization,” in Advances in Neural Information Processing Systems , 2018, pp. 9672–9683
2018
Later among the works it cites.
K. Zhang, Z. Yang, H. Liu, T. Zhang, and T. Basar, “Fully decentralized multi-agent reinforcement learning with networked agents,” in International Conference on Machine Learning , 2018, pp. 5867–5876
2018
Later among the works it cites.
K. Zhang, Z. Yang, and T. Basar, “Networked multi-agent reinforcement learning in continuous spaces,” in 2018 IEEE Conference on Decision and Control (CDC) . IEEE, 2018, pp. 2771–2776
2018
Later among the works it cites.
K. Zhang, Z. Yang, and T. Başar, “Networked multi-agent reinforcement learning in continuous spaces,” in CDC , 2018
2018
Later among the works it cites.