Fetching the paper…
Reading the bibliography…
We extend trust region policy optimization (TRPO) to multi-agent reinforcement learning (MARL) problems.
1911
Earlier work this paper cites.
H. Kahn and T. E. Harris, “Estimation of particle transmission by random sampling,” National Bureau of Standards applied mathematics series , vol. 12, pp. 27–30, 1951
1951
Earlier work this paper cites.
Y.-C. Ho and K.-C. Chu, “Team decision theory and information structures in optimal control problems–part i,” IEEE Transactions on Automatic Control , vol. 17, no. 1, pp. 15–22, 1972
1972
Earlier work this paper cites.
D. Lee, H. Yoon, and N. Hovakimyan, “Primal-dual algorithm for distributed reinforcement learning: Distributed gtd,” in 2018 IEEE Conference on Decision and Control (CDC) , 2018, pp. 1967–1972
1972
Earlier work this paper cites.
S. P. Singh, T. S. Jaakkola, and M. I. Jordan, “Learning without state-estimation in partially observable markovian decision processes,” in In Proceedings of the 11st International Conference on Machine Learning (ICML) , 1994, pp. 284–292
1994
Earlier work this paper cites.
M. L. Littman, “Markov games as a framework for multi-agent reinforcement learning,” in Proceedings of the Eleventh International Conference on International Conference on Machine Learning , ser. ICML’94. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 1994, p. 157–163
1994
Earlier work this paper cites.
D. P. Bertsekas, Dynamic Programming and Optimal Control , 1st ed. Athena Scientific, 1995
1995
Earlier work this paper cites.
M. Tan, Multi-Agent Reinforcement Learning: Independent vs. Cooperative Agents . San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 1997, p. 487–494
1997
Earlier work this paper cites.
J. Ooi, S. Verbout, J. Ludwig, and G. Wornell, “A separation theorem for periodic sharing information patterns in decentralized control,” IEEE Transactions on Automatic Control , vol. 42, no. 11, pp. 1546–1550, 1997
1997
Earlier work this paper cites.
C. Claus and C. Boutilier, “The dynamics of reinforcement learning in cooperative multiagent systems,” in Proceedings of the Fifteenth National/Tenth Conference on Artificial Intelligence/Innovative Applications of Artificial Intelligence , ser. AAAI ’98/IAAI ’98. USA: American Association for Artificial Intelligence, 1998, p. 746–752
1998
Earlier work this paper cites.
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in the 12th Conference on Neural Information Processing Systems , ser. NeurIPS’99. Cambridge, MA, USA: MIT Press, 1999, p. 1057–1063
1999
Earlier work this paper cites.
V. S. Borkar and S. P. Meyn, “The o.d.e. method for convergence of stochastic approximation and reinforcement learning,” SIAM Journal on Control and Optimization , vol. 38, no. 2, pp. 447–469, 2000. [Online]. Available: https://doi.org/10.1137/S0363012997331639
2000
Earlier work this paper cites.
S. Boyd and L. Vandenberghe, Convex optimization . Cambridge university press, 2004
2004
Earlier work this paper cites.
S. Boyd, A. Ghosh, B. Prabhakar, and D. Shah, “Randomized gossip algorithms,” IEEE Transactions on Information Theory , vol. 52, no. 6, pp. 2508–2530, 2006
2006
Earlier work this paper cites.
J. Nocedal and S. J. Wright, Numerical Optimization , 2nd ed. New York, NY, USA: Springer, 2006
2006
Earlier work this paper cites.
M. Pipattanasomporn, H. Feroze, and S. Rahman, “Multi-agent systems in a distributed smart grid: Design and implementation,” in 2009 IEEE/PES Power Systems Conference and Exposition , 2009, pp. 1–8
2009
Earlier work this paper cites.
A. Nedic and A. Ozdaglar, “Distributed subgradient methods for multi-agent optimization,” IEEE Transactions on Automatic Control , vol. 54, no. 1, pp. 48–61, 2009
2009
Earlier work this paper cites.
N. Takahashi, I. Yamada, and A. H. Sayed, “Diffusion least-mean squares with adaptive combiners: Formulation and performance analysis,” IEEE Transactions on Signal Processing , vol. 58, no. 9, pp. 4795–4810, 2010
2010
Earlier work this paper cites.
A. Nayyar, A. Mahajan, and D. Teneketzis, “Optimal control strategies in delayed sharing information structures,” IEEE Transactions on Automatic Control , vol. 56, no. 7, pp. 1606–1620, 2011
2011
Earlier work this paper cites.
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein, “Distributed optimization and statistical learning via the alternating direction method of multipliers,” Found. Trends Mach. Learn. , vol. 3, no. 1, p. 1–122, Jan. 2011. [Online]. Available: https://doi.org/10.1561/2200000016
2011
Earlier work this paper cites.
D. P. Kroese, T. Taimre, and Z. I. Botev, Handbook of Monte Carlo Methods . Hoboken, NJ, USA: Wiley, 2011
2011
Earlier work this paper cites.
L. Matignon, G. J. Laurent, and N. Le Fort-Piat, “Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems,” The Knowledge Engineering Review , vol. 27, no. 1, p. 1–31, 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
A. Nayyar, A. Mahajan, and D. Teneketzis, “Decentralized stochastic control with partial history sharing: A common information approach,” IEEE Transactions on Automatic Control , vol. 58, no. 7, pp. 1644–1658, 2013
2013
Earlier work this paper cites.
S. Kar, J. M. F. Moura, and H. V. Poor, “ 𝒬𝒟 {{\cal Q}{\cal D}} -learning: A collaborative distributed strategy for multi-agent reinforcement learning through consensus + innovations {\rm consensus}+{\rm innovations} ,” IEEE Transactions on Signal Processing , vol. 61, no. 7, pp. 1848–1862, 2013
2013
Earlier work this paper cites.
E. Wei and A. Ozdaglar, “On the o(1=k) convergence of asynchronous distributed alternating direction method of multipliers,” in 2013 IEEE Global Conference on Signal and Information Processing , 2013, pp. 551–554
2013
Cited alongside, same era.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in Proceedings of the 31st International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, E. P. Xing and T. Jebara, Eds., vol. 32, no. 1. Bejing, China: PMLR, 22–24 Jun 2014, pp. 387–395. [Online]. Available: https://proceedings.mlr.press/v32/silver14.html
2014
Cited alongside, same era.
J. Schulman, S. Levine, P. Moritz, M. Jordan, and P. Abbeel, “Trust region policy optimization,” in Proceedings of the 32nd International Conference on International Conference on Machine Learning (ICML) , 2015, pp. 1889–1897
2015
Cited alongside, same era.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature , vol. 521, pp. 436 – 444, 2015
H.-T. Wai, Z. Yang, Z. Wang, and M. Hong, “Multi-agent reinforcement learning via double averaging primal-dual optimization,” in the 32nd Conference on Neural Information Processing Systems , ser. NeurIPS’18. Red Hook, NY, USA: Curran Associates Inc., 2018, p. 9672–9683
2018
Later among the works it cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction , 2nd ed. The MIT Press, 2018. [Online]. Available: http://incompleteideas.net/book/the-book-2nd.html
2018
Later among the works it cites.
K. Zhang, Z. Yang, H. Liu, T. Zhang, and T. Basar, “Fully decentralized multi-agent reinforcement learning with networked agents,” in Proceedings of the 35th International Conference on Machine Learning (ICML) , 10–15 Jul 2018, pp. 5872–5881
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
S. Valcarcel Macua, J. Chen, S. Zazo, and A. H. Sayed, “Distributed policy evaluation under multiple behavior strategies,” IEEE Transactions on Automatic Control , vol. 60, no. 5, pp. 1260–1274, 2015
2015
Cited alongside, same era.
S. Magnússon, P. C. Weeraddana, and C. Fischione, “A distributed approach for the optimal power-flow problem based on admm and sequential convex approximations,” IEEE Transactions on Control of Network Systems , vol. 2, no. 3, pp. 238–253, 2015
2015
Cited alongside, same era.
2016
Cited alongside, same era.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis, “Mastering the game of Go with deep neural networks and tree search,” Nature , vol. 529, no. 7587, pp. 484–489, Jan 2016
2016
Cited alongside, same era.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” Journal of Machine Learning Research , vol. 17, no. 1, pp. 1334 – 1373, Jan 2016
2016
Cited alongside, same era.
S. Sukhbaatar, A. Szlam, and R. Fergus, “Learning multiagent communication with backpropagation,” in the 30th Conference on Neural Information Processing Systems , ser. NeurIPS’16. Red Hook, NY, USA: Curran Associates Inc., 2016, pp. 2252–2260
2016
Cited alongside, same era.
M. S. Stankovi?? and S. S. Stankovi??, “Multi-agent temporal-difference learning with linear function approximation: Weak convergence under time-varying network topologies,” in 2016 American Control Conference (ACC) , 2016, pp. 167–172
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2019
Later among the works it cites.
S. Iqbal and F. Sha, “Actor-attention-critic for multi-agent reinforcement learning,” in Proceedings of the 36th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97. PMLR, 09–15 Jun 2019, pp. 2961–2970. [Online]. Available: https://proceedings.mlr.press/v97/iqbal19a.html
2019
Later among the works it cites.
T. T. Doan, S. T. Maguluri, and J. Romberg, “Finite-time analysis of distributed TD (0) with linear function approximation on multi-agent reinforcement learning,” in ICML , 2019
2019
Later among the works it cites.
Y. Zhang and M. M. Zavlanos, “Distributed off-policy actor-critic reinforcement learning with policy consensus,” in 2019 IEEE 58th Conference on Decision and Control (CDC) , 2019, pp. 4674–4679
2019
Later among the works it cites.
X. Wang, L. Ke, Z. Qiao, and X. Chai, “Large-scale traffic signal control using a novel multiagent reinforcement learning,” IEEE Transactions on Cybernetics , pp. 1–14, 2020
2020
Closest in time.
W. Mao, K. Zhang, E. Miehling, and T. Başar, “Information state embedding in partially observable cooperative multi-agent reinforcement learning,” in 2020 59th IEEE Conference on Decision and Control (CDC) , 2020, pp. 6124–6131
2020
Closest in time.
T. Chu, S. Chinchali, and S. Katti, “Multi-agent reinforcement learning for networked system control,” in International Conference on Learning Representations , 2020. [Online]. Available: https://openreview.net/forum?id=Syx7A3NFvH
2020
Closest in time.
T. T. Nguyen, N. D. Nguyen, and S. Nahavandi, “Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications,” IEEE Transactions on Cybernetics , vol. 50, no. 9, pp. 3826–3839, 2020
2020
Closest in time.
X. Liu and Y. Tan, “Attentive relational state representation in decentralized multiagent reinforcement learning,” IEEE Transactions on Cybernetics , pp. 1–13, 2020
2020
Closest in time.
D. Silver, S. Singh, D. Precup, and R. S. Sutton, “Reward is enough,” Artificial Intelligence , vol. 299, p. 103535, 2021. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0004370221000862
2021
Closest in time.
J. Jumper, R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, B. Romera-Paredes, S. Nikolov, R. Jain, J. Adler, T. Back, S. Petersen, D. Reiman, E. Clancy, M. Zielinski, M. Steinegger, M. Pacholska, T. Berghammer, S. Bodenstein, D. Silver, O. Vinyals, A. W. Senior, K. Kavukcuoglu, P. Kohli, and D. Hassabis, “Highly accurate protein structure prediction with AlphaFold,” Nature , vol. 596, no. 7873, pp. 583–589, 2021
2021
Closest in time.
J. Chai, W. Li, Y. Zhu, D. Zhao, Z. Ma, K. Sun, and J. Ding, “Unmas: Multiagent reinforcement learning for unshaped cooperative scenarios,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–12, 2021
2021
Closest in time.
K. T. X. Y. Peng Yang, Qi Yang, “Parallel exploration via negatively correlated search,” Frontiers of Computer Science , vol. 15, no. 155333, 2021
2021
Closest in time.
G. Hu, Y. Zhu, D. Zhao, M. Zhao, and J. Hao, “Event-triggered communication network with limited-bandwidth constraint for multi-agent reinforcement learning,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–13, 2021
2021
Closest in time.
H. Tavafoghi, Y. Ouyang, and D. Teneketzis, “A unified approach to dynamic decision problems with asymmetric information: Nonstrategic agents,” IEEE Transactions on Automatic Control , vol. 67, no. 3, pp. 1105–1119, 2022
2022
Closest in time.
Z. Pu, H. Wang, Z. Liu, J. Yi, and S. Wu, “Attention enhanced reinforcement learning for multi agent cooperation,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–15, 2022
2022
Closest in time.
D. Xie and X. Zhong, “Semicentralized deep deterministic policy gradient in cooperative starcraft games,” IEEE Transactions on Neural Networks and Learning Systems , vol. 33, no. 4, pp. 1584–1593, 2022
2022
Closest in time.
C. Yu, A. Velu, E. Vinitsky, J. Gao, Y. Wang, A. Bayen, and Y. Wu, “The surprising effectiveness of PPO in cooperative multi-agent games,” in 36th Conference on Neural Information Processing Systems, Datasets and Benchmarks Track , 2022. [Online]. Available: https://openreview.net/forum?id=YVXaxB6L2Pl
2022
Closest in time.
L. Matignon, L. Jeanpierre, and A.-I. Mouaddib, “Coordinated multi-robot exploration under communication constraints using decentralized markov decision processes,” in Proceedings of the Twenty-Sixth AAAI Conference on Artificial Intelligence , ser. AAAI’12. AAAI Press, 2012, p. 2017–2023
2023
Closest in time.
P. Sunehag, G. Lever, A. Gruslys, W. M. Czarnecki, V. Zambaldi, M. Jaderberg, M. Lanctot, N. Sonnerat, J. Z. Leibo, K. Tuyls, and T. Graepel, “Value-decomposition networks for cooperative multi-agent learning based on team reward,” in Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems , ser. AAMAS ’18. Richland, SC: International Foundation for Autonomous Agents and Multiagent Systems, 2018, p. 2085–2087
2087
Closest in time.