Fetching the paper…
Reading the bibliography…
Multi-agent reinforcement learning (MARL) has emerged as a useful approach to solving decentralised decision-making problems at scale.
1908
Earlier work this paper cites.
1909
Earlier work this paper cites.
B. Efron, Bootstrap Methods: Another Look at the Jackknife . New York, NY: Springer New York, 1992, pp. 569–593. [Online]. Available: https://doi.org/10.1007/978-1-4612-4380-9_41
1992
Earlier work this paper cites.
F. Christianos, G. Papoudakis, M. A. Rahman, and S. V. Albrecht, “Scaling multi-agent reinforcement learning with selective parameter sharing,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 1989–1998. [Online]. Available: https://proceedings.mlr.press/v139/christianos21a.html
1998
Earlier work this paper cites.
E. D. Dolan and J. J. Moré, “Benchmarking optimization software with performance profiles,” 2001. [Online]. Available: https://arxiv.org/abs/cs/0102001
2001
Earlier work this paper cites.
2006
Earlier work this paper cites.
2006
Earlier work this paper cites.
2011
Earlier work this paper cites.
S. Whiteson, B. Tanner, M. E. Taylor, and P. Stone, “Protecting against evaluation overfitting in empirical reinforcement learning,” in 2011 IEEE symposium on adaptive dynamic programming and reinforcement learning (ADPRL) . IEEE, 2011, pp. 120–127
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
A. Tampuu, T. Matiisen, D. Kodelja, I. Kuzovkin, K. Korjus, J. Aru, J. Aru, and R. Vicente, “Multiagent cooperation and competition with deep reinforcement learning,” 2015
2015
Earlier work this paper cites.
A. Tampuu, T. Matiisen, D. Kodelja, I. Kuzovkin, K. Korjus, J. Aru, J. Aru, and R. Vicente, “Multiagent cooperation and competition with deep reinforcement learning,” PLOS ONE , vol. 12, 11 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Sukhbaatar, a. szlam, and R. Fergus, “Learning multiagent communication with backpropagation,” in Advances in Neural Information Processing Systems , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, Eds., vol. 29. Curran Associates, Inc., 2016. [Online]. Available: https://proceedings.neurips.cc/paper/2016/file/55b1927fdafef39c48e5b73b5d61ea60-Paper.pdf
2016
Earlier work this paper cites.
J. Foerster, I. A. Assael, N. de Freitas, and S. Whiteson, “Learning to communicate with deep multi-agent reinforcement learning,” in Advances in Neural Information Processing Systems , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, Eds., vol. 29. Curran Associates, Inc., 2016. [Online]. Available: https://proceedings.neurips.cc/paper/2016/file/c7635bfd99248a2cdef8249ef7bfbef4-Paper.pdf
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. Lowe, Y. I. Wu, A. Tamar, J. Harb, O. Pieter Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative-competitive environments,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Omidshafiei, J. Pazis, C. Amato, J. P. How, and J. Vian, “Deep decentralized multi-task multi-agent reinforcement learning under partial observability,” in ICML , 2017, pp. 2681–2690. [Online]. Available: http://proceedings.mlr.press/v70/omidshafiei17a.html
2017
Earlier work this paper cites.
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch, “Multi-agent actor-critic for mixed cooperative-competitive environments,” in NIPS , 2017, pp. 6382–6393. [Online]. Available: http://papers.nips.cc/paper/7217-multi-agent-actor-critic-for-mixed-cooperative-competitive-environments
2017
Earlier work this paper cites.
J. Foerster, N. Nardelli, G. Farquhar, T. Afouras, P. H. S. Torr, P. Kohli, and S. Whiteson, “Stabilising experience replay for deep multi-agent reinforcement learning,” in Proceedings of the 34th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, D. Precup and Y. W. Teh, Eds., vol. 70. PMLR, 06–11 Aug 2017, pp. 1146–1155. [Online]. Available: https://proceedings.mlr.press/v70/foerster17b.html
2017
Earlier work this paper cites.
P. Henderson, “Reproducibility and reusability in deep reinforcement learning,” Master’s thesis, McGill University, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger, “Deep Reinforcement Learning that Matters,” 2018
2018
Earlier work this paper cites.
C. Colas, O. Sigaud, and P.-Y. Oudeyer, “How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments,” 2018
2018
Earlier work this paper cites.
J. N. Foerster, “Deep multi-agent reinforcement learning,” Ph.D. dissertation, University of Oxford, 2018
2018
Earlier work this paper cites.
T. Rashid, M. Samvelyan, C. Schroeder, G. Farquhar, J. Foerster, and S. Whiteson, “Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning,” in International Conference on Machine Learning . PMLR, 2018, pp. 4295–4304
2018
Earlier work this paper cites.
J. Foerster, G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson, “Counterfactual multi-agent policy gradients,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
C. Colas, O. Sigaud, and P.-Y. Oudeyer, “Gep-pg: Decoupling exploration and exploitation in deep reinforcement learning algorithms,” in International conference on machine learning . PMLR, 2018, pp. 1039–1048
2018
Earlier work this paper cites.
E. Wei, D. Wicke, D. Freelan, and S. Luke, “Multiagent soft q-learning,” in 2018 AAAI Spring Symposia, Stanford University, Palo Alto, California, USA, March 26-28, 2018 . AAAI Press, 2018. [Online]. Available: https://aaai.org/ocs/index.php/SSS/SSS18/paper/view/17508
2018
Earlier work this paper cites.
J. N. Foerster, G. Farquhar, T. Afouras, N. Nardelli, and S. Whiteson, “Counterfactual multi-agent policy gradients,” in AAAI . AAAI Press, 2018, pp. 2974–2982
2018
Earlier work this paper cites.
T. Rashid, M. Samvelyan, C. Schroeder, G. Farquhar, J. Foerster, and S. Whiteson, “QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning,” in Proceedings of the 35th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, J. Dy and A. Krause, Eds., vol. 80. PMLR, 10–15 Jul 2018, pp. 4295–4304. [Online]. Available: https://proceedings.mlr.press/v80/rashid18a.html
2018
Earlier work this paper cites.
——, “A Hitchhiker’s Guide to Statistical Comparisons of Reinforcement Learning Algorithms,” 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
N. Carion, N. Usunier, G. Synnaeve, and A. Lazaric, “A structured prediction approach for generalization in cooperative multi-agent reinforcement learning,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru, “Model cards for model reporting,” in Proceedings of the conference on fairness, accountability, and transparency , 2019, pp. 220–229
2019
Earlier work this paper cites.
S. Omidshafiei, C. Papadimitriou, G. Piliouras, K. Tuyls, M. Rowland, J.-B. Lespiau, W. M. Czarnecki, M. Lanctot, J. Perolat, and R. Munos, “ α \alpha -rank: Multi-agent evaluation by evolution,” Scientific reports , vol. 9, no. 1, pp. 1–29, 2019
2019
Earlier work this paper cites.
A. Singh, T. Jain, and S. Sukhbaatar, “Learning when to communicate at scale in multiagent cooperative and competitive tasks,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=rye7knCqK7
2019
Earlier work this paper cites.
S. Iqbal and F. Sha, “Actor-attention-critic for multi-agent reinforcement learning,” in Proceedings of the 36th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97. PMLR, 09–15 Jun 2019, pp. 2961–2970. [Online]. Available: https://proceedings.mlr.press/v97/iqbal19a.html
2019
Cited alongside, same era.
S. Q. Zhang, Q. Zhang, and J. Lin, “Efficient communication in multi-agent reinforcement learning via variance based control,” in Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. Alché-Buc, E. Fox, and R. Garnett, Eds., vol. 32. Curran Associates, Inc., 2019. [Online]. Available: https://proceedings.neurips.cc/paper/2019/file/14cfdb59b5bda1fc245aadae15b1984a-Paper.pdf
2019
Cited alongside, same era.
A. Malysheva, D. Kudenko, and A. Shpilman, “Magnet: Multi-agent graph network for deep multi-agent reinforcement learning,” in 2019 XVI International Symposium "Problems of Redundancy in Information and Control Systems" (REDUNDANCY) , 2019, pp. 171–176
2019
Cited alongside, same era.
J. Z. Leibo, E. Duéñez-Guzmán, A. S. Vezhnevets, J. P. Agapiou, P. Sunehag, R. Koster, J. Matyas, C. Beattie, I. Mordatch, and T. Graepel, “Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting Pot,” 2021
2021
Later among the works it cites.
J. Chen, Y. Zhang, Y. Xu, H. Ma, H. Yang, J. Song, Y. Wang, and Y. Wu, “Variational automatic curriculum learning for sparse-reward cooperative multi-agent problems,” in Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. Curran Associates, Inc., 2021, pp. 9681–9693. [Online]. Available: https://proceedings.neurips.cc/paper/2021/file/503e7dbbd6217b9a591f3322f39b5a6c-Paper.pdf
2021
Later among the works it cites.
M. Chen, Y. Li, E. Wang, Z. Yang, Z. Wang, and T. Zhao, “Pessimism meets invariance: Provably efficient offline mean-field multi-agent rl,” in Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. Curran Associates, Inc., 2021, pp. 17 913–17 926. [Online]. Available: https://proceedings.neurips.cc/paper/2021/file/9559fc73b13fa721a816958488a5b449-Paper.pdf
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Mao, Z. Zhang, Z. Xiao, and Z. Gong, “Modelling the dynamic joint policy of teammates with attention multi-agent DDPG,” in AAMAS . International Foundation for Autonomous Agents and Multiagent Systems, 2019, pp. 1108–1116
2019
Cited alongside, same era.
M. Samvelyan, T. Rashid, C. S. de Witt, G. Farquhar, N. Nardelli, T. G. J. Rudner, C.-M. Hung, P. H. S. Torr, J. N. Foerster, and S. Whiteson, “The starcraft multi-agent challenge,” in AAMAS , 2019, pp. 2186–2188. [Online]. Available: http://dl.acm.org/citation.cfm?id=3332052
2019
Cited alongside, same era.
N. Jaques, A. Lazaridou, E. Hughes, C. Gulcehre, P. Ortega, D. Strouse, J. Z. Leibo, and N. De Freitas, “Social influence as intrinsic motivation for multi-agent deep reinforcement learning,” in Proceedings of the 36th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97. PMLR, 09–15 Jun 2019, pp. 3040–3049. [Online]. Available: https://proceedings.mlr.press/v97/jaques19a.html
2019
Cited alongside, same era.
Y. Du, L. Han, M. Fang, J. Liu, T. Dai, and D. Tao, “Liir: Learning individual intrinsic reward in multi-agent reinforcement learning,” in Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. Alché-Buc, E. Fox, and R. Garnett, Eds., vol. 32. Curran Associates, Inc., 2019. [Online]. Available: https://proceedings.neurips.cc/paper/2019/file/07a9d3fed4c5ea6b17e80258dee231fa-Paper.pdf
2019
Cited alongside, same era.
A. Mahajan, T. Rashid, M. Samvelyan, and S. Whiteson, “Maven: Multi-agent variational exploration,” in Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. Alché-Buc, E. Fox, and R. Garnett, Eds., vol. 32. Curran Associates, Inc., 2019. [Online]. Available: https://proceedings.neurips.cc/paper/2019/file/f816dc0acface7498e10496222e9db10-Paper.pdf
2019
Cited alongside, same era.
C. Schroeder de Witt, J. Foerster, G. Farquhar, P. Torr, W. Boehmer, and S. Whiteson, “Multi-agent common knowledge reinforcement learning,” in Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. Alché-Buc, E. Fox, and R. Garnett, Eds., vol. 32. Curran Associates, Inc., 2019. [Online]. Available: https://proceedings.neurips.cc/paper/2019/file/f968fdc88852a4a3a27a81fe3f57bfc5-Paper.pdf
2019
Cited alongside, same era.
N. Carion, N. Usunier, G. Synnaeve, and A. Lazaric, “A structured prediction approach for generalization in cooperative multi-agent reinforcement learning,” in Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. Alché-Buc, E. Fox, and R. Garnett, Eds., vol. 32. Curran Associates, Inc., 2019. [Online]. Available: https://proceedings.neurips.cc/paper/2019/file/3c3c139bd8467c1587a41081ad78045e-Paper.pdf
2019
Cited alongside, same era.
A. Das, T. Gervet, J. Romoff, D. Batra, D. Parikh, M. Rabbat, and J. Pineau, “TarMAC: Targeted multi-agent communication,” in Proceedings of the 36th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97. PMLR, 09–15 Jun 2019, pp. 1538–1546. [Online]. Available: https://proceedings.mlr.press/v97/das19a.html
2019
Cited alongside, same era.
K. Son, D. Kim, W. J. Kang, D. E. Hostallero, and Y. Yi, “QTRAN: Learning to factorize with transformation for cooperative multi-agent reinforcement learning,” in Proceedings of the 36th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97. PMLR, 09–15 Jun 2019, pp. 5887–5896. [Online]. Available: https://proceedings.mlr.press/v97/son19a.html
2019
Cited alongside, same era.
2021
Later among the works it cites.
S. Li, J. K. Gupta, P. Morales, R. Allen, and M. J. Kochenderfer, “Deep implicit coordination graphs for multi-agent reinforcement learning,” in Proceedings of the 20th International Conference on Autonomous Agents and MultiAgent Systems , ser. AAMAS ’21. Richland, SC: International Foundation for Autonomous Agents and Multiagent Systems, 2021, p. 764–772
2021
Later among the works it cites.
W.-F. Sun, C.-K. Lee, and C.-Y. Lee, “Dfac framework: Factorizing the value function via quantile mixture for multi-agent distributional q-learning,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 9945–9954. [Online]. Available: https://proceedings.mlr.press/v139/sun21c.html
2021
Later among the works it cites.
J. Wang, Z. Ren, B. Han, J. Ye, and C. Zhang, “Towards understanding cooperative multi-agent q-learning with value factorization,” in Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. Curran Associates, Inc., 2021, pp. 29 142–29 155. [Online]. Available: https://proceedings.neurips.cc/paper/2021/file/f3f1fa1e4348bfbebdeee8c80a04c3b9-Paper.pdf
2021
Later among the works it cites.
K. M. Lee, S. G. Subramanian, and M. Crowley, “Investigation of independent reinforcement learning algorithms in multi-agent environments,” in Deep RL Workshop NeurIPS 2021 , 2021. [Online]. Available: https://openreview.net/forum?id=8MkKGZ2AlmJ
2021
Later among the works it cites.
L. Chenghao, T. Wang, C. Wu, Q. Zhao, J. Yang, and C. Zhang, “Celebrating diversity in shared multi-agent reinforcement learning,” in Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. Curran Associates, Inc., 2021, pp. 3991–4002. [Online]. Available: https://proceedings.neurips.cc/paper/2021/file/20aee3a5f4643755a79ee5f6a73050ac-Paper.pdf
2021
Later among the works it cites.
T. Wang, T. Gupta, A. Mahajan, B. Peng, S. Whiteson, and C. Zhang, “Rode: Learning roles to decompose multi-agent tasks,” in ICLR , 2021. [Online]. Available: https://openreview.net/forum?id=TTUVg6vkNjK
2021
Later among the works it cites.
Y. Xiao, X. Lyu, and C. Amato, “Local advantage actor-critic for robust multi-agent deep reinforcement learning,” in MRS . IEEE, 2021, pp. 155–163
2021
Later among the works it cites.
Z. Xu, D. Li, Y. Bai, and G. Fan, “MMD-MIX: value function factorisation with maximum mean discrepancy for cooperative multi-agent reinforcement learning,” in International Joint Conference on Neural Networks, IJCNN 2021, Shenzhen, China, July 18-22, 2021 . IEEE, 2021, pp. 1–7. [Online]. Available: https://doi.org/10.1109/IJCNN52387.2021.9533636
2021
Later among the works it cites.
J. Jiang and Z. Lu, “The emergence of individuality,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 4992–5001. [Online]. Available: https://proceedings.mlr.press/v139/jiang21g.html
2021
Later among the works it cites.
T. Rashid, G. Farquhar, B. Peng, and S. Whiteson, “Weighted qmix: Expanding monotonic value function factorisation for deep multi-agent reinforcement learning.” NeurIPS, 2021
2021
Later among the works it cites.
J. Su, S. C. Adams, and P. A. Beling, “Value-decomposition multi-agent actor-critics,” in Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021 . AAAI Press, 2021, pp. 11 352–11 360. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/17353
2021
Later among the works it cites.
L. Pan, T. Rashid, B. Peng, L. Huang, and S. Whiteson, “Regularized softmax deep multi-agent q-learning,” in Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. Curran Associates, Inc., 2021, pp. 1365–1377. [Online]. Available: https://proceedings.neurips.cc/paper/2021/file/0a113ef6b61820daa5611c870ed8d5ee-Paper.pdf
2021
Later among the works it cites.
I.-J. Liu, U. Jain, R. A. Yeh, and A. Schwing, “Cooperative exploration for multi-agent deep reinforcement learning,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 6826–6836. [Online]. Available: https://proceedings.mlr.press/v139/liu21j.html
2021
Later among the works it cites.
I. Saeed, A. C. Cullen, S. M. Erfani, and T. Alpcan, “Domain-aware multiagent reinforcement learning in navigation,” in International Joint Conference on Neural Networks, IJCNN 2021, Shenzhen, China, July 18-22, 2021 . IEEE, 2021, pp. 1–8. [Online]. Available: https://doi.org/10.1109/IJCNN52387.2021.9533975
2021
Later among the works it cites.
2021
Later among the works it cites.
L. Zheng, J. Chen, J. Wang, J. He, Y. Hu, Y. Chen, C. Fan, Y. Gao, and C. Zhang, “Episodic multi-agent reinforcement learning with curiosity-driven exploration,” in Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. Curran Associates, Inc., 2021, pp. 3757–3769. [Online]. Available: https://proceedings.neurips.cc/paper/2021/file/1e8ca836c962598551882e689265c1c5-Paper.pdf
2021
Later among the works it cites.
E. Marchesini and A. Farinelli, “Centralizing state-values in dueling networks for multi-robot reinforcement learning mapless navigation,” in IEEE/RSJ International Conference on Intelligent Robots and Systems, IROS 2021, Prague, Czech Republic, September 27 - Oct. 1, 2021 . IEEE, 2021, pp. 4583–4588. [Online]. Available: https://doi.org/10.1109/IROS51168.2021.9636349
2021
Later among the works it cites.
J. Wang, Z. Ren, T. Liu, Y. Yu, and C. Zhang, “{QPLEX}: Duplex dueling multi-agent q-learning,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=Rcmk0xxIQV
2021
Later among the works it cites.
J. G. Kuba, M. Wen, L. Meng, s. gu, H. Zhang, D. Mguni, J. Wang, and Y. Yang, “Settling the variance of multi-agent policy gradients,” in Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. Curran Associates, Inc., 2021, pp. 13 458–13 470. [Online]. Available: https://proceedings.neurips.cc/paper/2021/file/6fe6a8a6e6cb710584efc4af0c34ce50-Paper.pdf
2021
Later among the works it cites.
B. Peng, T. Rashid, C. Schroeder de Witt, P.-A. Kamienny, P. Torr, W. Boehmer, and S. Whiteson, “Facmac: Factored multi-agent centralised policy gradients,” in Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. Curran Associates, Inc., 2021, pp. 12 208–12 221. [Online]. Available: https://proceedings.neurips.cc/paper/2021/file/65b9eea6e1cc6bb9f0cd2a47751a186f-Paper.pdf
2021
Later among the works it cites.
2021
Later among the works it cites.
R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. Bellemare, “Deep reinforcement learning at the edge of the statistical precipice,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
R. Agarwal, M. Schwarzer, P. S. Castro, A. Courville, and M. G. Bellemare, “Deep Reinforcement Learning at the Edge of the Statistical Precipice,” 2022
2022
Closest in time.
J. Hu, S. Jiang, S. A. Harding, H. Wu, and S. wei Liao, “Rethinking the implementation tricks and monotonicity constraint in cooperative multi-agent reinforcement learning,” 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
L. Yuan, J. Wang, F. Zhang, C. Wang, Z. Zhang, Y. Yu, and C. Zhang, “Multi-agent incentive communication via decentralized teammate modeling,” 2022
2022
Closest in time.
D. H. Mguni, T. Jafferjee, J. Wang, N. Perez-Nieves, O. Slumbers, F. Tong, Y. Li, J. Zhu, Y. Yang, and J. Wang, “LIGS: Learnable intrinsic-reward generation selection for multi-agent learning,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=CpTuR2ECuW
2022
Closest in time.
Y. Wang, fangwei zhong, J. Xu, and Y. Wang, “Tom2c: Target-oriented multi-agent communication and cooperation with theory of mind,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=2t7CkQXNpuq
2022
Closest in time.
J. G. Kuba, R. Chen, M. Wen, Y. Wen, F. Sun, J. Wang, and Y. Yang, “Trust region policy optimisation in multi-agent reinforcement learning,” in International Conference on Learning Representations , 2022. [Online]. Available: https://openreview.net/forum?id=EcGGFkNTxdJ
2022
Closest in time.
S. A. Stavroulakis and B. Sengupta, “Reinforcement learning for location-aware warehouse scheduling,” in ICLR 2022 Workshop on Generalizable Policy Learning in Physical World , 2022. [Online]. Available: https://openreview.net/forum?id=Bt-gaVaVJ-9
2022
Closest in time.
A. Castagna and I. Dusparic, “Multi-agent transfer learning in reinforcement learning-based ride-sharing systems,” in Proceedings of the 14th International Conference on Agents and Artificial Intelligence, ICAART 2022, Volume 2, Online Streaming, February 3-5, 2022 , A. P. Rocha, L. Steels, and H. J. van den Herik, Eds. SCITEPRESS, 2022, pp. 120–130. [Online]. Available: https://doi.org/10.5220/0010785200003116
2022
Closest in time.
M. Zawalski, B. Osinski, H. Michalewski, and P. Milos, “Off-policy correction for multi-agent reinforcement learning,” in AAMAS . International Foundation for Autonomous Agents and Multiagent Systems (IFAAMAS), 2022, pp. 1774–1776
2022
Closest in time.
R. Avalos, M. Reymond, A. Nowé, and D. M. Roijers, “Local advantage networks for cooperative multi-agent reinforcement learning,” in AAMAS . International Foundation for Autonomous Agents and Multiagent Systems (IFAAMAS), 2022, pp. 1524–1526
2022
Closest in time.
Y. X. Xueguang Lyu, “A deeper understanding of state-based critics in multi-agent reinforcement learning,” Proceedings of the AAAI Conference on Artificial Intelligence , 2022. [Online]. Available: https://par.nsf.gov/biblio/10315765
2022
Closest in time.
K. Cobbe, C. Hesse, J. Hilton, and J. Schulman, “Leveraging procedural generation to benchmark reinforcement learning,” in International conference on machine learning . PMLR, 2020, pp. 2048–2056
2056
Closest in time.
P. Sunehag, G. Lever, A. Gruslys, W. M. Czarnecki, V. F. Zambaldi, M. Jaderberg, M. Lanctot, N. Sonnerat, J. Z. Leibo, K. Tuyls, and T. Graepel, “Value-decomposition networks for cooperative multi-agent learning based on team reward,” in Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems, AAMAS 2018, Stockholm, Sweden, July 10-15, 2018 , E. André, S. Koenig, M. Dastani, and G. Sukthankar, Eds. International Foundation for Autonomous Agents and Multiagent Systems Richland, SC, USA / ACM, 2018, pp. 2085–2087. [Online]. Available: http://dl.acm.org/citation.cfm?id=3238080
2087
Closest in time.