Fetching the paper…
Reading the bibliography…
Multi-agent football poses an unsolved challenge in AI research.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Robocup: The robot world cup initiative. In Proceedings of the first international conference on Autonomous agents . 340–347
Hiroaki Kitano, Minoru Asada, Yasuo Kuniyoshi, Itsuki Noda, and Eiichi Osawa. 1997 · 1997
Earlier work this paper cites.
Actor-critic algorithms
Vijay Konda and John Tsitsiklis. 1999 · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping. In Icml , Vol. 99. 278–287
Andrew Y Ng, Daishi Harada, and Stuart Russell. 1999 · 1999
Earlier work this paper cites.
The complexity of decentralized control of Markov decision processes
Daniel S Bernstein, Robert Givan, Neil Immerman, and Shlomo Zilberstein. 2002 · 2002
Earlier work this paper cites.
Reinforcement learning for robocup soccer keepaway
Peter Stone, Richard S Sutton, and Gregory Kuhlmann. 2005 · 2005
Earlier work this paper cites.
Trueskill™: A Bayesian skill rating system. In Proceedings of the 19th international conference on neural information processing systems . 569–576
Ralf Herbrich, Tom Minka, and Thore Graepel. 2006 · 2006
Earlier work this paper cites.
Half field offense in RoboCup soccer: A multiagent reinforcement learning case study. In RoboCup 2006: Robot Soccer World Cup X 10 . Springer, 72–85
Shivaram Kalyanakrishnan, Yaxin Liu, and Peter Stone. 2007 · 2006
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Fictitious self-play in extensive-form games. In International conference on machine learning . PMLR, 805–813
Johannes Heinrich, Marc Lanctot, and David Silver. 2015 · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. 2015 · 2015
Earlier work this paper cites.
Vizdoom: A doom-based ai research platform for visual reinforcement learning. In 2016 IEEE Conference on Computational Intelligence and Games (CIG) . IEEE, 1–8
Michał Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Jaśkowski. 2016 · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Training agent for first-person shooter game with actor-critic curriculum learning
Yuxin Wu and Yuandong Tian. 2016 · 2016
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolat, David Silver, and Thore Graepel. 2017 · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi I Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2017
Earlier work this paper cites.
Value-decomposition networks for cooperative multi-agent learning
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2017
Earlier work this paper cites.
Counterfactual multi-agent policy gradients. In Proceedings of the AAAI conference on artificial intelligence , Vol. 32
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. 2018 · 2018
Earlier work this paper cites.
Recurrent experience replay in distributed reinforcement learning. In International conference on learning representations
Steven Kapturowski, Georg Ostrovski, John Quan, Remi Munos, and Will Dabney. 2018 · 2018
Earlier work this paper cites.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning. In International Conference on Machine Learning . PMLR, 4295–4304
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. 2018 · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Cited alongside, same era.
Open-ended learning in symmetric zero-sum games. In International Conference on Machine Learning . PMLR, 434–443
David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech Czarnecki, Julien Perolat, Max Jaderberg, and Thore Graepel. 2019 · 2019
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemyslaw Debiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Cited alongside, same era.
First return, then explore
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth O Stanley, and Jeff Clune. 2021 · 2021
Later among the works it cites.
Learning Diverse Policies in MOBA Games via Macro-Goals
Yiming Gao, Bei Shi, Xueying Du, Liang Wang, Guangwei Chen, Zhenjie Lian, Fuhao Qiu, Guoan Han, Weixuan Wang, Deheng Ye, et al · 2021
Later among the works it cites.
TiKick: Towards Playing Multi-agent Football Full Games from Single-agent Demonstrations
Shiyu Huang, Wenze Chen, Longfei Zhang, Ziyang Li, Fengming Zhu, Deheng Ye, Ting Chen, and Jun Zhu. 2021 · 2021
Later among the works it cites.
Celebrating Diversity in Shared Multi-Agent Reinforcement Learning
Chenghao Li, Chengjie Wu, Tonghan Wang, Jun Yang, Qianchuan Zhao, and Chongjie Zhang. 2021 · 2021
Later among the works it cites.
From Motor Control to Team Play in Simulated Humanoid Football
Siqi Liu, Guy Lever, Zhe Wang, Josh Merel, SM Eslami, Daniel Hennes, Wojciech M Czarnecki, Yuval Tassa, Shayegan Omidshafiei, Abbas Abdolmaleki, et al · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Combo-Action: Training Agent For FPS Game with Auxiliary Tasks. AAAI
Shiyu Huang, Hang Su, Jun Zhu, and Ting Chen. 2019 · 2019
Cited alongside, same era.
Google research football: A novel reinforcement learning environment
Karol Kurach, Anton Raichuk, Piotr Stanczyk, Michal Zajkac, Olivier Bachem, Lasse Espeholt, Carlos Riquelme, Damien Vincent, Marcin Michalski, Olivier Bousquet, et al · 2019
Cited alongside, same era.
Emergent coordination through competition
Siqi Liu, Guy Lever, Josh Merel, Saran Tunyasuvunakool, Nicolas Heess, and Thore Graepel. 2019 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning. In International Conference on Machine Learning . PMLR, 5887–5896
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi. 2019 · 2019
Cited alongside, same era.
On bonus based exploration methods in the arcade learning environment. In International Conference on Learning Representations
Adrien Ali Taiga, William Fedus, Marlos C Machado, Aaron Courville, and Marc G Bellemare. 2019 · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Solution ranked 35th in Kaggle Football Competition
Sarvar Anvarov. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Unifying behavioral and response diversity for open-ended learning in zero-sum games
Xiangyu Liu, Hangtian Jia, Ying Wen, Yaodong Yang, Yujing Hu, Yingfeng Chen, Changjie Fan, and Zhipeng Hu. 2021a · 2021
Later among the works it cites.
rSoccer: A Framework for Studying Reinforcement Learning in Small and Very Small Size Robot Soccer
Felipe B. Martins, Mateus G. Machado, Hansenclever F. Bassani, Pedro H. M. Braga, and Edna S. Barros. 2021 · 2021
Later among the works it cites.
The Surprising Effectiveness of MAPPO in Cooperative, Multi-Agent Games
Chao Yu, Akash Velu, Eugene Vinitsky, Yu Wang, Alexandre Bayen, and Yi Wu. 2021 · 2021
Later among the works it cites.
Douzero: Mastering doudizhu with self-play deep reinforcement learning. In International Conference on Machine Learning . PMLR, 12333–12344
Daochen Zha, Jingru Xie, Wenye Ma, Sheng Zhang, Xiangru Lian, Xia Hu, and Ji Liu. 2021 · 2021
Later among the works it cites.
DGPO: Discovering Multiple Strategies with Diversity-Guided Policy Optimization
Wenze Chen, Shiyu Huang, Yuan Chiang, Ting Chen, and Jun Zhu. 2022 · 2022
Later among the works it cites.
Multi-Level Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
Lei Feng, Yuxuan Xie, Bing Liu, and Shuyan Wang. 2022 · 2022
Later among the works it cites.
Revisiting some common practices in cooperative multi-agent reinforcement learning
Wei Fu, Chao Yu, Zelai Xu, Jiaqi Yang, and Yi Wu. 2022 · 2022
Later among the works it cites.
PerfectDou: Dominating DouDizhu with Perfect Information Distillation
Yang Guan, Minghuan Liu, Weijun Hong, Weinan Zhang, Fei Fang, Guangjun Zeng, and Yue Lin. 2022 · 2022
Later among the works it cites.
JiDi Olympics Football
JiDi. 2022 · 2022
Later among the works it cites.
Steven Kapturowski, Víctor Campos, Ray Jiang, Nemanja Rakićević, Hado van Hasselt, Charles Blundell, and Adrià Puigdomènech Badia. 2022 · 2022
Later among the works it cites.
NeuPL: Neural Population Learning
Siqi Liu, Luke Marris, Daniel Hennes, Josh Merel, Nicolas Heess, and Thore Graepel. 2022 · 2022
Later among the works it cites.
Counter-Strike Deathmatch with Large-Scale Behavioural Cloning
Tim Pearce and Jun Zhu. 2022 · 2022
Later among the works it cites.
A novel optimization perspective to the problem of designing sequences of tasks in a reinforcement learning framework
Ruggiero Seccia, Francesco Foglino, Matteo Leonetti, and Simone Sagratella. 2022 · 2022
Later among the works it cites.
Monotonic Improvement Guarantees under Non-stationarity for Decentralized PPO
Mingfei Sun, Sam Devlin, Jacob Beck, Katja Hofmann, and Shimon Whiteson. 2022 · 2022
Later among the works it cites.
Individual Reward Assisted Multi-Agent Reinforcement Learning. In International Conference on Machine Learning . PMLR, 23417–23432
Li Wang, Yupeng Zhang, Yujing Hu, Weixun Wang, Chongjie Zhang, Yang Gao, Jianye Hao, Tangjie Lv, and Changjie Fan. 2022 · 2022
Later among the works it cites.
Multi-Agent Reinforcement Learning is a Sequence Modeling Problem
Muning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang, Ying Wen, Jun Wang, and Yaodong Yang. 2022 · 2022
Later among the works it cites.
From Local to Global: A Curriculum Learning Approach for Reinforcement Learning-based Traffic Signal Control. In 2022 IEEE 2nd International Conference on Software Engineering and Artificial Intelligence (SEAI) . IEEE, 253–258
Nianzhao Zheng, Jialong Li, Zhenyu Mao, and Kenji Tei. 2022 · 2022
Later among the works it cites.
Leveraging procedural generation to benchmark reinforcement learning. In International conference on machine learning . PMLR, 2048–2056
Karl Cobbe, Chris Hesse, Jacob Hilton, and John Schulman. 2020 · 2056
Closest in time.