Fetching the paper…
Reading the bibliography…
Competitive Self-Play (CSP) based Multi-Agent Reinforcement Learning (MARL) has shown phenomenal breakthroughs recently.
An iterative method of solving a game
Julia Robinson · 1951
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Using MPI-2: Advanced Features of the Message-Passing Interface
William Gropp, Ewing Lusk, and Rajeev Thakur · 1999
Earlier work this paper cites.
Graphical models
Michael I Jordan et al · 2004
Earlier work this paper cites.
A concise introduction to multiagent systems and distributed artificial intelligence
Nikos Vlassis · 2007
Earlier work this paper cites.
Robust strategies and counter-strategies: Building a champion level computer poker player
Michael Bradley Johanson · 2007
Earlier work this paper cites.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione · 2008
Earlier work this paper cites.
A unified view of large-scale zero-sum equilibrium computation
Kevin Waugh and J Andrew Bagnell · 2014
Earlier work this paper cites.
Designing collective behavior in a termite-inspired robot construction team
Werfel, Justin, Petersen, Kirstin, Nagpal, and Radhika · 2014
Earlier work this paper cites.
Programmable self-assembly in a thousand-robot swarm
Michael Rubenstein, Alejandro Cornejo, and Radhika Nagpal · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Massively parallel methods for deep reinforcement learning
Arun Nair, Praveen Srinivasan, Sam Blackwell, Cagdas Alcicek, Rory Fearon, Alessandro De Maria, Vedavyas Panneershelvam, Mustafa Suleyman, Charles Beattie, Stig Petersen, et al · 2015
Earlier work this paper cites.
Fictitious self-play in extensive-form games
Johannes Heinrich, Marc Lanctot, and David Silver · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Vizdoom: A doom-based ai research platform for visual reinforcement learning
Michał Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Jaśkowski · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Reinforcement learning through asynchronous advantage actor-critic on a gpu
Mohammad Babaeizadeh, Iuri Frosio, Stephen Tyree, Jason Clemons, and Jan Kautz · 2016
Earlier work this paper cites.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Microservices architecture enables devops: Migration to a cloud-native architecture
Armin Balalaie, Abbas Heydarnoori, and Pooyan Jamshidi · 2016
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2017
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Julien Perolat, David Silver, Thore Graepel, et al · 2017
Earlier work this paper cites.
Single sample fictitious play
Brian Swenson, Soummya Kar, and Joao Xavier · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Efficient parallel methods for deep reinforcement learning
Alfredo V Clemente, Humberto N Castejón, and Arjun Chandra · 2017
Cited alongside, same era.
Elf: An extensive, lightweight and flexible research platform for real-time strategy games
Yuandong Tian, Qucheng Gong, Wenling Shang, Yuxin Wu, and C. Lawrence Zitnick · 2017
Cited alongside, same era.
Openai baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov · 2017
Cited alongside, same era.
Seed rl: Scalable and efficient deep-rl with accelerated central inference
Lasse Espeholt, Raphaël Marinier, Piotr Stanczyk, Ke Wang, and Marcin Michalski · 2019
Later among the works it cites.
Deep counterfactual regret minimization
Noam Brown, Adam Lerer, Sam Gross, and Tuomas Sandholm · 2019
Later among the works it cites.
Arena: a toolkit for multi-agent reinforcement learning
Qing Wang, Jiechao Xiong, Lei Han, Meng Fang, Xinghai Sun, Zhuobin Zheng, Peng Sun, and Zhengyou Zhang · 2019
Later among the works it cites.
Multi-agent deep reinforcement learning for liquidation strategy analysis
Wenhang Bao and Xiao-yang Liu · 2019
Later among the works it cites.
Mastering complex control in moba games with deep reinforcement learning
Deheng Ye, Zhao Liu, Mingfei Sun, Bei Shi, Peilin Zhao, Hao Wu, Hongsheng Yu, Shaojie Yang, Xipeng Wu, Qingwei Guo, et al · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nccl 2.0
Sylvain Jeaugey · 2017
Cited alongside, same era.
Kubernetes in action
Marko Luka · 2017
Cited alongside, same era.
Training agent for first-person shooter game with actor-critic curriculum learning
Yuxin Wu and Yuandong Tian · 2017
Cited alongside, same era.
Sampled fictitious play is hannan consistent
Zifan Li and Ambuj Tewari · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Cited alongside, same era.
Pommerman: A multi-agent playground
Cinjon Resnick, Wes Eldridge, David Ha, Denny Britz, Jakob Foerster, Julian Togelius, Kyunghyun Cho, and Joan Bruna · 2018
Cited alongside, same era.
Distributed prioritized experience replay
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado Van Hasselt, and David Silver · 2018
Cited alongside, same era.
Closest in time.
Reverb: An efficient data storage and transport system for ml research
Albin Cassirer, Gabriel Barth-Maron, Thibault Sottiaux, Manuel Kroiss, Eugene Brevdo · 2020
Closest in time.
Acme: A research framework for distributed reinforcement learning
Matt Hoffman, Bobak Shahriari, John Aslanides, Gabriel Barth-Maron, Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, et al · 2020
Closest in time.
Arena: A general evaluation platform and building toolkit for multi-agent intelligence
Yuhang Song, Andrzej Wojcicki, Thomas Lukasiewicz, Jianyi Wang, Abi Aryan, Zhenghua Xu, Mai Xu, Zihan Ding, and Lianlong Wu · 2020
Closest in time.
https://zguide.zeromq.org/
Ømq - the guide · 2020
Closest in time.
https://developers.google.com/protocol-buffers
Protocol buffers are a language-neutral, platform-neutral extensible mechanism for serializing structured data · 2020
Closest in time.
https://grpc.io/
A high-performance, open source universal rpc framework · 2020
Closest in time.
https://www.fullstackpython.com/jinja2.html
Jinja2 · 2020
Closest in time.
https://www.kubeflow.org/
kubeflow · 2020
Closest in time.
https://cloud.tencent.com/
Tencent cloud · 2020
Closest in time.
https://intl.cloud.tencent.com/product/tke
Tencent kubernetes engine · 2020
Closest in time.
https://intl.cloud.tencent.com/product/cvm
Tencent cloud cvm · 2020
Closest in time.
https://intl.cloud.tencent.com/product/cfs
Cloud file storage · 2020
Closest in time.
Lei Han, Jiechao Xiong, Peng Sun, Xinghai Sun, Meng Fang, Qingwei Guo, Qiaobo Chen, Tengfei Shi, and Zhengyou Zhang · 2020
Closest in time.
http://vizdoom.cs.put.edu.pl/competitions/vdaic-2016-cig
Vizdoom cig 2016 competition · 2020
Closest in time.
https://github.com/mihahauke/VDAIC2017
ViZDoom testing code · 2020
Closest in time.
https://github.com/mwydmuch/ViZDoom/blob/master/doc/Types.md#gamestate
ViZDoom doc · 2020
Closest in time.
https://en.wikipedia.org/wiki/Command:_Modern_Air_Naval_Operations
Command: Modern air naval operations · 2020
Closest in time.
https://github.com/microsoft/AirSim
Welcome to airsim · 2020
Closest in time.
https://en.wikipedia.org/wiki/Drone_racing
Drone racing · 2020
Closest in time.