Fetching the paper…
Reading the bibliography…
To achieve general intelligence, agents must learn how to interact with others in a shared environment: this is the challenge of multiagent reinforcement learning (MARL).
Iterative solutions of games by fictitious play
G. W. Brown · 1951
Earlier work this paper cites.
Extensive games and the problem of information
H. W. Kuhn · 1953
Earlier work this paper cites.
Evolutionarily stable strategies and game dynamics
Taylor and Jonker · 1978
Earlier work this paper cites.
Fast algorithms for finding randomized strategies in game trees
D. Koller, N. Megiddo, and B. von Stengel · 1994
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
Gambling in a rigged casino: The adversarial multi-armed bandit problem
P. Auer, N. Cesa-Bianchi, Y. Freund, and R. E. Schapire · 1995
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
C. Claus and C. Boutilier · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. Sutton and A. Barto · 1998
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
S. Hart and A. Mas-Colell · 2000
Earlier work this paper cites.
A reinforcement procedure leading to correlated equilibrium
Sergiu Hart and Andreu Mas-Colell · 2001
Earlier work this paper cites.
Friend-or-foe Q-learning in general-sum games
Michael L. Littman · 2001
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Michael Bowling and Manuela Veloso · 2002
Earlier work this paper cites.
On the global convergence of stochastic fictitious play
Josef Hofbauer and William H. Sandholm · 2002
Earlier work this paper cites.
Analyzing complex strategic interactions in multi-agent games
W. E. Walsh, R. Das, G. Tesauro, and J.O. Kephart · 2002
Earlier work this paper cites.
Correlated Q-learning
Amy Greenwald and Keith Hall · 2003
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
H.B. McMahan, G. Gordon, and A. Blum · 2003
Earlier work this paper cites.
A cognitive hierarchy model of games
Colin F. Camerer, Teck-Hua Ho, and Juin-Kuan Chong · 2004
Earlier work this paper cites.
Reinforcement learning for stochastic cooperative multi-agent systems
M. Lauer and M. Riedmiller · 2004
Earlier work this paper cites.
Coordinating multiagent teams in uncertain domains using distributed POMDPs
Ranjit Nair · 2004
Earlier work this paper cites.
A framework for sequential planning in multiagent settings
Gmytrasiewicz and Doshi · 2005
Earlier work this paper cites.
Cognition and behavior in two-person guessing games: An experimental study
M. A. Costa-Gomes and V. P. Crawford · 2006
Earlier work this paper cites.
Generalised weakened fictitious play
David S. Leslie and Edmund J. Collins · 2006
Earlier work this paper cites.
The parallel Nash memory for asymmetric games
F.A. Oliehoek, E.D. de Jong, and N. Vlassis · 2006
Earlier work this paper cites.
Methods for empirical game-theoretic analysis
Michael P. Wellman · 2006
Earlier work this paper cites.
Learning, regret minimization, and equilibria
A. Blum and Y. Mansour · 2007
Earlier work this paper cites.
A new algorithm for generating equilibria in massive zero-sum games
N. Burch M. Zinkevich, M. Bowling · 2007
Earlier work this paper cites.
Adversarial planning through strategy simulation
F. Sailer, M. Buro, and M. Lanctot · 2007
Earlier work this paper cites.
If multi-agent learning is the answer, what is the question?
Yoav Shoham, Rob Powers, and Trond Grenager · 2007
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
L. Busoniu, R. Babuska, and B. De Schutter · 2008
Earlier work this paper cites.
Computing robust counter-strategies
M. Johanson, M. Zinkevich, and M. Bowling · 2008
Earlier work this paper cites.
Not all agents are equal: Scaling up distributed pomdps for agent networks
Janusz Marecki, Tapana Gupta, Pradeep Varakantham, Milind Tambe, and Makoto Yokoo · 2008
Earlier work this paper cites.
Theoretical advantages of lenient learners: An evolutionary game theoretic perspective
Liviu Panait, Karl Tuyls, and Sean Luke · 2008
Earlier work this paper cites.
Replicator dynamics in discrete and continuous strategy spaces
K. Tuyls and R. Westra · 2008
Earlier work this paper cites.
Game theory of mind
Wako Yoshida, Ray J. Dolan, and Karl J. Friston · 2008
Earlier work this paper cites.
Regret minimization in games with incomplete information
M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione · 2008
Earlier work this paper cites.
Algorithms for Abstracting and Solving Imperfect Information Games
A. Gilpin · 2009
Earlier work this paper cites.
An evolutionary game theoretic analysis of poker strategies
Marc Ponsen, Karl Tuyls, Michael Kaisers, and Jan Ramon · 2009
Cited alongside, same era.
Stronger CDA strategies through empirical game-theoretic analysis and reinforcement learning
L. Julian Schvartzman and Michael P. Wellman · 2009
Cited alongside, same era.
Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations
Y. Shoham and K. Leyton-Brown · 2009
Cited alongside, same era.
Frequency adjusted multi-agent Q-learning
Michael Kaisers and Karl Tuyls · 2010
Cited alongside, same era.
Understanding the success of perfect information Monte Carlo sampling in game tree search
J. Long, N. R. Sturtevant, M. Buro, and T. Furtak · 2010
Cited alongside, same era.
Statsmodels: Econometric and statistical modeling with python
Skipper Seabold and Josef Perktold · 2010
Algorithms for computing strategies in two-player simultaneous move games
Branislav Bošanský, Viliam Lisý, Marc Lanctot, Jiří Čermák, and Mark H.M. Winands · 2016
Later among the works it cites.
Learning to communicate with deep multi-agent reinforcement learning
Jakob N. Foerster, Yannis M. Assael, Nando de Freitas, and Shimon Whiteson · 2016
Later among the works it cites.
Deep learning for predicting human strategic behavior
Jason S. Hartford, James R. Wright, and Kevin Leyton-Brown · 2016
Later among the works it cites.
Deep reinforcement learning in parameterized action space
Matthew Hausknecht and Peter Stone · 2016
Later among the works it cites.
Cooperation and communication in multiagent deep reinforcement learning
Matthew John Hausknecht · 2016
Later among the works it cites.
Opponent modeling in deep reinforcement learning
He He, Jordan Boyd-Graber, Kevin Kwok, , and Hal Daumé III · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Beyond equilibrium: Predicting human behavior in normal form games
James R. Wright and Kevin Leyton-Brown · 2010
Cited alongside, same era.
Bayesian theory of mind: Modeling joint belief-desire attribution
C.L. Baker, R.R. Saxe, and J.B. Tenenbaum · 2011
Cited alongside, same era.
Accelerating best response calculation in large extensive games
Michael Johanson, Michael Bowling, Kevin Waugh, and Martin Zinkevich · 2011
Cited alongside, same era.
The world of independent learners is not Markovian
Guillaume J. Laurent, Laëtitia Matignon, and Nadine Le Fort-Piat · 2011
Cited alongside, same era.
Protecting against evaluation overfitting in empirical reinforcement learning
S. Whiteson, B. Tanner, M. E. Taylor, and P. Stone · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
Later among the works it cites.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z. Leibo, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Robust Strategies and Counter-Strategies: From Superhuman to Optimal Play
Michael Bradley Johanson · 2016
Later among the works it cites.
Coordinate to cooperate or compete: abstract goals and joint intentions in social interaction
M. Kleiman-Weiner, M. K. Ho, J. L. Austerweil, M. L. Littman, and J. B. Tenenbaum · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. G. Bellemare · 2016
Later among the works it cites.
A Concise Introduction to Decentralized POMDPs
Frans A. Oliehoek and Christopher Amato · 2016
Later among the works it cites.
Understanding and improving convolutional neural networks via concatenated rectified linear units
Wenling Shang, Kihyuk Sohn, Diogo Almeida, and Honglak Lee · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Later among the works it cites.
Learning multiagent communication with backpropagation
S. Sukhbaatar, A. Szlam, and R. Fergus · 2016
Later among the works it cites.
Rock, paper, starcraft: Strategy selection in real-time strategy games
Anderson Tavares, Hector Azpurua, Amanda Santos, and Luiz Chaimowicz · 2016
Later among the works it cites.
Mason Wright · 2016
Later among the works it cites.
Poker-CNN: A pattern learning strategy for making draws and bets in poker games using convolutional networks
Nikolai Yakovenko, Liangliang Cao, Colin Raffel, and James Fan · 2016
Later among the works it cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yuri Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2017
Closest in time.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2017
Closest in time.
Transfer in reinforcement learning with successor features and generalised policy improvement
André Barreto, Will Dabney, Rémi Munos, Jonathan Hunt, Tom Schaul, David Silver, and Hado van Hasselt · 2017
Closest in time.
Safe and nested subgame solving for imperfect-information games
Noam Brown and Tuomas Sandholm · 2017
Closest in time.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Closest in time.
Stabilising experience replay for deep multi-agent reinforcement learning
Jakob Foerster, Nantas Nardelli, Gregory Farquhar, Triantafyllos Afouras, Philip H. S. Torr, Pushmeet Kohli, and Shimon Whiteson · 2017
Closest in time.
Solving for best responses and equilibria in extensive-form games with reinforcement learning methods
Amy Greenwald, Jiacui Li, and Eric Sodomka · 2017
Closest in time.
The Reactor: A sample-efficient actor-critic architecture
Audrunas Gruslys, Mohammad Gheshlaghi Azar, Marc G. Bellemare, and Remi Munos · 2017
Closest in time.
How evolution learns to generalise: Using the principles of learning theory to understand the evolution of developmental organisation
Kostas Kouvaris, Jeff Clune, Loizos Kounios, Markus Brede, and Richard A. Watson · 2017
Closest in time.
Multi-agent cooperation and the emergence of (natural) language
Angeliki Lazaridou, Alexander Peysakhovich, and Marco Baroni · 2017
Closest in time.
Multi-agent reinforcement learning in sequential social dilemmas
Joel Z. Leibo, Vinicius Zambaldia, Marc Lanctot, Janusz Marecki, and Thore Graepel · 2017
Closest in time.
Multi-agent actor-critic for mixed cooperative-competitive environments
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch · 2017
Closest in time.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisý, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Closest in time.
Deep decentralized multi-task multi-agent reinforcement learning under partial observability
Shayegan Omidshafiei, Jason Pazis, Christopher Amato, Jonathan P. How, and John Vian · 2017
Closest in time.
Predicting human behavior in unrepeated, simultaneous-move games
James R.Wright and Kevin Leyton-Brown · 2017
Closest in time.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis · 2017
Closest in time.
Multiagent cooperation and competition with deep reinforcement learning
Ardi Tampuu, Tambet Matiisen, Dorian Kodelja, Ilya Kuzovkin, Kristjan Korjus, Juhan Aru, Jaan Aru, and Raul Vicente · 2017
Closest in time.