Fetching the paper…
Reading the bibliography…
The key challenge in multiagent learning is learning a best response to the behaviour of other agents, which may be non-stationary: if the other agents adapt their strategy as well, the learning target moves.
Iterative solution of games by fictitious play
George W. Brown · 1951
Earlier work this paper cites.
A Markovian decision process
Richard Bellman · 1957
Earlier work this paper cites.
Subjectivity and correlation in randomized strategies
Robert J. Aumann · 1974
Earlier work this paper cites.
Prospect theory: An analysis of decision under risk
Daniel Kahneman and Amos Tversky · 1979
Earlier work this paper cites.
The evolution of cooperation
Robert Axelrod and William D. Hamilton · 1981
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Herbert Robbins · 1985
Earlier work this paper cites.
The complexity of Markov decision processes
Christos H. Papadimitriou and John N. Tsitsiklis · 1987
Earlier work this paper cites.
A general theory of equilibrium selection in games
John C. Harsanyi and Reinhard Selten · 1988
Earlier work this paper cites.
Learning from delayed rewards
John Watkins · 1989
Earlier work this paper cites.
Game Theory
Drew Fudenberg and Jean Tirole · 1991
Earlier work this paper cites.
Q-learning
Christopher Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Multi-Agent Reinforcement Learning: Independent vs. Cooperative Agents
Ming Tan · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L. Littman · 1994
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Learning to coordinate without sharing information
Sandip Sen, Mahendra Sekaran, and John Hale · 1994
Earlier work this paper cites.
On Players’ Models of Other Players: Theory and Experimental Evidence
Dale O. Stahl and P.W. Wilson · 1995
Earlier work this paper cites.
Evolutionary game theory
Jörgen W. Weibull · 1995
Earlier work this paper cites.
Planning, learning and coordination in multiagent decision processes
Craig Boutilier · 1996
Earlier work this paper cites.
Algorithms for sequential decision making
Michael L. Littman · 1996
Earlier work this paper cites.
Learning in the presence of concept drift and hidden contexts
Gerhard Widmer and Miroslav Kubat · 1996
Earlier work this paper cites.
Learning Through Reinforcement and Replicator Dynamics
Tilman Börgers and Rajiv Sarin · 1997
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Elevator Group Control Using Multiple Reinforcement Learning Agents
Robert H. Crites and Andrew G. Barto · 1998
Earlier work this paper cites.
Negotiation decision functions for autonomous agents
Peyman Faratin, Carles Sierra, and Nicholas R. Jennings · 1998
Earlier work this paper cites.
Multiagent Reinforcement Learning: Theoretical Framework and an Algorithm
Junling Hu and Michael P. Wellman · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie P. Kaelbling, Michael L. Littman, and Anthony R. Cassandra · 1998
Earlier work this paper cites.
Determining successful negotiation strategies: an evolutionary approach
Noyda Matos, Carles Sierra, and Nicholas R. Jennings · 1998
Earlier work this paper cites.
Reinforcement Learning An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Statistical learning theory
Vladimir Naumovich Vapnik · 1998
Earlier work this paper cites.
Interactive epistemology I: knowledge
Robert J. Aumann · 1999
Earlier work this paper cites.
Experience-weighted attraction learning in normal form games
Colin F. Camerer and T Hua Ho · 1999
Earlier work this paper cites.
An Environment Model for Nonstationary Reinforcement Learning
Samuel P. M. Choi, Dit-Yan Yeung, and Nevin L. Zhang · 1999
Earlier work this paper cites.
Rational Coordination in Multi-Agent Environments
Piotr J. Gmytrasiewicz and Edmund H. Durfee · 2000
Earlier work this paper cites.
What is rational about Nash equilibria?
Mathias Risse · 2000
Earlier work this paper cites.
Convergence results for single-step on-policy reinforcement-learning algorithms
Satinder Singh, Tommi Jaakkola, Michael L. Littman, and Csaba Szepesvári · 2000
Earlier work this paper cites.
Hidden-mode markov decision processes for nonstationary sequential decision making
Samuel P. M. Choi, Dit-Yan Yeung, and Nevin L. Zhang · 2001
Earlier work this paper cites.
Cognition and Behavior in Normal–Form Games: An Experimental Study
Miguel Costa Gomes, Vincent P. Crawford, and B. Broseta · 2001
Earlier work this paper cites.
Ten little treasures of game theory and ten intuitive contradictions
Jacob K. Goeree and C.A. Holt · 2001
Earlier work this paper cites.
Market performance of adaptive trading agents in synchronous double auctions
Wei-Tek Hsu and Von-Wun Soo · 2001
Earlier work this paper cites.
Automated negotiation: Prospects, methods and challenges
Nicholas R. Jennings, Peyman Faratin, Alessio R. Lomuscio, Simon Parsons, Michael J. Wooldridge, and Carles Sierra · 2001
Earlier work this paper cites.
Strategic Negotiation in Multiagent Environments
Sarit Kraus · 2001
Earlier work this paper cites.
Friend-or-foe Q-learning in general-sum games
Michael L. Littman · 2001
Earlier work this paper cites.
Implicit Negotiation in Repeated Games
Michael L. Littman and Peter Stone · 2001
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Peter Auer · 2002
Earlier work this paper cites.
Multiagent learning using a variable learning rate
Michael Bowling and Manuela Veloso · 2002
Earlier work this paper cites.
Sophisticated Experience-Weighted Attraction Learning and Strategic Teaching in Repeated Games
Colin F. Camerer, Teck-Hua Ho, and Juin-Kuan Chong · 2002
Earlier work this paper cites.
A multiagent reinforcement learning algorithm using extended optimal response
Nobuo Suematsu and Akira Hayashi · 2002
Earlier work this paper cites.
Pricing in agent economies using multi-agent q-learning
Gerald Tesauro and Jeffrey O. Kephart · 2002
Earlier work this paper cites.
Analyzing complex strategic interactions in multi-agent systems
William E. Walsh, Rajarshi Das, Gerald Tesauro, and Jeffrey O Kephart · 2002
Earlier work this paper cites.
R-MAX a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I. Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
Behavioral Game Theory: Experiments in Strategic Interaction (Roundtable Series in Behavioral Economics)
Colin F. Camerer · 2003
Earlier work this paper cites.
Correlated Q-learning
Amy Greenwald and Keith Hall · 2003
Earlier work this paper cites.
On agent-mediated electronic commerce
Minghua He, Nicholas R. Jennings, and Ho-fung Leung · 2003
Earlier work this paper cites.
Nash Q-learning for general-sum stochastic games
Junling Hu and Michael P. Wellman · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Integrating mobile and intelligent agents in advanced e-commerce: A survey
Ryszard Kowalczyk, Mihaela Ulieru, and Rainer Unland · 2003
Earlier work this paper cites.
Learning To Cooperate in a Social Dilemma: A Satisficing Approach to Bargaining
Jeffrey L. Stimpson and Michael A. Goodrich · 2003
Earlier work this paper cites.
Extending Q-learning to general adaptive multi-agent systems
Gerald Tesauro · 2003
Earlier work this paper cites.
Performance bounded reinforcement learning in strategic interactions
Bikramjit Banerjee and Jing Peng · 2004
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games
Daniel S. Bernstein, Eric A Hansen, Shlomo Zilberstein, and Christopher Amato · 2004
Earlier work this paper cites.
Convergence and no-regret in multiagent learning
Michael Bowling · 2004
Earlier work this paper cites.
A cognitive hierarchy model of games
Colin F. Camerer, Teck-Hua Ho, and Juin-Kuan Chong · 2004
Earlier work this paper cites.
Learning an opponent’s preferences to make effective multi-issue negotiation trade-offs
Robert M. Coehoorn and Nicholas R. Jennings · 2004
Earlier work this paper cites.
Predicting agents tactics in automated negotiation
Chongming Hou · 2004
Earlier work this paper cites.
Run the GAMUT: a comprehensive approach to evaluating game-theoretic algorithms
Eugene Nudelman, Jennifer Wortman, Yoav Shoham, and Kevin Leyton-Brown · 2004
Earlier work this paper cites.
New criteria and a new algorithm for learning in multi-agent systems
Rob Powers and Yoav Shoham · 2004
Earlier work this paper cites.
Best-response multiagent learning in non-stationary environments
Michael Weinberg and Jeffrey S. Rosenschein · 2004
Earlier work this paper cites.
Efficient learning of multi-step best response
Bikramjit Banerjee and Jing Peng · 2005
Earlier work this paper cites.
Colored trails: a formalism for investigating decision-making in strategic environments
Ya’akov Gal, Barbara J Grosz, Sarit Kraus, Avi Pfeffer, and Stuart Shieber · 2005
Earlier work this paper cites.
A framework for sequential planning in multiagent settings
Piotr J. Gmytrasiewicz and Prashant Doshi · 2005
Earlier work this paper cites.
Effective Short-Term Opponent Exploitation in Simplified Poker
Bret Hoehn, Finnegan Southey, Robert C. Holte, and Valeriy Bulitko · 2005
Earlier work this paper cites.
Non-stationary Policy Learning in 2-player Zero Sum Games
Steven Jensen, Daniel Boley, Maria Gini, and Paul Schrater · 2005
Earlier work this paper cites.
Individual Q-learning in normal form games
David S. Leslie and E. J. Collins · 2005
Cited alongside, same era.
Cooperative Multi-Agent Learning: The State of the Art
Liviu Panait and Sean Luke · 2005
Cited alongside, same era.
Learning against opponents with bounded memory
Rob Powers and Yoav Shoham · 2005
Cited alongside, same era.
AWESOME: A general multiagent learning algorithm that converges in self-play and learns a best response against stationary opponents
Vincent Conitzer and Tuomas Sandholm · 2006
Cited alongside, same era.
Dealing with non-stationary environments using context detection
Bruno C. Da Silva, Eduardo W. Basso, Ana L.C. Bazzan, and Paulo M. Engel · 2006
Cited alongside, same era.
On the Difficulty of Achieving Equilibrium in Interactive POMDPs
Prashant Doshi and Piotr J. Gmytrasiewicz · 2006
Negotiating concurrently with unknown opponents in complex, real-time domains
Colin R. Williams, Valentin Robu, Enrico H. Gerding, and Nicholas R. Jennings · 2012
Later among the works it cites.
Heuristic search of multiagent influence space
Stefan J. Witwicki, Frans A. Oliehoek, and Leslie P. Kaelbling · 2012
Later among the works it cites.
A framework for modeling population strategies by depth of reasoning
Michael Wunder, John Robert Yaros, Michael Kaisers, and Michael L. Littman · 2012
Later among the works it cites.
Addressing the Policy-bias of Q-learning by Repeating Updates
Sherief Abdallah and Michael Kaisers · 2013
Later among the works it cites.
A game-theoretic model and best-response learning method for ad hoc coordination in multiagent systems
Stefano V. Albrecht and Subramanian Ramamoorthy · 2013
Later among the works it cites.
Predicting the performance of opponent models in automated negotiation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Multi-armed Bandit, Dynamic Environments and Meta-Bandits
Cédric Hartland, Sylvain Gelly, Nicolas Baskiotis, Olivier Teytaud, and Michèle Sebag · 2006
Cited alongside, same era.
Learning to cooperate in multi-agent social dilemmas
Enrique Munoz de Cote, Alessandro Lazaric, and Marcello Restelli · 2006
Cited alongside, same era.
Lenience towards teammates helps in cooperative multiagent learning
Liviu Panait, Keith Sullivan, and Sean Luke · 2006
Cited alongside, same era.
Particle filtering for dynamic agent modelling in simplified poker
Nolan Bard and Michael Bowling · 2007
Cited alongside, same era.
Predicting and Preventing Coordination Problems in Cooperative Q-learning Systems
Nancy Fulda and Dan Ventura · 2007
Cited alongside, same era.
Computing Robust Counter-Strategies
Michael Johanson, Martin A. Zinkevich, and Michael Bowling · 2007
Cited alongside, same era.
Tim Baarslag, Mark J.C. Hendrikx, Koen V. Hindriks, and Catholijn M. Jonker · 2013
Later among the works it cites.
Online implicit agent modelling
Nolan Bard, Michael Johanson, Neil Burch, and Michael Bowling · 2013
Later among the works it cites.
Teamwork with Limited Knowledge of Teammates
Samuel Barrett, Peter Stone, Sarit Kraus, and Avi Rosenfeld · 2013
Later among the works it cites.
Multiagent learning in the presence of memory-bounded agents
Doran Chakraborty and Peter Stone · 2013
Later among the works it cites.
Targeted opponent modeling of memory-bounded agents
Doran Chakraborty, Noa Agmon, and Peter Stone · 2013
Later among the works it cites.
How much does it help to know what she knows you know? An agent-based simulation study
Harmen de Weerd, Rineke Verbrugge, and Bart Verheij · 2013
Later among the works it cites.
Deep Learning Methods and Applications
Li Deng and Dong Yu · 2013
Later among the works it cites.
Modeling non-stationary opponents
Pablo Hernandez-Leal, Enrique Munoz de Cote, and L. Enrique Sucar · 2013
Later among the works it cites.
Power TAC: A competitive economic simulation of the smart grid
Wolfgang Ketter, John Collins, and Prashant P. Reddy · 2013
Later among the works it cites.
Learning in non-stationary MDPs as transfer learning
M. M. Hassan Mahmud and Subramanian Ramamoorthy · 2013
Later among the works it cites.
Optimal Regret Bounds for Selecting the State Representation in Reinforcement Learning
Odalric-Ambrym Maillard, Phuong Nguyen, Ronald Ortner, and Daniil Ryabko · 2013
Later among the works it cites.
Are You Thinking What I’m Thinking? An Evaluation of a Simplified Theory of Mind , pages 44–57
David V. Pynadath, Ning Wang, and Stacy C. Marsella · 2013
Later among the works it cites.
Teaching on a Budget: Agents advising agents in reinforcement learning
Lisa Torrey and Matthew E. Taylor · 2013
Later among the works it cites.
Multiagent Systems
Gerhard Weiss, editor · 2013
Later among the works it cites.
Cooperating with Unknown Teammates in Complex Domains: A Robot Soccer Case Study of Ad Hoc Teamwork
Samuel Barrett and Peter Stone · 2014
Later among the works it cites.
Stochastic multi-armed-bandit problem with non-stationary rewards
Omar Besbes, Yonatan Gur, and Assaf Zeevi · 2014
Later among the works it cites.
Towards minimizing disappointment in repeated games
Jacob W. Crandall · 2014
Later among the works it cites.
Fast adaptive learning in repeated stochastic games by game abstraction
Mohamed Elidrisi, Nicholas Johnson, Maria Gini, and Jacob W. Crandall · 2014
Later among the works it cites.
A survey on concept drift adaptation
Joao Gama, Indre Zliobaite, Albert Bifet, Mykola Pechenizkiy, and Abdelhamid Bouchachia · 2014
Later among the works it cites.
Using a priori information for fast learning against non-stationary opponents
Pablo Hernandez-Leal, Enrique Munoz de Cote, and L. Enrique Sucar · 2014
Later among the works it cites.
CLEANing the reward: counterfactual actions to remove exploratory action noise in multiagent learning
Chris HolmesParker, Matthew E. Taylor, Adrian Agogino, and Kagan Tumer · 2014
Later among the works it cites.
Decentralised Multi-Agent Reinforcement Learning for Dynamic and Uncertain Environments
Andrei Marinescu, Ivana Dusparic, Adam Taylor, Vinny Cahill, and Siobhán Clarke · 2014
Later among the works it cites.
Application impact of multi-agent systems and technologies: a survey
J. P. Müller and K. Fischer · 2014
Later among the works it cites.
Regret bounds for restless Markov bandits
Ronald Ortner, Daniil Ryabko, Peter Auer, and Rémi Munos · 2014
Later among the works it cites.
Level-0 meta-models for predicting human behavior in games
James Robert Wright and Kevin Leyton-Brown · 2014
Later among the works it cites.
Decision-theoretic Clustering of Strategies
Nolan Bard, Deon Nicholas, Csaba Szepesvári, and Michael Bowling · 2015
Later among the works it cites.
Evolutionary Dynamics of Multi-Agent Learning: A Survey
Daan Bloembergen, Karl Tuyls, Daniel Hennes, and Michael Kaisers · 2015
Later among the works it cites.
Negotiating with other minds: the role of recursive theory of mind in negotiation with incomplete information
Harmen de Weerd, Rineke Verbrugge, and Bart Verheij · 2015
Later among the works it cites.
Toward natural turn-taking in a virtual human negotiation agent
David DeVault, Johnathan Mell, and Jonathan Gratch · 2015
Later among the works it cites.
Defender strategies in domains involving frequent adversary interaction
Fei Fang, Peter Stone, and Milind Tambe · 2015
Later among the works it cites.
Bayesian Reinforcement Learning: A Survey
Mohammed Ghavamzadeh, Shie Mannor, Joelle Pineau, and Aviv Tamar · 2015
Later among the works it cites.
Negotiation as a challenge problem for virtual humans
Jonathan Gratch, David DeVault, Gale M. Lucas, and Stacy Marsella · 2015
Later among the works it cites.
Bidding in Non-Stationary Energy Markets
Pablo Hernandez-Leal, Matthew E. Taylor, L. Enrique Sucar, and Enrique Munoz de Cote · 2015
Later among the works it cites.
P-MARL: Prediction-Based Multi-Agent Reinforcement Learning for Non-Stationary Environments
Andrei Marinescu, Ivana Dusparic, Adam Taylor, Vinny Cahill, and Siobhán Clarke · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
Influence-optimistic local values for multiagent planning
Frans A. Oliehoek, Matthijs T. J. Spaan, and Stefan J. Witwicki · 2015
Later among the works it cites.
Individual Planning in Agent Populations: Exploiting Anonymity and Frame-Action Hypergraphs
Ekhlas Sonu, Yingke Chen, and Prashant Doshi · 2015
Later among the works it cites.
A Multiagent Approach to Variable-Rate Electric Vehicle Charging Coordination
Konstantina Valogianni, Wolfgang Ketter, and John Collins · 2015
Later among the works it cites.
A Reinforcement Learning Approach for Interdomain Routing with Link Prices
Peter Vrancx, Pasquale Gurzi, Abdel Rodriguez, Kris Steenhaut, and Ann Nowe · 2015
Later among the works it cites.
Multiagent Learning of Coordination in Loosely Coupled Multiagent Systems
Chao Yu, Minjie Zhang, Fenghui Ren, and Guozhen Tan · 2015
Later among the works it cites.
Bayesian-based preference prediction in bilateral multi-issue negotiation between intelligent agents
Jihang Zhang, Fenghui Ren, and Minjie Zhang · 2015
Later among the works it cites.
Addressing Environment Non-Stationarity by Repeating Q-learning Updates
Sherief Abdallah and Michael Kaisers · 2016
Later among the works it cites.
Learning about the opponent in automated bilateral negotiation: a comprehensive survey of opponent modeling techniques
Tim Baarslag, Mark J.C. Hendrikx, Koen V. Hindriks, and Catholijn M. Jonker · 2016
Later among the works it cites.
Deep Reinforcement Learning Variants of Multi-Agent Learning Algorithms
Alvaro O. Castaneda · 2016
Later among the works it cites.
Who speaks for AI?
Eric Eaton, Peter Stone, Toby Walsh, Michael Wooldridge, Tom Dietterich, Maria Gini, Barbara J. Grosz, Charles L Isbell, Subbarao Kambhampati, Michael L. Littman, Francesca Rossi, and Stuart J. Russell · 2016
Later among the works it cites.
Learning to communicate with deep multi-agent reinforcement learning
Jakob N. Foerster, Yannis M Assael, Nando De Freitas, and Shimon Whiteson · 2016
Later among the works it cites.
Opponent modeling in deep reinforcement learning
He He, Jordan Boyd-Graber, Kevin Kwok, and Hal Daume · 2016
Later among the works it cites.
Ad hoc teamwork by learning teammates’ task
Francisco S. Melo and Alberto Sardinha · 2016
Later among the works it cites.
Bayesian Policy Reuse
Benjamin Rosman, Majd Hawasly, and Subramanian Ramamoorthy · 2016
Later among the works it cites.
Solving transition-independent multi-agent MDPs with sparse interactions
Joris Scharpff, Diedrerik M. Roijers, Frans A. Oliehoek, Matthijs T. J. Spaan, and Mathijs de Weerdt · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, A Huang, C J Maddison, A Guez, L Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Later among the works it cites.
Automatic Curriculum Graph Generation for Reinforcement Learning Agents
Maxwell Svetlik, Matteo Leonetti, Jivko Sinapov, Rishi Shah, Nick Walker, and Peter Stone · 2016
Later among the works it cites.
Theoretically grounded policy advice from multiple teachr in reinfocement learning settings with Applications to negative transfer
Yusen Zhan, Haitham Bou Ammar, and Matthew E. Taylor · 2016
Later among the works it cites.
When will negotiation agents be able to represent us? the challenges and opportunities for autonomous negotiators
Tim Baarslag, Michael Kaisers, Enrico H. Gerding, Catholijn M. Jonker, and Jonathan Gratch · 2017
Closest in time.
Quickest Change Detection Approach to Optimal Control in Markov Decision Processes with Model Changes
Taposh Banerjee, Miao Liu, and Jonathan P How · 2017
Closest in time.
Coordinated Versus Decentralized Exploration In Multi-Agent Multi-Armed Bandits
Mithun Chakraborty, Sanmay Das, Brendan Juba, and Kai Yee Phoebe Chua · 2017
Closest in time.
Simultaneously Learning and Advising in Multiagent Reinforcement Learning
Felipe Leno da Silva, Ruben Glatt, and Anna Helena Reali Costa · 2017
Closest in time.
Safely using predictions in general-sum normal form games
Steven Damer and Maria Gini · 2017
Closest in time.
Cooperative Multi-agent Control using deep reinforcement learning
Jayesh K Gupta, Maxim Egorov, and Mykel J Kochenderfer · 2017
Closest in time.
Identifying Unknown Unknowns in the Open World: Representations and Policies for Guided Exploration
Himabindu Lakkaraju, Ece Kamar, Rich Caruana, and Eric Horvitz · 2017
Closest in time.
Multi-agent Reinforcement Learning in Sequential Social Dilemmas
Joel Z. Leibo, V Zambaldi, M Lanctot, and J Marecki · 2017
Closest in time.
Allocating training instances to learning agents for team formation
Somchaya Liemhetcharat and Manuela Veloso · 2017
Closest in time.
Autonomous Task Sequencing for Customized Curriculum Design in Reinforcement Learning
Sanmit Narvekar, Jivko Sinapov, and Peter Stone · 2017
Closest in time.
Multiagent cooperation and competition with deep reinforcement learning
Ardi Tampuu, Tambet Matiisen, Dorian Kodelja, Ilya Kuzovkin, Kristjan Korjus, Juhan Aru, Jaan Aru, and Raul Vicente · 2017
Closest in time.
The minds of many: opponent modelling in a stochastic game
Friedrich Van der Osten, Michael Kirley, and Tim Miller · 2017
Closest in time.
Is multiagent deep reinforcement learning the answer or the question? a brief survey
Pablo Hernandez-Leal, Bilal Kartal, and Matthew E Taylor · 2018
Closest in time.
Thanh Thi Nguyen, Ngoc Duy Nguyen, and Saeid Nahavandi · 2018
Closest in time.