Fetching the paper…
Reading the bibliography…
Stochastic dynamic teams and games are rich models for decentralized systems and challenging testing grounds for multi-agent learning.
Central limit theorem for nonstationary markov chains. 1
Roland L. Dobrushin · 1956
Earlier work this paper cites.
Equilibrium in a stochastic n n -person game
Arlington M Fink et al · 1964
Earlier work this paper cites.
Team decision theory and information structures
Yu-Chi Ho · 1980
Earlier work this paper cites.
Cooling schedules for optimal annealing
Bruce Hajek · 1988
Earlier work this paper cites.
Cooperation and bounded recall
Robert J Aumann and Sylvain Sorin · 1989
Earlier work this paper cites.
Learning from Delayed Rewards
Christopher Watkins · 1989
Earlier work this paper cites.
Learning how to cooperate: Optimal play in repeated coordination games
V.P. Crawford and H. Haller · 1990
Earlier work this paper cites.
Game theory, 1991
Drew Fudenberg and Jean Tirole · 1991
Earlier work this paper cites.
Q-Learning
Christopher Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
The evolution of conventions
H. Peyton Young · 1993
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Learning to coordinate without sharing information
Sandip Sen, Mahendra Sekaran, and John Hale · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and Q-learning
John N. Tsitsiklis · 1994
Earlier work this paper cites.
Non zero-sum stochastic games in admission, service and routing control in queueing systems
Eitan Altman · 1996
Earlier work this paper cites.
Discrete-Time Markov Control Processes: Basic Optimality Criteria
Onesimo Hernandez-Lerma and Jean B. Lasserre · 1996
Earlier work this paper cites.
A generalized reinforcement-learning model: Convergence and applications
Michael L Littman and Csaba Szepesvári · 1996
Earlier work this paper cites.
Potential games
D. Monderer and L.S. Shapley · 1996
Earlier work this paper cites.
Stochastic approximation with two time scales
Vivek S Borkar · 1997
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
Caroline Claus and Craig Boutilier · 1998
Earlier work this paper cites.
Multiagent reinforcement learning: theoretical framework and an algorithm
Junling Hu, Michael P Wellman, et al · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Individual Strategy and Social Structure: An Evolutionary Theory of Institutions
H. P. Young · 1998
Cited alongside, same era.
An algorithm for distributed reinforcement learning in cooperative multi-agent systems
Martin Lauer and Martin Riedmiller · 2000
Cited alongside, same era.
Convergence results for single-step on-policy reinforcement-learning algorithms
Satinder Singh, Tommi Jaakkola, Michael L Littman, and Csaba Szepesvári · 2000
Cited alongside, same era.
Multiagent systems: A survey from a machine learning perspective
Peter Stone and Manuela Veloso · 2000
Cited alongside, same era.
Friend-or-foe q-learning in general-sum games
Michael L Littman · 2001
Cited alongside, same era.
Value-function reinforcement learning in markov games
Michael L Littman · 2001
Cited alongside, same era.
On the structure of weakly acyclic games
A. Fabrikant, A. D. Jaggard, and M. Schapira · 2010
Later among the works it cites.
Distributed q-learning for interference control in ofdma-based femtocell networks
Ana Galindo-Serrano and Lorenza Giupponi · 2010
Later among the works it cites.
Designing decentralized controllers for distributed-air-jet mems-based micromanipulators by reinforcement learning
Laëtitia Matignon, Guillaume J Laurent, Nadine Le Fort-Piat, and Yves-André Chapuis · 2010
Later among the works it cites.
Revisiting log-linear learning: Asynchrony, completeness and payoff-based implementation
Jason R Marden and Jeff S Shamma · 2012
Later among the works it cites.
Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems
Laetitia Matignon, Guillaume J Laurent, and Nadine Le Fort-Piat · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning in Markovian evolutionary games
Vivek Borkar · 2002
Cited alongside, same era.
Multiagent learning using a variable learning rate
Michael Bowling and Manuela Veloso · 2002
Cited alongside, same era.
Reinforcement learning of coordination in cooperative multi-agent systems
Spiros Kapetanakis and Daniel Kudenko · 2002
Cited alongside, same era.
Reinforcement learning to play an optimal Nash equilibrium in team Markov games
Xiaofeng Wang and Tuomas Sandholm · 2002
Cited alongside, same era.
Correlated q-learning
Amy Greenwald, Keith Hall, and Roberto Serrano · 2003
Cited alongside, same era.
Nash Q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Cited alongside, same era.
Bary S.R. Pradelski and H. Peyton Young · 2012
Later among the works it cites.
Aspiration learning in coordination games
G. C. Chasparis, A. Arapostathis, and J. S. Shamma · 2013
Later among the works it cites.
A stochastic games framework for verification and control of discrete time stochastic hybrid systems
Jerry Ding, Maryam Kamgarpour, Sean Summers, Alessandro Abate, John Lygeros, and Claire Tomlin · 2013
Later among the works it cites.
Hybrid Learning in Stochastic Games and Its Application in Network Security
F. L. Lewis and D. Liu · 2013
Later among the works it cites.
Stochastic Networked Control Systems
Serdar Yüksel and Tamer Başar · 2013
Later among the works it cites.
Achieving pareto optimality through distributed learning
Jason R. Marden, H. Peyton Young, and Lucy Y. Pao · 2014
Later among the works it cites.
Q-learning based power control algorithm for d2d communication
Shiwen Nie, Zhiqiang Fan, Ming Zhao, Xinyu Gu, and Lin Zhang · 2016
Later among the works it cites.
A survey on applications of model-free strategy learning in cognitive wireless networks
Wenbo Wang, Andres Kwasinski, Dusit Niyato, and Zhu Han · 2016
Later among the works it cites.
Lenient learning in independent-learner stochastic cooperative games
Ermo Wei and Sean Luke · 2016
Later among the works it cites.
Decentralized Q-learning for stochastic teams and games
Gürdal Arslan and Serdar Yüksel · 2017
Later among the works it cites.
A survey of learning in multiagent environments: Dealing with non-stationarity
Pablo Hernandez-Leal, Michael Kaisers, Tim Baarslag, and Enrique Munoz de Cote · 2017
Later among the works it cites.
Decentralized reinforcement learning of robot behaviors
David L Leottau, Javier Ruiz-del Solar, and Robert Babuška · 2018
Later among the works it cites.
Fully decentralized multi-agent reinforcement learning with networked agents
Kaiqing Zhang, Zhuoran Yang, Han Liu, Tong Zhang, and Tamer Basar · 2018
Later among the works it cites.
Marl-based distributed cache placement for wireless networks
Xiaosheng Lin, Yuhao Tang, Xianfu Lei, Junjuan Xia, Qingfeng Zhou, Huijun Wu, and Liseng Fan · 2019
Closest in time.
Reinforcement learning for decentralized stochastic control
Bora Yongacoglu, Gürdal Arslan, and Serdar Yüksel · 2019
Closest in time.
Reinforcement learning for decentralized stochastic control and coordination games, 2019
Bora Yongacoglu, Gürdal Arslan, and Serdar Yüksel · 2019
Closest in time.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar · 2019
Closest in time.