Fetching the paper…
Reading the bibliography…
We study multi-agent reinforcement learning (MARL) in a stochastic network of agents.
Aloha packet system with and without slots and capture
L. G. Roberts · 1975
Earlier work this paper cites.
Asynchronous stochastic approximation and Q-learning
J. N. Tsitsiklis · 1994
Earlier work this paper cites.
Reinforcement learning with soft state aggregation
S. P. Singh, T. Jaakkola, and M. I. Jordan · 1995
Earlier work this paper cites.
Neuro-dynamic programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
C. Claus and C. Boutilier · 1998
Earlier work this paper cites.
The complexity of optimal queuing network control
C. H. Papadimitriou and J. N. Tsitsiklis · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour · 1999
Earlier work this paper cites.
Distributed control of spatially invariant systems
B. Bamieh, F. Paganini, and M. A. Dahleh · 2002
Earlier work this paper cites.
Nonequilibrium phase transition in a model for the propagation of innovations among economic agents
M. Llas, P. M. Gleiser, J. M. López, and A. Díaz-Guilera · 2003
Earlier work this paper cites.
The power of epidemics: Robust communication for large-scale distributed systems
W. Vogels, R. van Renesse, and K. Birman · 2003
Earlier work this paper cites.
State abstraction discovery from irrelevant state variables
N. K. Jong and P. Stone · 2005
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
L. Li, T. J. Walsh, and M. L. Littman · 2006
Earlier work this paper cites.
Optimal backpressure routing for wireless networks with multi-receiver diversity
M. J. Neely · 2006
Earlier work this paper cites.
Dynamic Programming and Optimal Control, Vol. II
D. P. Bertsekas · 2007
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
L. Bu, R. Babu, B. De Schutter, et al · 2008
Earlier work this paper cites.
Epidemic thresholds in real networks
D. Chakrabarti, Y. Wang, C. Wang, J. Leskovec, and C. Faloutsos · 2008
Earlier work this paper cites.
Optimal control of spatially distributed systems
N. Motee and A. Jadbabaie · 2008
Earlier work this paper cites.
Networks, crowds, and markets: Reasoning about a highly connected world
D. Easley, J. Kleinberg, et al · 2012
Earlier work this paper cites.
PAC bounds for discounted MDPs
T. Lattimore and M. Hutter · 2012
Cited alongside, same era.
Independent reinforcement learners in cooperative Markov games: a survey regarding coordination problems
L. Matignon, G. J. Laurent, and N. Le Fort-Piat · 2012
Cited alongside, same era.
Markov Chains: Gibbs Fields, Monte Carlo Simulation, and Queues
P. Bremaud · 2013
Cited alongside, same era.
Correlation decay method for decision, optimization, and inference in large-scale networks
D. Gamarnik · 2013
Cited alongside, same era.
Correlation decay in random decision networks
D. Gamarnik, D. A. Goldberg, and T. Weber · 2014
Cited alongside, same era.
Abstraction selection in model-based reinforcement learning
N. Jiang, A. Kulesza, and S. Singh · 2015
Cited alongside, same era.
Mean field multi-agent reinforcement learning
Y. Yang, R. Luo, M. Li, M. Zhou, W. Zhang, and J. Wang · 2018
Later among the works it cites.
Fully decentralized multi-agent reinforcement learning with networked agents
K. Zhang, Z. Yang, H. Liu, T. Zhang, and T. Başar · 2018
Later among the works it cites.
Finite-time analysis of distributed TD(0) with linear function approximation on multi-agent reinforcement learning
T. Doan, S. Maguluri, and J. Romberg · 2019
Later among the works it cites.
Finite-time analysis and restarting scheme for linear two-time-scale stochastic approximation, 2019
T. T. Doan · 2019
Later among the works it cites.
Q-learning with UCB exploration is sample efficient for infinite-horizon MDP
K. Dong, Y. Wang, X. Chen, and L. Wang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Cited alongside, same era.
Improved bounds on the epidemic threshold of exact sis models on complex networks
N. A. Ruhi, C. Thrampoulidis, and B. Hassibi · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Cited alongside, same era.
Control of robotic mobility-on-demand systems: a queueing-theoretical perspective
R. Zhang and M. Pavone · 2016
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. Wu, A. Tamar, J. Harb, O. P. Abbeel, and I. Mordatch · 2017
Cited alongside, same era.
A unified switching system perspective and O.D.E. analysis of q-learning algorithms
D. hwan Lee and N. He · 2019
Later among the works it cites.
Reinforcement learning and deep learning based lateral control for autonomous driving [application notes]
D. Li, D. Zhao, Q. Zhang, and Y. Chen · 2019
Later among the works it cites.
Exploiting fast decaying and locality in multi-agent MDP with tree dependence structure
G. Qu and N. Li · 2019
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation and TD learning
R. Srikant and L. Ying · 2019
Later among the works it cites.
Reinforcement learning in stationary mean-field games
J. Subramanian and A. Mahajan · 2019
Later among the works it cites.
M. J. Wainwright · 2019
Later among the works it cites.
Two time-scale off-policy TD learning: Non-asymptotic analysis over Markovian samples
T. Xu, S. Zou, and Y. Liang · 2019
Later among the works it cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
K. Zhang, Z. Yang, and T. Başar · 2019
Later among the works it cites.
Temporal starvation in multi-channel csma networks: an analytical framework
A. Zocca · 2019
Later among the works it cites.
Q-learning for mean-field controls, 2020
H. Gu, X. Guo, X. Wei, and R. Xu · 2020
Closest in time.
Q-learning for mean-field controls
H. Gu, X. Guo, X. Wei, and R. Xu · 2020
Closest in time.
Scalable multi-agent reinforcement learning for networked systems with average reward
G. Qu, Y. Lin, A. Wierman, and N. Li · 2020
Closest in time.
Finite-time analysis of asynchronous stochastic approximation and q q -learning
G. Qu and A. Wierman · 2020
Closest in time.
Scalable reinforcement learning of localized policies for multi-agent networked systems
G. Qu, A. Wierman, and N. Li · 2020
Closest in time.
A finite time analysis of two time-scale actor critic methods, 2020
Y. Wu, W. Zhang, P. Xu, and Q. Gu · 2020
Closest in time.