Fetching the paper…
Reading the bibliography…
We study reinforcement learning (RL) in a setting with a network of agents whose states and actions interact in a local manner where the objective is to find localized policies such that the (discounted) global reward is maximized.
ALOHA packet system with and without slots and capture
L. Roberts · 1975
Earlier work this paper cites.
Restless bandits: Activity allocation in a changing world
P. Whittle · 1988
Earlier work this paper cites.
Stability properties of constrained queueing systems and scheduling policies for maximum throughput in multihop radio networks
L. Tassiulas and A. Ephremides · 1990
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
M. Tan · 1993
Earlier work this paper cites.
Bifurcation analysis of periodic SEIR and SIR epidemic models
Y. A. Kuznetsov and C. Piccardi · 1994
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Stable linear approximations to dynamic programming for stochastic control problems with local transitions
B. Van Roy and J. N. Tsitsiklis · 1995
Earlier work this paper cites.
Neuro-dynamic programming , volume 5
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
The dynamics of reinforcement learning in cooperative multiagent systems
C. Claus and C. Boutilier · 1998
Earlier work this paper cites.
Solving very large weakly coupled Markov decision processes
N. Meuleau, M. Hauskrecht, K.-E. Kim, L. Peshkin, L. P. Kaelbling, T. L. Dean, and C. Boutilier · 1998
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
R. S. Sutton, A. G. Barto, et al · 1998
Earlier work this paper cites.
Efficient reinforcement learning in factored MDPs
M. Kearns and D. Koller · 1999
Earlier work this paper cites.
The complexity of optimal queuing network control
C. H. Papadimitriou and J. N. Tsitsiklis · 1999
Earlier work this paper cites.
Average cost temporal-difference learning
J. N. Tsitsiklis and B. Van Roy · 1999
Earlier work this paper cites.
A survey of computational complexity results in systems and control
V. D. Blondel and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
Actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. S. Sutton, D. A. McAllester, S. P. Singh, and Y. Mansour · 2000
Earlier work this paper cites.
Value-function reinforcement learning in Markov games
M. L. Littman · 2001
Earlier work this paper cites.
Distributed control of spatially invariant systems
B. Bamieh, F. Paganini, and M. A. Dahleh · 2002
Earlier work this paper cites.
On average versus discounted reward temporal-difference learning
J. N. Tsitsiklis and B. Van Roy · 2002
Earlier work this paper cites.
Efficient solution algorithms for factored MDPs
C. Guestrin, D. Koller, R. Parr, and S. Venkataraman · 2003
Earlier work this paper cites.
Nash Q-learning for general-sum stochastic games
J. Hu and M. P. Wellman · 2003
Earlier work this paper cites.
Nonequilibrium phase transition in a model for the propagation of innovations among economic agents
M. Llas, P. M. Gleiser, J. M. López, and A. Díaz-Guilera · 2003
Earlier work this paper cites.
The power of epidemics: Robust communication for large-scale distributed systems
W. Vogels, R. van Renesse, and K. Birman · 2003
Earlier work this paper cites.
Dynamic programming and optimal control
D. P. Bertsekas · 2005
Cited alongside, same era.
Networked distributed POMDPs: A synthesis of distributed constraint optimization and POMDPs
R. Nair, P. Varakantham, M. Tambe, and M. Yokoo · 2005
Cited alongside, same era.
A characterization of convex problems in decentralized control
M. Rotkowitz and S. Lall · 2005
Cited alongside, same era.
G. Qu, C. Yu, S. Low, and A. Wierman · 2006
Cited alongside, same era.
A comprehensive survey of multiagent reinforcement learning
L. Bu, R. Babu, B. De Schutter, et al · 2008
Cited alongside, same era.
Epidemic thresholds in real networks
D. Chakrabarti, Y. Wang, C. Wang, J. Leskovec, and C. Faloutsos · 2008
A concise introduction to decentralized POMDPs
F. A. Oliehoek and C. Amato · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Later among the works it cites.
Optimal control of multiroom HVAC system: An event-based approach
Z. Wu, Q.-S. Jia, and X. Guan · 2016
Later among the works it cites.
Control of robotic mobility-on-demand systems: a queueing-theoretical perspective
R. Zhang and M. Pavone · 2016
Later among the works it cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. Wu, A. Tamar, J. Harb, O. P. Abbeel, and I. Mordatch · 2017
Later among the works it cites.
Distributed reinforcement learning via gossip
A. Mathkar and V. S. Borkar · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Information, physics, and computation
M. Mezard and A. Montanari · 2009
Cited alongside, same era.
Stochastic epidemic models: a survey
T. Britton · 2010
Cited alongside, same era.
On the achievable throughput of CSMA under imperfect carrier sensing
T. H. Kim, J. Ni, R. Srikant, and N. H. Vaidya · 2011
Cited alongside, same era.
Independent reinforcement learners in cooperative Markov games: a survey regarding coordination problems
L. Matignon, G. J. Laurent, and N. Le Fort-Piat · 2012
Cited alongside, same era.
Optimal CSMA: a survey
S.-Y. Yun, Y. Yi, J. Shin, et al · 2012
Cited alongside, same era.
Correlation decay method for decision, optimization, and inference in large-scale networks
D. Gamarnik · 2013
Cited alongside, same era.
Later among the works it cites.
On the dynamics of deterministic epidemic propagation over networks
W. Mei, S. Mohagheghi, S. Zampieri, and F. Bullo · 2017
Later among the works it cites.
Decentralized and distributed temperature control via HVAC systems in energy efficient buildings
X. Zhang, W. Shi, B. Yan, A. Malkawi, and N. Li · 2017
Later among the works it cites.
Is Q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Later among the works it cites.
Multi-agent reinforcement learning via double averaging primal-dual optimization
H.-T. Wai, Z. Yang, Z. Wang, and M. Hong · 2018
Later among the works it cites.
Fully decentralized multi-agent reinforcement learning with networked agents
K. Zhang, Z. Yang, H. Liu, T. Zhang, and T. Basar · 2018
Later among the works it cites.
Global optimality guarantees for policy gradient methods
J. Bhandari and D. Russo · 2019
Closest in time.
Reinforcement learning and deep learning based lateral control for autonomous driving [application notes]
D. Li, D. Zhao, Q. Zhang, and Y. Chen · 2019
Closest in time.
Exploiting fast decaying and locality in multi-agent MDP with tree dependence structure
G. Qu and N. Li · 2019
Closest in time.
Finite-time error bounds for linear stochastic approximation and TD learning
R. Srikant and L. Ying · 2019
Closest in time.
The gap between model-based and model-free methods on the linear quadratic regulator: An asymptotic viewpoint
S. Tu and B. Recht · 2019
Closest in time.
Temporal starvation in multi-channel CSMA networks: an analytical framework
A. Zocca · 2019
Closest in time.
Finite-sample analysis for SARSA with linear function approximation
S. Zou, T. Xu, and Y. Liang · 2019
Closest in time.
Sample complexity of asynchronous q-learning: Sharper analysis and variance reduction
G. Li, Y. Wei, Y. Chi, Y. Gu, and Y. Chen · 2020
Closest in time.
On maintaining linear convergence of distributed learning and optimization under limited communication
S. Magnússon, H. Shokri-Ghadikolaei, and N. Li · 2020
Closest in time.
Optimal, near-optimal, and robust epidemic control
D. H. Morris, F. W. Rossine, J. B. Plotkin, and S. A. Levin · 2020
Closest in time.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2021
Closest in time.
Multi-agent reinforcement learning in stochastic networked systems
Y. Lin, G. Qu, L. Huang, and A. Wierman · 2021
Closest in time.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
K. Zhang, Z. Yang, and T. Başar · 2021
Closest in time.