Fetching the paper…
Reading the bibliography…
One of the challenges for multi-agent reinforcement learning (MARL) is designing efficient learning algorithms for a large system in which each agent has only limited or partial information of the entire system.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
1911
Earlier work this paper cites.
1912
Earlier work this paper cites.
Dawson D (1993) Measure-valued Markov processes. École d’été de probabilités de Saint-Flour XXI-1991 , 1–260 (Springer)
1991
Earlier work this paper cites.
Konda VR, Tsitsiklis JN (2000) Actor-critic algorithms. Advances in Neural Information Processing Systems , volume 12, 1008–1014
2000
Earlier work this paper cites.
Sutton RS, McAllester DA, Singh SP, Mansour Y (2000) Policy gradient methods for reinforcement learning with function approximation. Advances in Neural Information Processing Systems , volume 99, 1057–1063
2000
Earlier work this paper cites.
Kakade S, Langford J (2002) Approximately optimal approximate reinforcement learning. International Conference on Machine Learning , 267–274 (PMLR)
2002
Earlier work this paper cites.
Rabbat M, Nowak R (2004) Distributed optimization in sensor networks. International Symposium on Information Processing in Sensor Networks , 20–27
2004
Earlier work this paper cites.
2004
Earlier work this paper cites.
Micchelli CA, Xu Y, Zhang H (2006) Universal kernels. Journal of Machine Learning Research 7(12):2651–2667
2006
Earlier work this paper cites.
Rahimi A, Recht B (2008) Uniform approximation of functions with random bases. Annual Allerton Conference on Communication, Control, and Computing , 555–561 (IEEE)
2008
Earlier work this paper cites.
Glorot X, Bengio Y (2010) Understanding the difficulty of training deep feedforward neural networks. International Conference on Artificial Intelligence and Statistics , 249–256
2010
Earlier work this paper cites.
Cao Y, Yu W, Ren W, Chen G (2012) An overview of recent progress in the study of distributed multi-agent coordination. IEEE Transactions on Industrial informatics 9(1):427–438
2012
Earlier work this paper cites.
Sra S, Nowozin S, Wright SJ (2012) Optimization for Machine Learning (MIT Press)
2012
Earlier work this paper cites.
El-Tantawy S, Abdulhai B, Abdelgawad H (2013) Multi-agent reinforcement learning for integrated network of adaptive traffic signal controllers (MARLIN-ATSC): Methodology and large-scale application on downtown Toronto. IEEE Transactions on Intelligent Transportation Systems 14(3):1140–1150
2013
Earlier work this paper cites.
Gamarnik D (2013) Correlation decay method for decision, optimization, and inference in large-scale networks. Theory Driven by Influential Applications , 108–121 (INFORMS)
2013
Earlier work this paper cites.
Gamarnik D, Goldberg DA, Weber T (2014) Correlation decay in random decision networks. Mathematics of Operations Research 39(2):229–261
2014
Earlier work this paper cites.
Iyer K, Johari R, Sundararajan M (2014) Mean-field equilibria of dynamic auctions with learning. Management Science 60(12):2949–2970
2014
Earlier work this paper cites.
Carmona R, Fouque JP, Sun LH (2015) Mean-field games and systemic risk. Communications in Mathematical Sciences 13(4):911–933
2015
Cited alongside, same era.
Pirotta M, Restelli M, Bascetta L (2015) Policy gradient in lipschitz Markov decision processes. Machine Learning 100(2):255–283
2015
Cited alongside, same era.
2016
Cited alongside, same era.
Calderone D, Sastry SS (2017) Markov decision process routing games. International Conference on Cyber-Physical Systems , 273–280 (IEEE)
2017
Cited alongside, same era.
Ji Z, Telgarsky M, Xian R (2020) Neural tangent kernels, transportation mappings, and universal approximation. International Conference on Learning Representations
2020
Later among the works it cites.
Qu G, Wierman A, Li N (2020) Scalable reinforcement learning of localized policies for multi-agent networked systems. Learning for Dynamics and Control , 256–266 (PMLR)
2020
Later among the works it cites.
Vadori N, Ganesh S, Reddy P, Veloso M (2020) Calibration of shared equilibria in general sum partially observable markov games. Advances in Neural Information Processing Systems , volume 33, 14118–14128
2020
Later among the works it cites.
Wang L, Cai Q, Yang Z, Wang Z (2020) Neural policy gradient methods: Global optimality and rates of convergence. International Conference on Learning Representations
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
Lowe R, Wu YI, Tamar A, Harb J, Pieter Abbeel O, Mordatch I (2017) Multi-agent actor-critic for mixed cooperative-competitive environments. Advances in Neural Information Processing Systems , volume 30, 6382–6393
2017
Cited alongside, same era.
Bhandari J, Russo D, Singal R (2018) A finite time analysis of temporal difference learning with linear function approximation. Conference on Learning Theory , 1691–1692 (PMLR)
2018
Cited alongside, same era.
Foerster J, Farquhar G, Afouras T, Nardelli N, Whiteson S (2018) Counterfactual multi-agent policy gradients. AAAI Conference on Artificial Intelligence , volume 32
2018
Cited alongside, same era.
Guériau M, Dusparic I (2018) Samod: Shared autonomous mobility-on-demand using decentralized reinforcement learning. International Conference on Intelligent Transportation Systems , 1558–1563 (IEEE)
2018
Cited alongside, same era.
Jin J, Song C, Li H, Gai K, Wang J, Zhang W (2018) Real-time bidding with multi-agent reinforcement learning in display advertising. ACM International Conference on Information and Knowledge Management , 2193–2201
2018
Cited alongside, same era.
Rashid T, Samvelyan M, Schroeder C, Farquhar G, Foerster J, Whiteson S (2018) QMIX: Monotonic value function factorisation for deep multi-agent reinforcement learning. International Conference on Machine Learning , 4295–4304 (PMLR)
2018
Cited alongside, same era.
Allen-Zhu Z, Li Y, Song Z (2019) A convergence theory for deep learning via over-parameterization. International Conference on Machine Learning , 242–252 (PMLR)
2019
Cited alongside, same era.
Xu P, Gao F, Gu Q (2020) An improved convergence analysis of stochastic variance-reduced policy gradient. Uncertainty in Artificial Intelligence , 541–551 (PMLR)
2020
Later among the works it cites.
You X, Li X, Xu Y, Feng H, Zhao J, Yan H (2020) Toward packet routing with fully distributed multiagent deep reinforcement learning. IEEE Transactions on Systems, Man, and Cybernetics: Systems
2020
Later among the works it cites.
Agarwal A, Kakade SM, Lee JD, Mahajan G (2021) On the theory of policy gradient methods: Optimality, approximation, and distribution shift. Journal of Machine Learning Research 22(98):1–76
2021
Closest in time.
Aïd R, Dumitrescu R, Tankov P (2021) The entry and exit game in the electricity markets: a mean-field game approach. Journal of Dynamics & Games 8(4):331
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
Chen T, Zhang K, Giannakis GB, Basar T (2021) Communication-efficient policy gradient methods for distributed reinforcement learning. IEEE Transactions on Control of Network Systems
2021
Closest in time.
2021
Closest in time.
Gu H, Guo X, Wei X, Xu R (2021) Mean-field controls with Q-learning for cooperative MARL: convergence and complexity analysis. SIAM Journal on Mathematics of Data Science 3(4):1168–1196
2021
Closest in time.
2021
Closest in time.
Li Y, Tang Y, Zhang R, Li N (2021) Distributed reinforcement learning for decentralized linear quadratic control: A derivative-free policy optimization approach. IEEE Transactions on Automatic Control
2021
Closest in time.
Lin Y, Qu G, Huang L, Wierman A (2021) Multi-agent reinforcement learning in stochastic networked systems. Advances in Neural Information Processing Systems , volume 34
2021
Closest in time.
Zhou Z, Mertikopoulos P, Moustakas AL, Bambos N, Glynn P (2021) Robust power management via learning and game design. Operations Research 69(1):331–345
2021
Closest in time.
Sunehag P, Lever G, Gruslys A, Czarnecki WM, Zambaldi V, Jaderberg M, Lanctot M, Sonnerat N, Leibo JZ, Tuyls K, Thore G (2018) Value-decomposition networks for cooperative multi-agent learning based on team reward. International Conference on Autonomous Agents and Multi-agent Systems , volume 3, 2085–2087
2087
Closest in time.