Fetching the paper…
Reading the bibliography…
Mean field control (MFC) is an effective way to mitigate the curse of dimensionality of cooperative multi-agent reinforcement learning (MARL) problems.
Linear-quadratic mean-field reinforcement learning: convergence of policy gradient methods
René Carmona, Mathieu Laurière, and Zongjun Tan · 1910
Earlier work this paper cites.
Model-free mean-field reinforcement learning: mean-field MDP and mean-field Q-learning
René Carmona, Mathieu Laurière, and Zongjun Tan · 1910
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Multi-agent reinforcement learning: Independent vs. cooperative agents
Ming Tan · 1993
Earlier work this paper cites.
On-line Q-learning using connectionist systems , volume 37
Gavin A Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Mean-Field Controls with Q-learning for Cooperative MARL: Convergence and Complexity Analysis
Haotian Gu, Xin Guo, Xiaoli Wei, and Renyuan Xu · 2002
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Multi-agent machine learning: A reinforcement approach
Howard M Schwartz · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
State estimation for the individual and the population in mean field control with application to demand dispatch
Yue Chen, Ana Bušić, and Sean P Meyn · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Optimal resource allocation for competitive spreading processes on bilayer networks
Nicholas J Watkins, Cameron Nowzari, Victor M Preciado, and George J Pappas · 2016
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
Limit theory for controlled mckean–vlasov dynamics
Daniel Lacker · 2017
Cited alongside, same era.
Mean field control and mean field game models with several populations
Alain Bensoussan, Tao Huang, and Mathieu Laurière · 2018
Cited alongside, same era.
Probabilistic Theory of Mean Field Games with Applications II: Mean Field Games with Common Noise and Master Equations , volume 84
René Carmona and François Delarue · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
A Q-values sharing framework for multiple independent Q-learners
Changxi Zhu, Ho-fung Leung, Shuyue Hu, and Yi Cai · 2019
Later among the works it cites.
Unified reinforcement q-learning for mean field game and control problems
Andrea Angiuli, Jean-Pierre Fouque, and Mathieu Laurière · 2020
Later among the works it cites.
On the convergence of model free learning in mean field games
Romuald Elie, Julien Perolat, Mathieu Laurière, Matthieu Geist, and Olivier Pietquin · 2020
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Later among the works it cites.
An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods
Yanli Liu, Kaiqing Zhang, Tamer Basar, and Wotao Yin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Cited alongside, same era.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2018
Cited alongside, same era.
Learning deep mean field games for modeling large population behavior
Jiachen Yang, Xiaojing Ye, Rakshit Trivedi, Huan Xu, and Hongyuan Zha · 2018
Cited alongside, same era.
Deeppool: Distributed model-free algorithm for ride-sharing using deep reinforcement learning
Abubakr O Al-Abbasi, Arnob Ghosh, and Vaneet Aggarwal · 2019
Cited alongside, same era.
Learning mean-field games
Xin Guo, Anran Hu, Renyuan Xu, and Junzi Zhang · 2019
Cited alongside, same era.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi · 2019
Cited alongside, same era.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Cited alongside, same era.
Weighted qmix: Expanding monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Gregory Farquhar, Bei Peng, and Shimon Whiteson · 2020
Later among the works it cites.
Large-scale traffic signal control using a novel multiagent reinforcement learning
Xiaoqiang Wang, Liangjun Ke, Zhimin Qiao, and Xinghua Chai · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2021
Closest in time.
Efficient Model-Based Multi-Agent Mean-Field Reinforcement Learning
Barna Pasztor, Ilija Bogunovic, and Andreas Krause · 2021
Closest in time.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Kaiqing Zhang, Zhuoran Yang, and Tamer Başar · 2021
Closest in time.
Reinforcement learning for mean-field game
Mridul Agarwal, Vaneet Aggarwal, Arnob Ghosh, and Nilay Tiwari · 2022
Closest in time.