Fetching the paper…
Reading the bibliography…
We establish the convergence of the unified two-timescale Reinforcement Learning (RL) algorithm presented in a previous work by Angiuli et al.
Approximately solving mean field games via entropy-regularized deep reinforcement learning
Cui, K. and Koeppl, H. (2021) · 1917
Earlier work this paper cites.
Differential equations without uniqueness and classical topological dynamics
Sell, G. R. (1973) · 1973
Earlier work this paper cites.
Discrete-parameter Martingales
Neveu, J. (1975) · 1975
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H. (1989) · 1989
Earlier work this paper cites.
An analog scheme for fixed point computation. i. theory
Borkar, V. and Soumyanatha, K. (1997) · 1997
Earlier work this paper cites.
Stochastic approximation with two time scales
Borkar, V. S. (1997) · 1997
Earlier work this paper cites.
Asynchronous stochastic approximations
Borkar, V. S. (1998) · 1998
Earlier work this paper cites.
Actor-critic–type learning algorithms for markov decision processes
Konda, V. R. and Borkar, V. S. (1999) · 1999
Earlier work this paper cites.
The o.d.e. method for convergence of stochastic approximation and reinforcement learning
Borkar, V. S. and Meyn, S. P. (2000) · 2000
Earlier work this paper cites.
Large population stochastic dynamic games: closed-loop McKean-Vlasov systems and the Nash certainty equivalence principle
Huang, M., Malhamé, R. P., and Caines, P. E. (2006) · 2006
Earlier work this paper cites.
Mean field games
Lasry, J.-M. and Lions, P.-L. (2007) · 2007
Earlier work this paper cites.
A comprehensive survey of multiagent reinforcement learning
Busoniu, L., Babuska, R., and De Schutter, B. (2008) · 2008
Earlier work this paper cites.
An overview of multi-agent reinforcement learning from game theoretical perspective
Yang, Y. and Wang, J. (2020) · 2011
Earlier work this paper cites.
Mean field games and mean field type control theory
Bensoussan, A., Frehse, J., and Yam, S. C. P. (2013) · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016) · 2016
Cited alongside, same era.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Gu, S., Holly, E., Lillicrap, T., and Levine, S. (2017) · 2017
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Pérolat, J., Silver, D., and Graepel, T. (2017) · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
Vecerik, M., Hester, T., Scholz, J., Wang, F., Pietquin, O., Piot, B., Heess, N., Rothörl, T., Lampe, T., and Riedmiller, M. (2017) · 2017
Cited alongside, same era.
Probabilistic Theory of Mean Field Games with Applications I-II
Carmona, R. and Delarue, F. (2018) · 2018
Reinforcement learning in continuous time and space: A stochastic control approach
Wang, H., Zariphopoulou, T., and Zhou, X. Y. (2020) · 2020
Later among the works it cites.
Mean-field controls with q-learning for cooperative marl: convergence and complexity analysis
Gu, H., Guo, X., Wei, X., and Xu, R. (2021) · 2021
Later among the works it cites.
Efficient model-based multi-agent mean-field reinforcement learning
Pasztor, B., Bogunovic, I., and Krause, A. (2021) · 2021
Later among the works it cites.
Multi-agent reinforcement learning: A selective overview of theories and algorithms
Zhang, K., Yang, Z., and Başar, T. (2021) · 2021
Later among the works it cites.
Unified reinforcement q-learning for mean field game and control problems
Angiuli, A., Fouque, J.-P., and Laurière, M. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the properties of the softmax function with application in game theory and reinforcement learning
Gao, B. and Pavel, L. (2018) · 2018
Cited alongside, same era.
Decentralised learning in systems with many, many strategic agents
Mguni, D., Jennings, J., and de Cote, E. M. (2018) · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Cited alongside, same era.
Linear-quadratic mean-field reinforcement learning: Convergence of policy gradient methods
Carmona, R., Laurière, M., and Tan, Z. (2019) · 2019
Cited alongside, same era.
Learning mean-field games
Guo, X., Hu, A., Xu, R., and Zhang, J. (2019) · 2019
Cited alongside, same era.
Finite mean field games: fictitious play and convergence to a first order continuous mean field game
Hadikhanloo, S. and Silva, F. J. (2019) · 2019
Cited alongside, same era.
Reinforcement learning in stationary mean-field games
Subramanian, J. and Mahajan, A. (2019) · 2019
Cited alongside, same era.
Entropy regularization for mean field games with learning
Guo, X., Xu, R., and Zariphopoulou, T. (2022) · 2022
Later among the works it cites.
Learning mean field games: A survey
Laurière, M., Perrin, S., Geist, M., and Pietquin, O. (2022) · 2022
Later among the works it cites.
Mean-field markov decision processes with common noise and open-loop controls
Motte, M. and Pham, H. (2022) · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Later among the works it cites.
Q-learning in regularized mean-field games
Anahtarci, B., Kariksiz, C. D., and Saldi, N. (2023) · 2023
Closest in time.
Model-free mean-field reinforcement learning: mean-field mdp and mean-field q-learning
Carmona, R., Laurière, M., and Tan, Z. (2023) · 2023
Closest in time.
Actor-critic learning for mean-field control in continuous time
Frikha, N., Germain, M., Laurière, M., Pham, H., and Song, X. (2023) · 2023
Closest in time.
Quantitative propagation of chaos for mean field markov decision process with common noise
Motte, M. and Pham, H. (2023) · 2023
Closest in time.
Oracle-free reinforcement learning in mean-field games along a single sample path
Zaman, M. A. U., Koppel, A., Bhatt, S., and Basar, T. (2023) · 2023
Closest in time.