Fetching the paper…
Reading the bibliography…
We present the development and analysis of a reinforcement learning (RL) algorithm designed to solve continuous-space mean field game (MFG) and mean field control (MFC) problems in a unified manner.
A class of markov processes associated with nonlinear parabolic equations
McKean, H. P. (1966) · 1911
Earlier work this paper cites.
Approximately solving mean field games via entropy-regularized deep reinforcement learning
Cui, K. and Koeppl, H. (2021) · 1917
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Propagation of chaos for a class of nonlinear parabolic equations
McKean, H. P. (1967) · 1967
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Learning From Delayed Rewards
Watkins, C. (1989) · 1989
Earlier work this paper cites.
Stochastic approximation with two time scales
Borkar, V. S. (1997) · 1997
Earlier work this paper cites.
The actor-critic algorithm as multi-time-scale stochastic approximation
Borkar, V. S. and Konda, V. R. (1997) · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y. (1999) · 1999
Earlier work this paper cites.
On actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N. (2003) · 2003
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Hyvärinen, A. (2005) · 2005
Earlier work this paper cites.
Large population stochastic dynamic games: closed-loop McKean-Vlasov systems and the Nash certainty equivalence principle
Huang, M., Malhamé, R. P., and Caines, P. E. (2006) · 2006
Earlier work this paper cites.
Mean field games
Lasry, J.-M. and Lions, P.-L. (2007) · 2007
Earlier work this paper cites.
Stochastic Approximation: A Dynamical Systems Viewpoint
Borkar, V. S. (2008) · 2008
Cited alongside, same era.
Off-policy actor-critic
Degris, T., White, M., and Sutton, R. S. (2012) · 2012
Cited alongside, same era.
Reinforcement Learning in Continuous State and Action Spaces
van Hasselt, H. (2012) · 2012
Cited alongside, same era.
Mean field games and systemic risk
Carmona, R., Fouque, J., and Sun, L. (2015) · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Cited alongside, same era.
Variational inference with normalizing flows
Rezende, D. J. and Mohamed, S. (2015) · 2015
Cited alongside, same era.
CEMRACS 2017: Numerical probabilistic approach to MFG
Mean field games: A paradigm for individual-mass interactions
Malhamé, R. P. and Graves, C. (2020) · 2020
Later among the works it cites.
Reinforcement learning in continuous time and space: A stochastic control approach
Wang, H., Zariphopoulou, T., and Zhou, X. Y. (2020) · 2020
Later among the works it cites.
Mean field games flock! the reinforcement learning way
Perrin, S., Laurière, M., Pérolat, J., Geist, M., Élie, R., and Pietquin, O. (2021) · 2021
Later among the works it cites.
Entropy regularization for mean field games with learning
Guo, X., Xu, R., and Zariphopoulou, T. (2022) · 2022
Later among the works it cites.
Scalable deep reinforcement learning algorithms for mean field games
Laurière, M., Perrin, S., Girgin, S., Muller, P., Jain, A., Cabannes, T., Piliouras, G., Perolat, J., Elie, R., Pietquin, O., and Geist, M. (2022) · 2022
Later among the works it cites.
Deep learning for mean field games and mean field control with applications to finance
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Angiuli, A., Graves, C., Li, H., Chassagneux, J.-F., Delarue, F., and Carmona, R. (2019) · 2017
Cited alongside, same era.
Probabilistic Theory of Mean Field Games with Applications I-II
Carmona, R. and Delarue, F. (2018) · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Cited alongside, same era.
Learning mean-field games
Guo, X., Hu, A., Xu, R., and Zhang, J. (2019) · 2019
Cited alongside, same era.
Generative modeling by estimating gradients of the data distribution
Song, Y. and Ermon, S. (2019) · 2019
Cited alongside, same era.
Alternative way to derive the distribution of the multivariate ornstein–uhlenbeck process
Vatiwutipong, P. and Phewchean, N. (2019) · 2019
Cited alongside, same era.
Carmona, R. and Laurière, M. (2023) · 2023
Closest in time.
Model-free mean-field reinforcement learning: mean-field MDP and mean-field Q-learning
Carmona, R., Laurière, M., and Tan, Z. (2023) · 2023
Closest in time.
Global convergence of two-timescale actor-critic for solving linear quadratic regulator
Chen, X., Duan, J., Liang, Y., and Zhao, L. (2023) · 2023
Closest in time.
Actor-critic learning for mean-field control in continuous time
Frikha, N., Germain, M., Laurière, M., Pham, H., and Song, X. (2023) · 2023
Closest in time.
Q-learning in continuous time
Jia, Y. and Zhou, X. Y. (2023) · 2023
Closest in time.
Continuous-time q-learning for mckean-vlasov control problems
Wei, X. and Yu, X. (2023) · 2023
Closest in time.
Finite horizon deep reinforcement learning for mean field problems in continuous spaces
Angiuli, A., Fouque, J.-P., Hu, R., and Raydan, A. (2024) · 2024
Closest in time.