Fetching the paper…
Reading the bibliography…
We consider reinforcement learning (RL) methods for finding optimal policies in linear quadratic (LQ) mean field control (MFC) problems over an infinite horizon in continuous time, with common noise and entropy regularization.
Numerical Solution of Stochastic Differential Equations
Eckhard Platen Peter E. Kloeden · 1992
Earlier work this paper cites.
Policy gradient in continuous time
Rémi Munos · 2006
Earlier work this paper cites.
Probabilistic Theory of Mean Field Games: vol. I, Mean Field FBSDEs, Control, and Games
R. Carmona and F. Delarue · 2018
Earlier work this paper cites.
Probabilistic Theory of Mean Field Games: vol. II, Mean Field game with common noise and Master equations
R. Carmona and F. Delarue · 2018
Earlier work this paper cites.
Global convergence of policy gradient methods for the linear quadratic regulator
M. Fazel, R. Ge, S.M. Kakade, and M. Mesbahi · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. Sutton and A. Barto · 2018
Earlier work this paper cites.
A weak martingale approach to linear-quadratic mckean- vlasov stochastic control problems
M. Basei and H. Pham · 2019
Earlier work this paper cites.
Linear-quadratic mean-field reinforcement learning: convergence of policy gradient methods
R. Carmona, M. Laurière, and Z. Tan · 2019
Earlier work this paper cites.
Policy gradient-based algorithms for continuous-time linear quadratic control
J. Bu, A. Mesbahi, and M. Mesbahi · 2020
Cited alongside, same era.
Reinforcement learning in continuous time and space: A stochastic control approach
H. Wang, T. Zariphopoulou, and X.Y. Zhou · 2020
Cited alongside, same era.
Bounds on tail probabilities for quadratic forms in dependent sub-gaussian random variables
K. Zajkowski · 2020
Cited alongside, same era.
Mean field controls with Q-learning for cooperative MARL: convergence and complexity analysis
H. Gu, X. Guo, X. Wei, and R. Xu · 2021
Cited alongside, same era.
Policy gradient methods for the noisy linear quadratic regulator over a finite horizon
B. Hambly, R. Xu, and H. Yang · 2021
Cited alongside, same era.
Convergence and sample complexity of gradient methods for the model-free linear–quadratic regulator problem
H. Mohammadi, A. Zare, M. Soltanolkotabi, and M. R. Jovanović · 2022
Later among the works it cites.
Model-free mean-field reinforcement learning: Mean-field MDP and mean-field Q-learning
R. Carmona, M. Laurière, and Z. Tan · 2023
Later among the works it cites.
Actor-critic learning for mean-field control in continuous time
N. Frikha, M. Germain, M. Laurière, H. Pham, and X. Song · 2023
Later among the works it cites.
Fast policy learning for linear-quadratic control with entropy regularization
X. Guo, X. Li, and R. Xu · 2023
Later among the works it cites.
q q learning in continuous time
Y. Jia and X.Y. Zhou · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Policy gradient and actor–critic learning in continuous time and space: Theory and algorithms
Y. Jia and X.Y. Zhou · 2021
Cited alongside, same era.
Global convergence of policy gradient for linear-quadratic mean-field control/game in continuous time
W. Wang, J. Han, Z. Yang, and Z. Wang · 2021
Cited alongside, same era.
Unified reinforcement Q-learning for mean field game and control problems
A. Angiuli, J-.P. Fouque, and M. Laurière · 2022
Cited alongside, same era.
H. Pham and X. Warin · 2023
Later among the works it cites.
Convergence of policy gradient methods for finite-horizon exploratory linear-quadratic control problems
M. Giegrich, C. Reisinger, and Y. Zhang · 2024
Closest in time.
Optimal scheduling of entropy regularizer for continuous-time linear-quadratic reinforcement learning
L. Szpruch, T. Treetanthiploet, and Y. Zhang · 2024
Closest in time.