Fetching the paper…
Reading the bibliography…
This paper studies the q-learning, recently coined as the continuous time counterpart of Q-learning by Jia and Zhou (2023), for continuous time Mckean-Vlasov control problems in the setting of entropy-regularized reinforcement learning.
N. E. Karoui and S. Méléard (1990): Martingale measures and stochastic calculus
1990
Earlier work this paper cites.
D. Stroock and S. Varadhan (1997): Multidimensional diffusion processes. volume 233 of
1997
Earlier work this paper cites.
P.-L. Lions (2006): Cours au collège de france: Théorie des jeux à champ moyens. Audio Conference
2006
Earlier work this paper cites.
Y. Sun (2006): The exact law of large numbers via Fubini extension and characterization of insurable risks
2006
Earlier work this paper cites.
R. Carmona, J. P. Fouque and L. H. Sun (2015): Mean field games and systemic risk
2015
Earlier work this paper cites.
D. Lacker (2017): Limit theory for controlled McKean-Vlasov dynamics
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. S. Sutton and A. G. Barto (2018): Reinforcement learning: An introduction. MIT press
2018
Earlier work this paper cites.
K. Doya (2020). Reinforcement learning in continuous time and space
2020
Earlier work this paper cites.
H. Wang, T. Zariphopoulou and X. Y. Zhou (2020): Reinforcement learning in continuous time and space: A stochastic control approach
2020
Earlier work this paper cites.
H. Gu, X. Guo, X. Wei and R. Xu (2021): Mean-field controls with Q-learning for cooperative MARL: convergence and complexity analysis
2021
Cited alongside, same era.
H. Gu, X. Guo, X. Wei, R. Xu (2021): Mean-field multi-agent reinforcement learning: A decentralized network approach. Forthcoming in Mathematics of Operations Research
2021
Cited alongside, same era.
2021
Cited alongside, same era.
L. Szpruch, T. Treetanthiploet and Y. Zhang. (2021): Exploration-exploitation trade-off for continuous-time episodic reinforcement learning with linear-convex models
2021
Cited alongside, same era.
A. Angiuli, J. P. Fouque and M. Laurière. (2022). Unified reinforcement Q-learning for mean field game and control problems
A. Cosso, F Gozzi, I. Kharroubi, H. Pham and M. Rosestolato (2023): Optimal control of path-dependent McKean-Vlasov SDEs in infinite dimension
2023
Closest in time.
H. Gu, X. Guo, X. Wei and R. Xu (2023): Dynamic programming principles for mean-field controls with learning
2023
Closest in time.
X. Guo, A. Hu and Y. Zhang (2023): Reinforcement learning for linear-convex models with jumps via stability analysis of feedback controls
2023
Closest in time.
2023
Closest in time.
Y. Jia and X. Y. Zhou (2023): q-learning in continuous time
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
X. Guo, R. Xu and T. Zariphopoulou (2022): Entropy regularization for mean field games with learning
2022
Cited alongside, same era.
W.U. Mondal, M. Agarwal, V. Aggarwal and S.V. Ukkusuri (2022): On the approximation of cooperative heterogeneous multi-agent reinforcement learning (MARL) using mean field control (MFC)
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
R. Carmona, M. Laurir̀e. and Z. Tan. (2023): Model-free mean-field reinforcement learning: mean-field MDP and mean-field Q-learning
2023
Cited alongside, same era.
R. Carmona and F. Delarue (2018a): Probabilistic Theory of Mean Field Games with Applications, Vol I. Springer
Cited in the paper.
R. Carmona and F. Delarue (2018b): Probabilistic Theory of Mean Field Games with Applications, Vol II. Springer
Cited in the paper.
2023
Closest in time.
2023
Closest in time.
M. Giegrich, C. Reisinger and Y. Zhang (2024): Convergence of policy gradient methods for finite-horizon exploratory linear-quadratic control problems
2024
Closest in time.