Linear-quadratic mean-field reinforcement learning: convergence of policy gradient methods
Original
R. Carmona, M. Lauriere, and Z. Tan · 1910
Earlier work this paper cites.
Model-free mean-field reinforcement learning: Mean-field MDP and mean-field Q-learning
Original
R. Carmona, M. Lauriere, and Z. Tan · 1910
Earlier work this paper cites.
Convergence of dynamic programming models
H.J. Langen · 1981
Earlier work this paper cites.
Convergence of Stochastic Processes
D. Pollard · 1984
Earlier work this paper cites.
Adaptive Markov Control Processes
O. Hernández-Lerma · 1989
Earlier work this paper cites.
Recurrence conditions for Markov decision processes with Borel state space: a survey
O. Hernández-Lerma, R. Montes-De-Oca, and R. Cavazos-Cadena · 1991
Earlier work this paper cites.
Discrete-Time Markov Control Processes: Basic Optimality Criteria
O. Hernández-Lerma and J.B. Lasserre · 1996
Earlier work this paper cites.
Further Topics on Discrete-Time Markov Control Processes
O. Hernández-Lerma and J.B. Lasserre · 1999
Earlier work this paper cites.
Perturbation Analysis of Optimization Problems
J.F. Bonnans and A. Shapiro · 2000
Earlier work this paper cites.
Q-learning in regularized mean-field games
Original
B. Anahtarci, C.D. Kariksiz, and N. Saldi · 2003
Earlier work this paper cites.
Oblivious equilibrium: A mean field approximation for large-scale dynamic games
Gabriel Weintraub, Lanier Benkard, and Benjamin Van Roy · 2005
Earlier work this paper cites.
Infinite Dimensional Analysis
C.D. Aliprantis and K.C. Border · 2006
Earlier work this paper cites.
Ergodic properties of Markov processes
M. Hairer · 2006
Earlier work this paper cites.
Large population stochastic dynamic games: Closed loop McKean-Vlasov systems and the Nash certainty equivalence principle
M. Huang, R.P. Malhamé, and P.E. Caines · 2006
Earlier work this paper cites.
Large-population cost coupled LQG problems with nonuniform agents: Individual-mass behavior and decentralized ϵ \epsilon -Nash equilibria
M. Huang, P.E. Caines, and R.P. Malhamé · 2007
Earlier work this paper cites.
Mean field games
J. Lasry and P.Lions · 2007
Earlier work this paper cites.
Concentration inequalities for dependent random variables via the martingale method
L. Kontorovich and K. Ramanan · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
R. Munos and C. Szepesvári · 2008
Earlier work this paper cites.