Fetching the paper…
Reading the bibliography…
We study the policy iteration algorithm (PIA) for entropy-regularized stochastic control problems on an infinite time horizon with a large discount rate, focusing on two main scenarios.
Dynamic programming
R. Bellman · 1957
Earlier work this paper cites.
Dynamic programming and markov processes
R. A. Howard · 1960
Earlier work this paper cites.
On the convergence of policy iteration in stationary dynamic programming
M. L. Puterman and S. L. Brumelle · 1979
Earlier work this paper cites.
On the convergence of policy iteration for controlled diffusions
M. L. Puterman · 1981
Earlier work this paper cites.
N. V. Krylov, Nonlinear Elliptic and Parabolic Equations of the Second Order
1987
Earlier work this paper cites.
User’s guide to viscosity solutions of second order partial differential equations
M. G. Crandall, H. Ishii, and P.-L. Lions · 1992
Earlier work this paper cites.
A Weak Convergence Approach to the Theory of Large Deviations
P. Dupuis and R. S. Ellis · 1997
Earlier work this paper cites.
Second order elliptic equations and elliptic systems
Y.-Z. Chen and L.-C. Wu · 1998
Earlier work this paper cites.
Partial differential equations
L. C. Evans · 1998
Earlier work this paper cites.
Nearly optimal control laws for nonlinear systems with saturating actuators using a neural network HJB approach
M. Abu-Khalaf and F. L. Lewis · 2005
Earlier work this paper cites.
Controlled diffusion processes
N. V. Krylov · 2008
Earlier work this paper cites.
Adaptive optimal control for continuous-time linear systems based on policy iteration
D. Vrabie, O. Pastravanu, M. Abu-Khalaf, and F. L. Lewis · 2009
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 2014
Cited alongside, same era.
Value and policy iterations in optimal control and adaptive dynamic programming
D. P. Bertsekas · 2015
Cited alongside, same era.
On the policy improvement algorithm in continuous time
S. D. Jacka and A. Mijatović · 2017
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Exponential convergence and stability of Howard’s policy improvement algorithm for controlled diffusions
B. Kerimkulov, D. Šiška, and L. Szpruch · 2020
Curse of optimality, and how we break it
X. Y. Zhou · 2021
Later among the works it cites.
Rates of convergence for the policy iteration method for mean field games systems
F. Camilli and Q. Tang · 2022
Later among the works it cites.
Randomized optimal stopping problem in continuous time and reinforcement learning algorithm
Y. Dong · 2022
Later among the works it cites.
Exploratory LQG mean field games with entropy regularization
D. Firoozi and S. Jaimungal · 2022
Later among the works it cites.
Entropy regularization for mean field games with learning
X. Guo, R. Xu, and T. Zariphopoulou · 2022
Later among the works it cites.
Convergence of policy improvement for entropy-regularized stochastic control problems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reinforcement learning in continuous time and space: A stochastic control approach
H. Wang, T. Zariphopoulou, and X. Y. Zhou · 2020
Cited alongside, same era.
Continuous-time mean–variance portfolio selection: A reinforcement learning framework
H. Wang and X. Y. Zhou · 2020
Cited alongside, same era.
A policy iteration method for mean field games
S. Cacace, F. Camilli, and A. Goffi · 2021
Cited alongside, same era.
Policy iterations for reinforcement learning problems in continuous time and space—fundamental theory and methods
J. Lee and R. S. Sutton · 2021
Cited alongside, same era.
Hamilton-Jacobi equations—theory and applications
H. V. Tran · 2021
Cited alongside, same era.
Hamilton-Jacobi Based Policy-Iteration via Deep Operator Learning
J. Y. Lee, Y. Kim
Cited in the paper.
Y.-J. Huang, Z. Wang, and Z. Zhou · 2022
Later among the works it cites.
Exploratory HJB equations and their convergence
W. Tang, Y. P. Zhang, and X. Y. Zhou · 2022
Later among the works it cites.
Learning equilibrium mean‐variance strategy
M. Dai, Y. Dong, and Y. Jia · 2023
Later among the works it cites.
Actor-critic learning for mean-field control in continuous time
N. Frikha, M. Germain, M. Laurière, H. Pham, and X. Song · 2023
Later among the works it cites.
Continuous-time q-learning for mean-field control problems
X. Wei and X. Yu · 2023
Later among the works it cites.
Regret of exploratory policy improvement and q q -learning
W. Tang and X. Y. Zhou · 2024
Closest in time.