Fetching the paper…
Reading the bibliography…
In this paper we investigate the convergence of the Policy Iteration Algorithm (PIA) for a class of general continuous-time entropy-regularized stochastic control problems.
Bellman, R. (1955) Functional equations in the theory of dynamic programming
1955
Earlier work this paper cites.
Bellman, R. (1957) Dynamic Programming
1957
Earlier work this paper cites.
Howard, R.A. (1960) Dynamic programming and Markov processes
1960
Earlier work this paper cites.
Puterman, M. L., and Brumelle, S. L. (1979) On the convergence of policy iteration in stationary dynamic programming
1979
Earlier work this paper cites.
Puterman, M. L., (1981) On the convergence of policy iteration for controlled diffusions , J. Optim. Theory Appl. 33
1981
Earlier work this paper cites.
Bismut, J. M., (1984) Large Deviation and Malliavin Calculus
1984
Earlier work this paper cites.
Elworthy, K. D. and Li, X.M. (1994), Formulae for the derivatives of heat semigroups
1994
Earlier work this paper cites.
Krylov, N. V. (1996) Lectures on elliptic and parabolic equations in Hölder spaces
1996
Earlier work this paper cites.
2001
Earlier work this paper cites.
Ma, J. and Zhang, J. (2002), Representation Theorems for Backward Stochastic Differential Equations
2002
Earlier work this paper cites.
Nualart, D. (2006). Malliavin calculus and related topics
2006
Cited alongside, same era.
Bokanowski, O., Maroso, S. and Zidani, H., (2009) Some convergence results for Howard’s algorithm . SIAM Journal on Numerical Analysis
2009
Cited alongside, same era.
Jacka, S. and Mijatović, A., (2017) On the policy improvement algorithm in continuous time
2017
Cited alongside, same era.
Kerimkulov, B., Šiška, D., and Szpruch, L., (2020) Exponential convergence and stability of Howard’s policy improvement algorithm for controlled diffusions , SIAM J. Control Optim
2020
Cited alongside, same era.
Wang, H., Zariphopoulou, T. and Zhou, X.Y., (2020), Reinforcement learning in continuous time and space: a stochastic control approach , J. Mach. Learn. Res
2020
Cited alongside, same era.
Guo, X., Xu, R., and Zariphopoulou, T. (2022) Entropy regularization for mean field games with learning
2022
Later among the works it cites.
Tang, W., Zhang, Y. P., and Zhou, X. Y., (2022) Exploratory HJB equations and their convergence
2022
Later among the works it cites.
2023
Later among the works it cites.
Reisinger, C., Stockinger, W., and Zhang, Y. (2023) Linear convergence of a policy gradient method for some finite horizon continuous time control problems
2023
Later among the works it cites.
Dong, Y. (2024) Randomized optimal stopping problem in continuous time and reinforcement learning algorithm
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, H. and Zhou, X.Y., (2020), Continuous-time mean-variance portfolio selection: a reinforcement learning framework , Math. Finance
2020
Cited alongside, same era.
Ito, K., Reisinger, C. and Zhang, Y.,(2021), A neural network-based policy iteration algorithm with global H2-superlinear convergence for stochastic games on domains. Foundations of Computational Mathematics , 21(2),331-374
2021
Cited alongside, same era.
Kerimkulov, B., Šiška, D., and Szpruch, L., (2021) A modified MSA for stochastic control problems , Appl. Math. Optim
2021
Cited alongside, same era.
Reisinger, C. and Zhang, Y. (2021). Regularity and stability of feedback relaxed controls
2021
Cited alongside, same era.
Ma, J., Wang, G., Zhang, J., and Zhou, X.Y., Reinforcement Learning Algorithms for Entropy-Regularized HJB Equations with Model Uncertainty
Cited in the paper.
2024
Closest in time.
2024
Closest in time.
Huang, Y., Wang Z., and Zhou, Z. (2025), Convergence of Policy Iteration for Entropy-Regularized Stochastic Control Problems , SIAM J. Control Optim
2025
Closest in time.
Santos, M. S. and Rust, J.,(2004), Convergence properties of policy iteration . SIAM Journal on Control and Optimization, 42(6), 2094-2115
2094
Closest in time.