Fetching the paper…
Reading the bibliography…
We consider policy gradient methods for stochastic optimal control problem in continuous time.
The fundamental solution of the parabolic equation in a differentiable manifold
Seizô Itô · 1953
Earlier work this paper cites.
The fundamental solution of a linear parabolic equation containing a small parameter
DG Aronson · 1959
Earlier work this paper cites.
Pontryagin maximum principle
Richard E Kopp · 1962
Earlier work this paper cites.
Bounds for the fundamental solution of a parabolic equation
Donald Gary Aronson · 1967
Earlier work this paper cites.
Shadow prices and duality for a class of optimal control problems
Jean-Pierre Aubin and FH Clarke · 1979
Earlier work this paper cites.
Method of successive approximations for solution of optimal control problems
Felix L Chernousko and AA Lyubushin · 1982
Earlier work this paper cites.
Linear and quasi-linear equations of parabolic type
Olga A Ladyzenskaja, Vsevolod Alekseevich Solonnikov, and Nina N Uralceva · 1988
Earlier work this paper cites.
Numerical methods for stochastic control problems in continuous time
Harold J Kushner · 1990
Earlier work this paper cites.
A general stochastic maximum principle for optimal control problems
Shige Peng · 1990
Earlier work this paper cites.
Policy evaluation
James J Heckman · 1992
Earlier work this paper cites.
Backward stochastic differential equations and quasilinear parabolic partial differential equations
Etienne Pardoux and Shige Peng · 1992
Earlier work this paper cites.
Fokker-planck equation
Hannes Risken · 1996
Earlier work this paper cites.
Optimal control and viscosity solutions of Hamilton-Jacobi-Bellman equations
Martino Bardi and Italo Capuzzo Dolcetta · 1997
Earlier work this paper cites.
Actor-critic algorithms
Vijay Konda and John Tsitsiklis · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Stochastic controls: Hamiltonian systems and HJB equations
Jiongmin Yong and Xun Yu Zhou · 1999
Earlier work this paper cites.
Optimal and hierarchical controls in dynamic stochastic manufacturing systems: A survey
Suresh P Sethi, Houmin Yan, Hanqin Zhang, and Qing Zhang · 2002
Earlier work this paper cites.
Consistency of generalized finite difference schemes for the stochastic hjb equation
J Frédéric Bonnans and Housnaa Zidani · 2003
Earlier work this paper cites.
Policy gradient in continuous time
Rémi Munos · 2006
Earlier work this paper cites.
A stochastic control model for optimal timing of climate policies
Olivier Bahn, Alain Haurie, and Roland Malhamé · 2008
Earlier work this paper cites.
Partial differential equations of parabolic type
Avner Friedman · 2008
Cited alongside, same era.
Continuous-time stochastic control and optimization with financial applications
Huyên Pham · 2009
Cited alongside, same era.
Dynamic programming and optimal control: Volume I
Dimitri Bertsekas · 2012
Cited alongside, same era.
Deterministic and stochastic optimal control
Wendell H Fleming and Raymond W Rishel · 2012
Cited alongside, same era.
Partial differential equations. 3, Nonlinear equations
Michael Eugene Taylor · 2012
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the Polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
Signatured deep fictitious play for mean field games with common noise
Ming Min and Ruimeng Hu · 2021
Later among the works it cites.
Global convergence of policy gradient for linear-quadratic mean-field control/game in continuous time
Weichen Wang, Jiequn Han, Zhuoran Yang, and Zhaoran Wang · 2021
Later among the works it cites.
Actor-critic method for high dimensional static Hamilton–Jacobi–Bellman partial differential equations based on neural networks
Mo Zhou, Jiequn Han, and Jianfeng Lu · 2021
Later among the works it cites.
Curse of optimality, and how we break it
Xun Yu Zhou · 2021
Later among the works it cites.
Convergence analysis of machine learning algorithms for the numerical solution of mean field control and games: Ii—the finite horizon case
René Carmona and Mathieu Laurière · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dgm: A deep learning algorithm for solving partial differential equations
Justin Sirignano and Konstantinos Spiliopoulos · 2018
Cited alongside, same era.
Remarks on schauder estimates and existence of classical solutions for a class of uniformly parabolic Hamilton–Jacobi–Bellman integro-PDEs
Chenchen Mou · 2019
Cited alongside, same era.
A mean-field analysis of two-player zero-sum games
Carles Domingo-Enrich, Samy Jelassi, Arthur Mensch, Grant Rotskoff, and Joan Bruna · 2020
Cited alongside, same era.
Deep learning methods for mean field control problems with delay
Jean-Pierre Fouque and Zhaoyu Zhang · 2020
Cited alongside, same era.
Solving high-dimensional eigenvalue problems using deep neural networks: A diffusion monte carlo like approach
Jiequn Han, Jianfeng Lu, and Mo Zhou · 2020
Cited alongside, same era.
Exponential convergence and stability of Howard’s policy improvement algorithm for controlled diffusions
Bekzhan Kerimkulov, David Siska, and Lukasz Szpruch · 2020
Cited alongside, same era.
Later among the works it cites.
Exploratory LQG mean field games with entropy regularization
Dena Firoozi and Sebastian Jaimungal · 2022
Later among the works it cites.
Michael Giegrich, Christoph Reisinger, and Yufei Zhang · 2022
Later among the works it cites.
Newton method for stochastic control problems
Emmanuel Gobet and Maxime Grangereau · 2022
Later among the works it cites.
Entropy regularization for mean field games with learning
Xin Guo, Renyuan Xu, and Thaleia Zariphopoulou · 2022
Later among the works it cites.
Solving stochastic optimal control problem via stochastic maximum principle with deep learning method
Shaolin Ji, Shige Peng, Ying Peng, and Xichuan Zhang · 2022
Later among the works it cites.
Mean field control problems for vaccine distribution
Wonjun Lee, Siting Liu, Wuchen Li, and Stanley Osher · 2022
Later among the works it cites.
A neural network approach for high-dimensional optimal control applied to multiagent path finding
Derek Onken, Levon Nurbekyan, Xingjian Li, Samy Wu Fung, Stanley Osher, and Lars Ruthotto · 2022
Later among the works it cites.
Christoph Reisinger, Wolfgang Stockinger, and Yufei Zhang · 2022
Later among the works it cites.
The modified MSA, a gradient flow and convergence
Deven Sethi and David Šiška · 2022
Later among the works it cites.
Exploratory HJB equations and their convergence
Wenpin Tang, Yuming Paul Zhang, and Xun Yu Zhou · 2022
Later among the works it cites.
A machine learning enhanced algorithm for the optimal landing problem
Yaohua Zang, Jihao Long, Xuanxi Zhang, Wei Hu, E Weinan, and Jiequn Han · 2022
Later among the works it cites.
Recent developments in machine learning methods for stochastic control and games
Ruimeng Hu and Mathieu Lauriere · 2023
Closest in time.
Single timescale actor-critic method to solve the linear quadratic regulator with convergence guarantees
Mo Zhou and Jianfeng Lu · 2023
Closest in time.
Mo Zhou and Jianfeng Lu · 2024
Closest in time.
Mo Zhou, Stanley Osher, and Wuchen Li · 2024
Closest in time.