Fetching the paper…
Reading the bibliography…
We propose an actor-critic framework to solve the time-continuous stochastic optimal control problem.
Itô, S.: The fundamental solution of the parabolic equation in a differentiable manifold. Osaka Mathematical Journal 5
1953
Earlier work this paper cites.
Boltyanski, V., Gamkrelidze, R., Pontryagin, L.: On the theory of optimal processes. Dokl. Akad. Nauk SSSR 10
1956
Earlier work this paper cites.
Aronson, D.: The fundamental solution of a linear parabolic equation containing a small parameter. Illinois Journal of Mathematics 3
1959
Earlier work this paper cites.
Boltyanski, V., Gamkrelidze, R., Mishchenko, E., Pontryagin, L.: The maximum principle in the theory of optimal processes of control. IFAC Proceedings Volumes 1
1960
Earlier work this paper cites.
Kalman, R.E.: Contributions to the theory of optimal control. Bol. soc. mat. mexicana 5
1960
Earlier work this paper cites.
Bellman, R.: Dynamic programming. Science 153
1966
Earlier work this paper cites.
Aronson, D.G.: Bounds for the fundamental solution of a parabolic equation. Bulletin of the American Mathematical society 73
1967
Earlier work this paper cites.
Bobrow, J.E., Dubowsky, S., Gibson, J.S.: Time-optimal control of robotic manipulators along specified paths. The international journal of robotics research 4
1985
Earlier work this paper cites.
Ladyzenskaja, O.A., Solonnikov, V.A., Uralceva, N.N.: Linear and Quasi-linear Equations of Parabolic Type vol. 23. American Mathematical Soc. (1988)
1988
Earlier work this paper cites.
Kushner, H.J.: Numerical methods for stochastic control problems in continuous time. SIAM Journal on Control and Optimization 28
1990
Earlier work this paper cites.
Leonard, D., Van Long, N.: Optimal Control Theory and Static Optimization in Economics. Cambridge University Press (1992)
1992
Earlier work this paper cites.
Risken, H.: Fokker-planck equation. In: The Fokker-Planck Equation, pp. 63–95. Springer (1996)
1996
Earlier work this paper cites.
Bradtke, S.J., Barto, A.G.: Linear least-squares algorithms for temporal difference learning. Machine learning 22
1996
Earlier work this paper cites.
Munos, R., Bourgine, P.: Reinforcement learning for continuous stochastic control problems. Advances in Neural Information Processing Systems 10
1997
Earlier work this paper cites.
Yong, J., Zhou, X.Y.: Stochastic Controls: Hamiltonian Systems and HJB Equations vol. 43. Springer (1999)
1999
Earlier work this paper cites.
Konda, V., Tsitsiklis, J.: Actor-critic algorithms. Advances in Neural Information Processing Systems 12
1999
Earlier work this paper cites.
Sutton, R.S., McAllester, D., Singh, S., Mansour, Y.: Policy gradient methods for reinforcement learning with function approximation. Advances in Neural Information Processing Systems 12
1999
Earlier work this paper cites.
Forsyth, P.A., Labahn, G.: Numerical methods for controlled Hamilton–Jacobi–Bellman PDEs in finance. Journal of Computational Finance 11
2007
Earlier work this paper cites.
Krylov, N.V.: Controlled Diffusion Processes vol. 14. Springer (2008)
2008
Earlier work this paper cites.
Friedman, A.: Partial Differential Equations of Parabolic Type. Courier Dover Publications (2008)
2008
Earlier work this paper cites.
Stuart, A.M.: Inverse problems: a Bayesian perspective. Acta numerica 19
2010
Earlier work this paper cites.
Fleming, W.H., Rishel, R.W.: Deterministic and Stochastic Optimal Control vol. 1. Springer (2012)
2012
Earlier work this paper cites.
Longuski, J.M., Guzmán, J.J., Prussing, J.E.: Optimal Control with Aerospace Applications. Space Technology Library. Springer (2013)
2013
Earlier work this paper cites.
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. In: Bengio, Y., LeCun, Y. (eds.) 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings (2015)
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
Sutton, R.S., Barto, A.G.: Reinforcement Learning: An Introduction. MIT press (2018)
2018
Earlier work this paper cites.
Han, J., Jentzen, A., E, W.: Solving high-dimensional partial differential equations using deep learning. Proceedings of the National Academy of Sciences 115
2018
Cited alongside, same era.
Dalal, G., Thoppe, G., Szörényi, B., Mannor, S.: Finite sample analysis of two-timescale stochastic approximation with applications to reinforcement learning. In: Conference On Learning Theory, pp. 1199–1233 (2018). PMLR
2018
Cited alongside, same era.
Carmona, R., Delarue, F.: Probabilistic Theory of Mean Field Games with Applications I-II. Springer (2018)
2018
Cited alongside, same era.
Recht, B.: A tour of reinforcement learning: The view from continuous control. Annual Review of Control, Robotics, and Autonomous Systems 2
2019
Cited alongside, same era.
Mou, C.: Remarks on schauder estimates and existence of classical solutions for a class of uniformly parabolic Hamilton–Jacobi–Bellman integro-PDEs. Journal of Dynamics and Differential Equations 31
Onken, D., Nurbekyan, L., Li, X., Fung, S.W., Osher, S., Ruthotto, L.: A neural network approach for high-dimensional optimal control applied to multiagent path finding. IEEE Transactions on Control Systems Technology 31
2022
Later among the works it cites.
Jin, Z., Qiu, M., Tran, K., Yin, G.: A survey of numerical solutions for stochastic control problems: Some recent progress. Numerical Algebra, Control and Optimization 12
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2020
Cited alongside, same era.
Wang, H., Zariphopoulou, T., Zhou, X.Y.: Reinforcement learning in continuous time and space: A stochastic control approach. J. Mach. Learn. Res. 21
2020
Cited alongside, same era.
Domingo-Enrich, C., Jelassi, S., Mensch, A., Rotskoff, G., Bruna, J.: A mean-field analysis of two-player zero-sum games. Advances in Neural Information Processing Systems 33
2020
Cited alongside, same era.
Kerimkulov, B., Siska, D., Szpruch, L.: Exponential convergence and stability of Howard’s policy improvement algorithm for controlled diffusions. SIAM Journal on Control and Optimization 58
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Han, J., Lu, J., Zhou, M.: Solving high-dimensional eigenvalue problems using deep neural networks: A diffusion monte carlo like approach. Journal of Computational Physics 423
2020
Cited alongside, same era.
Gobet, E., Grangereau, M.: Newton method for stochastic control problems. SIAM Journal on Control and Optimization 60
2022
Later among the works it cites.
Firoozi, D., Jaimungal, S.: Exploratory LQG mean field games with entropy regularization. Automatica 139
2022
Later among the works it cites.
Guo, X., Xu, R., Zariphopoulou, T.: Entropy regularization for mean field games with learning. Mathematics of Operations Research (2022)
2022
Later among the works it cites.
Jia, Y., Zhou, X.Y.: Policy gradient and actor-critic learning in continuous time and space: Theory and algorithms. The Journal of Machine Learning Research 23
2022
Later among the works it cites.
Tang, W., Zhang, Y.P., Zhou, X.Y.: Exploratory HJB equations and their convergence. SIAM Journal on Control and Optimization 60
2022
Later among the works it cites.
Carmona, R., Laurière, M.: Convergence analysis of machine learning algorithms for the numerical solution of mean field control and games: Ii—the finite horizon case. The Annals of Applied Probability 32
2022
Later among the works it cites.
Laurière, M., Tangpi, L.: Convergence of large population games to mean field games with interaction through the controls. SIAM Journal on Mathematical Analysis 54
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Bokanowski, O., Prost, A., Warin, X.: Neural networks for first order hjb equations and application to front propagation with obstacle terms. Partial Differential Equations and Applications 4
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Darbon, J., Dower, P.M., Meng, T.: Neural network architectures using min-plus algebra for solving certain high-dimensional optimal control problems and Hamilton–Jacobi PDEs. Mathematics of Control, Signals, and Systems 35
2023
Later among the works it cites.
Laurière, M., Song, J., Tang, Q.: Policy iteration method for time-dependent mean field games systems with non-separable Hamiltonians. Applied Mathematics & Optimization 87
2023
Later among the works it cites.
2023
Later among the works it cites.
Zhou, M., Lu, J.: Single timescale actor-critic method to solve the linear quadratic regulator with convergence guarantees. Journal of Machine Learning Research 24
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.