Fetching the paper…
Reading the bibliography…
This paper studies stochastic control problems with the action space taken to be probability measures, with the objective penalised by the relative entropy.
Dynamic programming
R. Bellman · 1966
Earlier work this paper cites.
Linear and quasi-linear equations of parabolic type
O. A. Ladyzenskaja, V. A. Solonnikov, and N. N. Ural’ceva · 1968
Earlier work this paper cites.
Controlled diffusion processes
N. V. Krylov · 1980
Earlier work this paper cites.
Stochastic evolution equations
N. V. Krylov and B. L. Rozovskii · 1981
Earlier work this paper cites.
Continuous Exponential Martingales and BMO
N. Kazamaki · 1994
Earlier work this paper cites.
Dynamic programming and optimal control
D. P. Bertsekas · 1995
Earlier work this paper cites.
A computational fluid mechanics solution to the Monge–Kantorovich mass transfer problem
J.-D. Benamou and Y. Brenier · 2000
Earlier work this paper cites.
Reinforcement learning in continuous time and space
K. Doya · 2000
Earlier work this paper cites.
Generalization of an inequality by Talagrand and links with the logarithmic Sobolev inequality
F. Otto and C. Villani · 2000
Earlier work this paper cites.
Lectures on the calculus of variations and optimal control theory
L. C. Young · 2000
Earlier work this paper cites.
Stochastic control of partially observable systems
A. Bensoussan · 2004
Earlier work this paper cites.
Stochastic optimal control: the discrete-time case
D. P. Bertsekas and S. Shreve · 2004
Earlier work this paper cites.
Controlled Markov processes and viscosity solutions
W. H. Fleming and H. M. Soner · 2006
Earlier work this paper cites.
Optimal transport: old and new
C. Villani · 2008
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
B. D. Ziebart · 2010
Cited alongside, same era.
Applications of variational inequalities in stochastic control
A. Bensoussan and J.-L. Lions · 2011
Cited alongside, same era.
Central limit theorem for Markov processes with spectral gap in the Wasserstein metric
T. Komorowski and A. Walczuk · 2012
Cited alongside, same era.
Nonlinear stochastic evolution equations of second order with damping
E. Emmrich and D. Šiška · 2017
Cited alongside, same era.
Coupling and exponential ergodicity for stochastic differential equations driven by Lévy processes
M. B. Majka · 2017
Cited alongside, same era.
Backward stochastic differential equations
Exploration versus exploitation in reinforcement learning: a stochastic control approach
H. Wang, T. Zariphopoulou, and X. Y. Zhou · 2019
Later among the works it cites.
Exponential convergence and stability of Howard’s policy improvement algorithm for controlled diffusions
B. Kerimkulov, D. Šiška, and L. Szpruch · 2020
Closest in time.
Uniform in time weak propagation of chaos on the torus
F. Delarue and A. Tse · 2021
Closest in time.
Newton method for stochastic control problems
E. Gobet and M. Grangereau · 2021
Closest in time.
Mean-field Langevin dynamics and energy landscape of neural networks
K. Hu, Z. Ren, D. Šiška, and Ł. Szpruch · 2021
Closest in time.
A neural network-based policy iteration algorithm with global H 2 H^{2} -superlinear convergence for stochastic games on domains
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Zhang · 2017
Cited alongside, same era.
Probabilistic Theory of Mean Field Games with Applications I-II
R. Carmona and F. Delarue · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
A theory of regularized Markov decision processes
M. Geist, B. Scherrer, and O. Pietquin · 2019
Cited alongside, same era.
A stability approach for solving multidimensional quadratic BSDEs
J. Harter and A. Richou · 2019
Cited alongside, same era.
Mean-field Langevin system, optimal control and deep neural networks
K. Hu, A. Kazeykina, and Z. Ren · 2019
Cited alongside, same era.
Mean-field neural ODEs via relaxed optimal control
J.-F. Jabir, D. Šiška, and Ł. Szpruch · 2019
Cited alongside, same era.
K. Ito, C. Reisinger, and Y. Zhang · 2021
Closest in time.
A modified MSA for stochastic control problems
B. Kerimkulov, D. Šiška, and L. Szpruch · 2021
Closest in time.
Regularity and stability of feedback relaxed controls
C. Reisinger and Y. Zhang · 2021
Closest in time.
Weak quantitative propagation of chaos via differential calculus on the space of measures
J.-F. Chassagneux, L. Szpruch, and A. Tse · 2022
Closest in time.
Convergence of policy improvement for entropy-regularized stochastic control problems
Y.-J. Huang, Z. Wang, and Z. Zhou · 2022
Closest in time.
Exploratory HJB equations and their convergence
W. Tang, Y. P. Zhang, and X. Y. Zhou · 2022
Closest in time.
A learning scheme by sparse grids and Picard approximations for semilinear parabolic PDEs
J.-F. Chassagneux, J. Chen, N. Frikha, and C. Zhou · 2023
Closest in time.
Mirror descent for stochastic control problems with measure-valued controls
B. Kerimkulov, D. Šiška, Ł. Szpruch, and Y. Zhang · 2024
Closest in time.