Fetching the paper…
Reading the bibliography…
For a general entropy-regularized stochastic control problem on an infinite horizon, we prove that a policy iteration algorithm (PIA) converges to an optimal relaxed control.
A mathematical theory of communication
C. E. Shannon · 1948
Earlier work this paper cites.
Information theory and statistical mechanics
E. T. Jaynes · 1957
Earlier work this paper cites.
Information theory and statistical mechanics. II
E. T. Jaynes · 1957
Earlier work this paper cites.
On the convergence of policy iteration for controlled diffusions
M. L. Puterman · 1981
Earlier work this paper cites.
Second order elliptic equations and elliptic systems
Ya-Zhe Chen and Lan-Cheng Wu · 1991
Earlier work this paper cites.
Brownian motion and stochastic calculus
Ioannis Karatzas and Steven E. Shreve · 1991
Earlier work this paper cites.
Partial differential equations
Lawrence C. Evans · 1998
Earlier work this paper cites.
Elliptic partial differential equations of second order
David Gilbarg and Neil S. Trudinger · 1998
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D. Ziebart, Andrew L. Maas, J. Andrew Bagnell, and Anind K. Dey · 2008
Earlier work this paper cites.
Higher chain formula proved by combinatorics
Tsoy-Wo Ma · 2009
Earlier work this paper cites.
Modeling interaction via the principle of maximum causal entropy
Brian D. Ziebart, J. Andrew Bagnell, and Anind K. Dey · 2010
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2016
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
On the policy improvement algorithm in continuous time
Saul D. Jacka and Aleksandar Mijatović · 2017
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Exponential convergence and stability of Howard’s policy improvement algorithm for controlled diffusions
Value iteration in continuous actions, states and time
Michael Lutter, Shie Mannor, Jan Peters, Dieter Fox, and Animesh Garg · 2021
Later among the works it cites.
Regularity and stability of feedback relaxed controls
Christoph Reisinger and Yufei Zhang · 2021
Later among the works it cites.
Exploratory LQG mean field games with entropy regularization
Dena Firoozi and Sebastian Jaimungal · 2022
Closest in time.
Entropy regularization for mean field games with learning
Xin Guo, Renyuan Xu, and Thaleia Zariphopoulou · 2022
Closest in time.
Exploratory hjb equations and their convergence
Wenpin Tang, Yuming Paul Zhang, and Xun Yu Zhou · 2022
Closest in time.
Learning equilibrium mean-variance strategy
Min Dai, Yuchao Dong, and Yanwei Jia · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bekzhan Kerimkulov, David Šiška, and Lukasz Szpruch · 2020
Cited alongside, same era.
Reinforcement learning in continuous time and space: a stochastic control approach
Haoran Wang, Thaleia Zariphopoulou, and Xun Yu Zhou · 2020
Cited alongside, same era.
Continuous-time mean-variance portfolio selection: a reinforcement learning framework
Haoran Wang and Xun Yu Zhou · 2020
Cited alongside, same era.
A modified MSA for stochastic control problems
B. Kerimkulov, D. Šiška, and L. Szpruch · 2021
Cited alongside, same era.
Policy iterations for reinforcement learning problems in continuous time and space—fundamental theory and methods
Jaeyoung Lee and Richard S Sutton · 2021
Cited alongside, same era.
Closest in time.
q-learning in continuous time
Yanwei Jia and Xun Yu Zhou · 2023
Closest in time.
Policy iteration for the deterministic control problems–a viscosity approach
Wenpin Tang, Hung Vinh Tran, and Yuming Paul Zhang · 2023
Closest in time.
Continuous-time reinforcement learning control: A review of theoretical results, insights on performance, and needs for new designs
Brent A Wallace and Jennie Si · 2023
Closest in time.