Fetching the paper…
Reading the bibliography…
We study the convergence of $q$-learning and related algorithms introduced by Jia and Zhou (J.
Dynamic programming
R. Bellman · 1957
Earlier work this paper cites.
Dynamic programming and Markov processes
R. A. Howard · 1960
Earlier work this paper cites.
Partial differential equations of parabolic type
A. Friedman · 1964
Earlier work this paper cites.
Survey of measurable selection theorems
D. H. Wagner · 1977
Earlier work this paper cites.
Stochastic filtering theory
G. Kallianpur · 1980
Earlier work this paper cites.
Counterexamples in probability
J. M. Stoyanov · 1987
Earlier work this paper cites.
Adaptive algorithms and stochastic approximations
A. Benveniste, M. Métivier, and P. Priouret · 1990
Earlier work this paper cites.
A general stochastic maximum principle for optimal control problems
S. Peng · 1990
Earlier work this paper cites.
Formulae for the derivatives of heat semigroups
K. D. Elworthy and X.-M. Li · 1994
Earlier work this paper cites.
Backward stochastic differential equations in finance
N. El Karoui, S. Peng, and M. C. Quenez · 1997
Earlier work this paper cites.
Learning agents for uncertain environments
S. Russell · 1998
Earlier work this paper cites.
Continuous martingales and Brownian motion
D. Revuz and M. Yor · 1999
Earlier work this paper cites.
Stochastic controls – Hamiltonian systems and HJB equations
J. Yong and X. Y. Zhou · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A. Ng and S. Russell · 2000
Earlier work this paper cites.
Controlled Markov processes and viscosity solutions
W. H. Fleming and H. M. Soner · 2006
Earlier work this paper cites.
Second-order backward stochastic differential equations and fully nonlinear parabolic PDEs
P. Cheridito, H. M. Soner, N. Touzi, and N. Victoir · 2007
Earlier work this paper cites.
G G -expectation, G G -Brownian motion and related stochastic calculus of Itô type
S. Peng · 2007
Earlier work this paper cites.
Continuous-time stochastic control and optimization with financial applications
H. Pham · 2009
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, and G. Ostrovski · 2015
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, and M. Lanctot · 2016
Cited alongside, same era.
Learning to navigate in complex environments
P. Mirowski, R. Pascanu, F. Viola, H. Soyer, A. J. Ballard, A. Banino, M. Denil, R. Goroshin, L. Sifre, and K. Kavukcuoglu · 2017
Cited alongside, same era.
Mastering the game of Go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, and A. Bolton · 2017
Cited alongside, same era.
Reinforcement learning: an introduction
R. S. Sutton and A. G. Barto · 2018
Cited alongside, same era.
Exponential convergence and stability of Howard’s policy improvement algorithm for controlled diffusions
Policy optimization for continuous reinforcement learning
H. Zhao, W. Tang, and D. D. Yao · 2023
Later among the works it cites.
The curse of optimality, and how to break it?
X. Y. Zhou · 2023
Later among the works it cites.
On the grid-sampling limit SDE
C. Bender and N. T. Thuan · 2024
Closest in time.
Reinforcement learning for fine-tuning text-to-image diffusion models
Y. Fan, O. Watkins, Y. Du, H. Liu, M. Ryu, C. Boutilier, P. Abbeel, M. Ghavamzadeh, K. Lee, and K. Lee · 2024
Closest in time.
Mean–variance portfolio selection by continuous-time reinforcement learning: Algorithms, regret analysis, and empirical study
Y. Huang, Y. Jia, and X. Y. Zhou · 2024
Closest in time.
Entropy annealing for policy mirror descent in continuous time and space
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Kerimkulov, D. Šiška, and L. Szpruch · 2020
Cited alongside, same era.
On the local Lipschitz stability of Bayesian inverse problems
B. Sprungk · 2020
Cited alongside, same era.
Reinforcement learning in continuous time and space: A stochastic control approach
H. Wang, T. Zariphopoulou, and X. Y. Zhou · 2020
Cited alongside, same era.
Continuous-time mean-variance portfolio selection: a reinforcement learning framework
H. Wang and X. Y. Zhou · 2020
Cited alongside, same era.
Log-concave sampling
S. Chewi · 2021
Cited alongside, same era.
Convergence of policy improvement for entropy-regularized stochastic control problems
Y.-J. Huang, Z. Wang, and Z. Zhou · 2022
Cited alongside, same era.
Policy evaluation and temporal-difference learning in continuous time and space: A martingale approach
Y. Jia and X. Y. Zhou · 2022
Cited alongside, same era.
D. Sethi, D. Šiška, and Y. Zhang · 2024
Closest in time.
Optimal scheduling of entropy regularizer for continuous-time linear-quadratic reinforcement learning
L. Szpruch, T. Treetanthiploet, and Y. Zhang · 2024
Closest in time.
H. Zhao, H. Chen, J. Zhang, D. D. Yao, and W. Tang · 2024
Closest in time.
Reward-directed score-based diffusion models via q q -learning
X. Gao, J. Zha, and X. Y. Zhou · 2025
Closest in time.
Sublinear regret for a class of continuous-time linear-quadratic reinforcement learning problems
Y. Huang, Y. Jia, and X. Y. Zhou · 2025
Closest in time.
Y. Huang and X. Y. Zhou · 2025
Closest in time.
Erratum to “q-learning in continuous time”
Y. Jia and X. Y. Zhou · 2025
Closest in time.
Understanding sampler stochasticity in training diffusion models for RLHF
J. Sheng, H. Zhao, H. Chen, D. D. Yao, and W. Tang · 2025
Closest in time.
Policy iteration for the deterministic control problems—a viscosity approach
W. Tang, H. V. Tran, and Y. P. Zhang · 2025
Closest in time.
Score as action: Fine tuning diffusion generative models by continuous-time reinforcement learning
H. Zhao, H. Chen, J. Zhang, D. Yao, and W. Tang · 2025
Closest in time.
ART for diffusion sampling: continuous-time control and actor-critic learning
Y. Huang, W. Tang, and X. Y. Zhou · 2026
Closest in time.
Accuracy of discretely sampled stochastic policies in continuous-time reinforcement learning
Y. Jia, D. Ouyang, and Y. Zhang · 2026
Closest in time.
Convergence analysis for entropy-regularized control problems: a probabilistic approach
J. Ma, G. Wang, and J. Zhang · 2026
Closest in time.