Fetching the paper…
Reading the bibliography…
We study policy gradient (PG) for reinforcement learning in continuous time and space under the regularized exploratory formulation developed by Wang et al.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Myopia and inconsistency in dynamic utility maximization
Strotz, R. H. (1955) · 1955
Earlier work this paper cites.
Stochastic optimization
Aleksandrov, V., Sysoev, V., and Shemeneva, V. (1968) · 1968
Earlier work this paper cites.
Neuronlike adaptive elements that can solve difficult learning control problems
Barto, A. G., Sutton, R. S., and Anderson, C. W. (1983) · 1983
Earlier work this paper cites.
Ergodic control of multidimensional diffusions I: The existence results
Borkar, V. S. and Ghosh, M. K. (1988) · 1988
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S. (1988) · 1988
Earlier work this paper cites.
Ergodic control of multidimensional diffusions II: Adaptive control
Borkar, V. S. and Ghosh, M. K. (1990) · 1990
Earlier work this paper cites.
Likelihood ratio gradient estimation for stochastic systems
Glynn, P. W. (1990) · 1990
Earlier work this paper cites.
On Bellman equations of ergodic control in
Bensoussan, A. and Frehse, J. (1992) · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Advantage updating
Baird, L. C. (1993) · 1993
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J. and Barto, A. G. (1996) · 1996
Earlier work this paper cites.
Reinforcement learning for continuous stochastic control problems
Munos, R. and Bourgine, P. (1997) · 1997
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y. (1999) · 1999
Earlier work this paper cites.
Stochastic Controls: Hamiltonian Systems and HJB Equations
Yong, J. and Zhou, X. Y. (1999) · 1999
Earlier work this paper cites.
Reinforcement learning in continuous time and space
Doya, K. (2000) · 2000
Earlier work this paper cites.
Optimal dynamic portfolio selection: Multiperiod mean-variance formulation
Li, D. and Ng, W.-L. (2000) · 2000
Earlier work this paper cites.
Comparing policy-gradient algorithms
Sutton, R. S., Singh, S., and McAllester, D. (2000) · 2000
Earlier work this paper cites.
Continuous-time mean-variance portfolio selection: A stochastic LQ framework
Zhou, X. Y. and Li, D. (2000) · 2000
Earlier work this paper cites.
Simulation-based optimization of Markov reward processes
Marbach, P. and Tsitsiklis, J. N. (2001) · 2001
Earlier work this paper cites.
Basei, M., Guo, X., Hu, A., and Zhang, Y. (2020) · 2006
Earlier work this paper cites.
Being serious about non-commitment: subgame perfect equilibrium in continuous time
Ekeland, I. and Lazrak, A. (2006) · 2006
Earlier work this paper cites.
Controlled Markov Processes and Viscosity Solutions
Fleming, W. H. and Soner, H. M. (2006) · 2006
Cited alongside, same era.
Policy gradient in continuous time
Munos, R. (2006) · 2006
Cited alongside, same era.
Learning to control a 6-degree-of-freedom walking robot
Wawrzynski, P. (2007) · 2007
Cited alongside, same era.
A convergent o(n) temporal-difference algorithm for off-policy learning with linear function approximation
Sutton, R. S., Szepesvári, C., and Maei, H. R. (2008) · 2008
Cited alongside, same era.
Natural actor–critic algorithms
Bhatnagar, S., Sutton, R. S., Ghavamzadeh, M., and Lee, M. (2009) · 2009
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
Maei, H. R., Szepesvari, C., Bhatnagar, S., Precup, D., Silver, D., and Sutton, R. S. (2009) · 2009
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. (2017) · 2017
Later among the works it cites.
A deeper look at experience replay
Zhang, S. and Sutton, R. S. (2017) · 2017
Later among the works it cites.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al. (2018) · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Later among the works it cites.
Making deep Q-learning methods robust to time discretization
Tallec, C., Blier, L., and Ollivier, Y. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Sutton, R. S., Maei, H. R., Precup, D., Bhatnagar, S., Silver, D., Szepesvári, C., and Wiewiora, E. (2009) · 2009
Cited alongside, same era.
Dynamic mean-variance asset allocation
Basak, S. and Chabakauri, G. (2010) · 2010
Cited alongside, same era.
Online actor–critic algorithm to solve the continuous-time infinite horizon optimal control problem
Vamvoudakis, K. G. and Lewis, F. L. (2010) · 2010
Cited alongside, same era.
Analysis and improvement of policy gradient estimation
Zhao, T., Hachiya, H., Niu, G., and Sugiyama, M. (2011) · 2011
Cited alongside, same era.
Ergodic control of diffusion processes
Arapostathis, A., Borkar, V. S., and Ghosh, M. K. (2012) · 2012
Cited alongside, same era.
Model-free reinforcement learning with continuous action in practice
Degris, T., Pilarski, P. M., and Sutton, R. S. (2012) · 2012
Cited alongside, same era.
Traffic-signal control reinforcement learning approach for continuous-time Markov games
Aragon-Gómez, R. and Clempner, J. B. (2020) · 2020
Later among the works it cites.
Learning equilibrium mean-variance strategy
Dai, M., Dong, Y., and Jia, Y. (2020) · 2020
Later among the works it cites.
Revisiting fundamentals of experience replay
Fedus, W., Ramachandran, P., Agarwal, R., Bengio, Y., Larochelle, H., Rowland, M., and Dabney, W. (2020) · 2020
Later among the works it cites.
Reinforcement learning in continuous time and space: A stochastic control approach
Wang, H., Zariphopoulou, T., and Zhou, X. Y. (2020) · 2020
Later among the works it cites.
Continuous-time mean–variance portfolio selection: A reinforcement learning framework
Wang, H. and Zhou, X. Y. (2020) · 2020
Later among the works it cites.
On nonlinear Feynman–Kac formulas for viscosity solutions of semilinear parabolic partial differential equations
Beck, C., Hutzenthaler, M., and Jentzen, A. (2021) · 2021
Closest in time.
A dynamic mean-variance analysis for log returns
Dai, M., Jin, H., Kou, S., and Xu, Y. (2021) · 2021
Closest in time.
Hamilton-Jacobi deep Q-Learning for deterministic continuous-time systems with Lipschitz continuous controls
Kim, J., Shin, J., and Yang, I. (2021) · 2021
Closest in time.
Deep reinforcement learning for autonomous driving: A survey
Kiran, B. R., Sobh, I., Talpaert, V., Mannion, P., Al Sallab, A. A., Yogamani, S., and Pérez, P. (2021) · 2021
Closest in time.
Policy iterations for reinforcement learning problems in continuous time and space—Fundamental theory and methods
Lee, J. and Sutton, R. S. (2021) · 2021
Closest in time.
Szpruch, L., Treetanthiploet, T., and Zhang, Y. (2021) · 2021
Closest in time.
Exploratory HJB equations and their convergence
Tang, W., Zhang, P. Y., and Zhou, X. Y. (2021) · 2021
Closest in time.
Global convergence of policy gradient for linear-quadratic mean-field control/game in continuous time
Wang, W., Han, J., Yang, Z., and Wang, Z. (2021) · 2021
Closest in time.
Continuous-time model-based reinforcement learning
Yildiz, C., Heinonen, M., and Lähdesmäki, H. (2021) · 2021
Closest in time.
State-dependent temperature control for Langevin diffusions
Gao, X., Xu, Z. Q., and Zhou, X. Y. (2022) · 2022
Closest in time.
Entropy regularization for mean field games with learning
Guo, X., Xu, R., and Zariphopoulou, T. (2022) · 2022
Closest in time.