Fetching the paper…
Reading the bibliography…
Diffusion models have become a popular choice for representing actor policies in behavior cloning and offline reinforcement learning.
Sur la théorie du mouvement brownien
Langevin, P · 1908
Earlier work this paper cites.
Dynamic programming and a new formalism in the calculus of variations
Bellman, R · 1954
Earlier work this paper cites.
Martingale approach to some limit theorems
Papanicolaou, G · 1977
Earlier work this paper cites.
Reinforcement learning methods for continuous-time markov decision problems
Bradtke, S. and Duff, M · 1994
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Reinforcement learning in continuous time and space
Doya, K · 2000
Earlier work this paper cites.
Optimal control theory: an introduction
Kirk, D. E · 2004
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Hyvärinen, A. and Dayan, P · 2005
Earlier work this paper cites.
Reinforcement learning with a gaussian mixture model
Agostini, A. and Celaya, E · 2010
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
Deterministic and stochastic optimal control , volume 1
Fleming, W. H. and Rishel, R. W · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Kober, J., Bagnell, J. A., and Peters, J · 2013
Earlier work this paper cites.
Analysis and geometry of Markov diffusion operators , volume 103
Bakry, D., Gentil, I., Ledoux, M., et al · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., et al · 2017
Cited alongside, same era.
Model-based action exploration for learning dynamic motion skills
Berseth, G., Kyriazis, A., Zinin, I., Choi, W., and van de Panne, M · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods, 2018
Fujimoto, S., van Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Umap: Uniform manifold approximation and projection for dimension reduction
McInnes, L., Healy, J., and Melville, J · 2018
Planning with diffusion for flexible behavior synthesis
Janner, M., Du, Y., Tenenbaum, J., and Levine, S · 2022
Later among the works it cites.
Policy gradient and actor-critic learning in continuous time and space: Theory and algorithms
Jia, Y. and Zhou, X. Y · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
Diffusion policies as an expressive policy class for offline reinforcement learning
Wang, Z., Hunt, J. J., and Zhou, M · 2022
Later among the works it cites.
Novel view synthesis with diffusion models
Watson, D., Chan, W., Martin-Brualla, R., Ho, J., Tagliasacchi, A., and Norouzi, M · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The limits and potentials of deep learning for robotics
Sünderhauf, N., Brock, O., Scheirer, W., Hadsell, R., Fox, D., Leitner, J., Upcroft, B., Abbeel, P., Burgard, W., Milford, M., et al · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Superhuman ai for multiplayer poker
Brown, N. and Sandholm, T · 2019
Cited alongside, same era.
Continuous control with deep reinforcement learning, 2019
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2019
Cited alongside, same era.
How to learn a useful critic? model-based action-gradient-estimator policy optimization
D’Oro, P. and Jaśkowski, W · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Cited alongside, same era.
dm control: Software and tasks for continuous control
Tunyasuvunakool, S., Muldal, A., Doron, Y., Liu, S., Bohez, S., Merel, J., Erez, T., Lillicrap, T., Heess, N., and Tassa, Y · 2020
Cited alongside, same era.
Daydreamer: World models for physical robot learning
Wu, P., Escontrela, A., Hafner, D., Goldberg, K., and Abbeel, P · 2022
Later among the works it cites.
Training diffusion models with reinforcement learning
Black, K., Janner, M., Du, Y., Kostrikov, I., and Levine, S · 2023
Closest in time.
Idql: Implicit q-learning as an actor-critic method with diffusion policies
Hansen-Estruch, P., Kostrikov, I., Janner, M., Kuba, J. G., and Levine, S · 2023
Closest in time.
q-learning in continuous time
Jia, Y. and Zhou, X. Y · 2023
Closest in time.
Efficient diffusion policies for offline reinforcement learning
Kang, B., Ma, X., Du, C., Pang, T., and Yan, S · 2023
Closest in time.
Lu, C., Chen, H., Chen, J., Su, H., Li, C., and Zhu, J · 2023
Closest in time.
Goal-conditioned imitation learning using score-based diffusion policies
Reuss, M., Li, M., Jia, X., and Lioutikov, R · 2023
Closest in time.
Fighting uncertainty with gradients: Offline reinforcement learning via diffusion score matching
Suh, H., Chou, G., Dai, H., Yang, L., Gupta, A., and Tedrake, R · 2023
Closest in time.
Policy optimization for continuous reinforcement learning
Zhao, H., Tang, W., and Yao, D · 2024
Closest in time.