Fetching the paper…
Reading the bibliography…
We propose a new reinforcement learning (RL) formulation for training continuous-time score-based diffusion models for generative AI to generate samples that maximize reward functions while keeping the generated distributions close to the unknown target data distributions.
Reverse-time diffusion equation models
Anderson, B. D. O. (1982) · 1982
Earlier work this paper cites.
Time reversal of diffusions
Haussmann, U. G. and E. Pardoux (1986) · 1986
Earlier work this paper cites.
Stochastic approximation and recursive algorithms
Kushner, H. and G. Yin (2003) · 2003
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Hyvärinen, A. and P. Dayan (2005) · 2005
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. (2009) · 2009
Earlier work this paper cites.
Divergence estimation for multidimensional densities via k k -nearest-neighbor distances
Wang, Q., S. R. Kulkarni, and S. Verdú (2009) · 2009
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent, P. (2011) · 2011
Earlier work this paper cites.
Adaptive algorithms and stochastic approximations
Benveniste, A., M. Métivier, and P. Priouret (2012) · 2012
Earlier work this paper cites.
Stochastic controls: Hamiltonian systems and HJB equations
Yong, J. and X. Y. Zhou (2012) · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and J. Ba (2015) · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., P. Fischer, and T. Brox (2015) · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., P. Moritz, S. Levine, M. Jordan, and P. Abbeel (2015) · 2015
Earlier work this paper cites.
GANs trained by a two time-scale update rule converge to a local Nash equilibrium
Heusel, M., H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter (2017) · 2017
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., A. Zhou, P. Abbeel, and S. Levine (2018) · 2018
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y. and S. Ermon (2019) · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., A. Jain, and P. Abbeel (2020) · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., C. Meng, and S. Ermon (2020) · 2020
Earlier work this paper cites.
Sliced score matching: A scalable approach to density and score estimation
Song, Y., S. Garg, J. Shi, and S. Ermon (2020) · 2020
Cited alongside, same era.
Reinforcement learning in continuous time and space: A stochastic control approach
Wang, H., T. Zariphopoulou, and X. Y. Zhou (2020) · 2020
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J., D. Russo, and R. Singal (2021) · 2021
Cited alongside, same era.
Diffusion models beat GANs on image synthesis
Dhariwal, P. and A. Nichol (2021) · 2021
Cited alongside, same era.
Maximum likelihood training of score-based diffusion models
Song, Y., C. Durkan, I. Murray, and S. Ermon (2021) · 2021
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations
Song, Y., J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole (2021) · 2021
Time reversal of diffusion processes under a finite entropy condition
Cattiaux, P., G. Conforti, I. Gentil, and C. Léonard (2023) · 2023
Later among the works it cites.
Reinforcement learning for fine-tuning text-to-image diffusion models
Fan, Y., O. Watkins, Y. Du, H. Liu, M. Ryu, C. Boutilier, P. Abbeel, M. Ghavamzadeh, K. Lee, and K. Lee (2023) · 2023
Later among the works it cites.
q-learning in continuous time
Jia, Y. and X. Y. Zhou (2023) · 2023
Later among the works it cites.
A theory of continuous generative flow networks
Lahlou, S., T. Deleu, P. Lemos, D. Zhang, A. Volokhova, A. Hernández-Garcıa, L. N. Ezzine, Y. Bengio, and N. Malkin (2023) · 2023
Later among the works it cites.
FP-Diffusion: Improving score-based diffusion models by enforcing the underlying score fokker-planck equation
Lai, C.-H., Y. Takida, N. Murata, T. Uesaka, Y. Mitsufuji, and S. Ermon (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Classifier-free diffusion guidance
Ho, J. and T. Salimans (2022) · 2022
Cited alongside, same era.
Equivariant diffusion for molecule generation in 3d
Hoogeboom, E., V. G. Satorras, C. Vignac, and M. Welling (2022) · 2022
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
Karras, T., M. Aittala, T. Aila, and S. Laine (2022) · 2022
Cited alongside, same era.
Convergence for score-based generative modeling with polynomial complexity
Lee, H., J. Lu, and Y. Tan (2022) · 2022
Cited alongside, same era.
DPM-solver: A fast ODE solver for diffusion probabilistic model sampling in around 10 steps
Lu, C., Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu (2022) · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with CLIP latents
Ramesh, A., P. Dhariwal, A. Nichol, C. Chu, and M. Chen (2022) · 2022
Cited alongside, same era.
Lee, K., H. Liu, M. Ryu, O. Watkins, Y. Du, C. Boutilier, P. Abbeel, M. Ghavamzadeh, and S. S. Gu (2023) · 2023
Later among the works it cites.
Diffusion models: A comprehensive survey of methods and applications
Yang, L., Z. Zhang, Y. Song, S. Hong, R. Xu, Y. Zhao, Y. Shao, W. Zhang, B. Cui, and M.-H. Yang (2023) · 2023
Later among the works it cites.
Fast sampling of diffusion models with exponential integrator
Zhang, Q. and Y. Chen (2023) · 2023
Later among the works it cites.
Training diffusion models with reinforcement learning
Black, K., M. Janner, Y. Du, I. Kostrikov, and S. Levine (2024) · 2024
Closest in time.
An overview of diffusion models: Applications, guided generation, statistical rates and optimization
Chen, M., S. Mei, J. Fan, and M. Wang (2024) · 2024
Closest in time.
Directly fine-tuning diffusion models on differentiable rewards
Clark, K., P. Vicol, K. Swersky, and D. J. Fleet (2024) · 2024
Closest in time.
Fine-tuning of diffusion models via stochastic control: entropy regularization and beyond
Tang, W. (2024) · 2024
Closest in time.
Fine-tuning of continuous-time diffusion models as entropy-regularized control
Uehara, M., Y. Zhao, K. Black, E. Hajiramezanali, G. Scalia, N. L. Diamant, A. M. Tseng, T. Biancalani, and S. Levine (2024) · 2024
Closest in time.
Improving GFlowNets for text-to-image diffusion alignment
Zhang, D., Y. Zhang, J. Gu, R. Zhang, J. Susskind, N. Jaitly, and S. Zhai (2024) · 2024
Closest in time.
Zhao, H., C. Haoxian, J. Zhang, D. Yao, and W. Tang (2024) · 2024
Closest in time.
Accuracy of discretely sampled stochastic policies in continuous-time reinforcement learning
Jia, Y., D. Ouyang, and Y. Zhang (2025) · 2025
Closest in time.
Erratum to “q-learning in continuous time”
Jia, Y. and X. Y. Zhou (2025) · 2025
Closest in time.