Fetching the paper…
Reading the bibliography…
Reinforcement Learning from human feedback (RLHF) has been shown a promising direction for aligning generative models with human intent and has also been explored in recent works for alignment of diffusion generative models.
Multidimensional diffusion processes , volume 233 of Grundlehren der Mathematischen Wissenschaften
Stroock, D. W. and Varadhan, S. R. S · 1979
Earlier work this paper cites.
Brownian motion and stochastic calculus , volume 113 of Graduate Texts in Mathematics
Karatzas, I. and Shreve, S. E · 1991
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Stochastic controls: Hamiltonian systems and HJB equations , volume 43
Yong, J. and Zhou, X. Y · 1999
Earlier work this paper cites.
Exponential ergodicity of non-lipschitz stochastic differential equations
Zhang, X · 2009
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent, P · 2011
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Improving image captioning with better use of captions
Shi, Z., Zhou, X., Qiu, X., and Zhu, X · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Earlier work this paper cites.
Reinforcement learning in continuous time and space: A stochastic control approach
Wang, H., Zariphopoulou, T., and Zhou, X. Y · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
A variational perspective on diffusion-based generative models and score matching
Huang, C.-W., Lim, J. H., and Courville, A. C · 2021
Cited alongside, same era.
Variational diffusion models
Kingma, D., Salimans, T., Poole, B., and Ho, J · 2021
Cited alongside, same era.
Normalizing flows for probabilistic modeling and inference
Papamakarios, G., Nalisnick, E., Rezende, D. J., Mohamed, S., and Lakshminarayanan, B · 2021
Cited alongside, same era.
Sampling is as easy as learning the score: theory for diffusion models with minimal data assumptions
Chen, S., Chewi, S., Li, J., Li, Y., Salim, A., and Zhang, A. R · 2022
Cited alongside, same era.
Training-free structured diffusion guidance for compositional text-to-image synthesis
Feng, W., He, X., Fu, T.-J., Jampani, V., Akula, A., Narayana, P., Basu, S., Wang, X. E., and Wang, W. Y · 2022
Cited alongside, same era.
Diffusion policies as an expressive policy class for offline reinforcement learning
Wang, Z., Hunt, J. J., and Zhou, M · 2022
Later among the works it cites.
Geodiff: A geometric diffusion model for molecular conformation generation
Xu, M., Yu, L., Song, Y., Shi, C., Ermon, S., and Tang, J · 2022
Later among the works it cites.
Fast sampling of diffusion models with exponential integrator
Zhang, Q. and Chen, Y · 2022
Later among the works it cites.
gddim: Generalized denoising diffusion implicit models
Zhang, Q., Tao, M., and Chen, Y · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Benchmarking spatial relationships in text-to-image generation
Gokhale, T., Palangi, H., Nushi, B., Vineet, V., Horvitz, E., Kamar, E., Baral, C., and Yang, Y · 2022
Cited alongside, same era.
Optimizing prompts for text-to-image generation
Hao, Y., Chi, Z., Dong, L., and Wei, F · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al · 2022
Cited alongside, same era.
Achieving mean–variance efficiency by continuous-time reinforcement learning
Huang, Y., Jia, Y., and Zhou, X · 2022
Cited alongside, same era.
Planning with diffusion for flexible behavior synthesis
Janner, M., Du, Y., Tenenbaum, J. B., and Levine, S · 2022
Cited alongside, same era.
Flow straight and fast: Learning to generate and transfer data with rectified flow
Liu, X., Gong, C., and Liu, Q · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Black, K., Janner, M., Du, Y., Kostrikov, I., and Levine, S · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Later among the works it cites.
Dai, M., Dong, Y., Jia, Y., and Zhou, X. Y · 2023
Later among the works it cites.
Optimizing ddpm sampling with shortcut fine-tuning
Fan, Y. and Lee, K · 2023
Later among the works it cites.
Dpok: Reinforcement learning for fine-tuning text-to-image diffusion models
Fan, Y., Watkins, O., Du, Y., Liu, H., Ryu, M., Boutilier, C., Abbeel, P., Ghavamzadeh, M., Lee, K., and Lee, K · 2023
Later among the works it cites.
Aligning text-to-image models using human feedback
Lee, K., Liu, H., Ryu, M., Watkins, O., Du, Y., Boutilier, C., Abbeel, P., Ghavamzadeh, M., and Gu, S. S · 2023
Later among the works it cites.
Instaflow: One step is enough for high-quality diffusion-based text-to-image generation
Liu, X., Zhang, X., Ma, J., Peng, J., et al · 2023
Later among the works it cites.
Stable bias: Analyzing societal representations in diffusion models
Luccioni, A. S., Akiki, C., Mitchell, M., and Jernite, Y · 2023
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al · 2024
Closest in time.
Fine-tuning of diffusion models via stochastic control: entropy regularization and beyond
Tang, W · 2024
Closest in time.
Policy optimization for continuous reinforcement learning
Zhao, H., Tang, W., and Yao, D · 2024
Closest in time.
Scores as actions: fine-tuning diffusion models by continuous-time reinforcement learning
Zhao, H., Chen, H., Zhang, J., Yao, D., and Tang, W · 2024
Closest in time.