Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF), which aligns a diffusion model with input prompt, has become a crucial step in building reliable generative AI models.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Stochastic controls: Hamiltonian systems and HJB equations , volume 43
Yong, J. and Zhou, X. Y · 1999
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. M. and Langford, J · 2002
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Vincent, P · 2011
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Improving image captioning with better use of captions
Shi, Z., Zhou, X., Qiu, X., and Zhu, X · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Earlier work this paper cites.
Reinforcement learning in continuous time and space: A stochastic control approach
Wang, H., Zariphopoulou, T., and Zhou, X. Y · 2020
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Variational diffusion models
Kingma, D., Salimans, T., Poole, B., and Ho, J · 2021
Earlier work this paper cites.
Normalizing flows for probabilistic modeling and inference
Papamakarios, G., Nalisnick, E., Rezende, D. J., Mohamed, S., and Lakshminarayanan, B · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
On-policy deep reinforcement learning for the average-reward criterion
Zhang, Y. and Ross, K. W · 2021
Cited alongside, same era.
Sampling is as easy as learning the score: theory for diffusion models with minimal data assumptions
Chen, S., Chewi, S., Li, J., Li, Y., Salim, A., and Zhang, A. R · 2022
Cited alongside, same era.
Optimizing prompts for text-to-image generation
Hao, Y., Chi, Z., Dong, L., and Wei, F · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., et al · 2022
Optimizing ddpm sampling with shortcut fine-tuning
Fan, Y. and Lee, K · 2023
Later among the works it cites.
Dpok: Reinforcement learning for fine-tuning text-to-image diffusion models
Fan, Y., Watkins, O., Du, Y., Liu, H., Ryu, M., Boutilier, C., Abbeel, P., Ghavamzadeh, M., Lee, K., and Lee, K · 2023
Later among the works it cites.
q-learning in continuous time
Jia, Y. and Zhou, X. Y · 2023
Later among the works it cites.
Aligning text-to-image models using human feedback
Lee, K., Liu, H., Ryu, M., Watkins, O., Du, Y., Boutilier, C., Abbeel, P., Ghavamzadeh, M., and Gu, S. S · 2023
Later among the works it cites.
Instaflow: One step is enough for high-quality diffusion-based text-to-image generation
Liu, X., Zhang, X., Ma, J., Peng, J., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
Karras, T., Aittala, M., Aila, T., and Laine, S · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Cited alongside, same era.
Flow straight and fast: Learning to generate and transfer data with rectified flow
Liu, X., Gong, C., and Liu, Q · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E. L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al · 2022
Cited alongside, same era.
Aligning text-to-image diffusion models with reward backpropagation
Prabhudesai, M., Goyal, A., Pathak, D., and Fragkiadaki, K · 2023
Later among the works it cites.
Song, Y., Dhariwal, P., Chen, M., and Sutskever, I · 2023
Later among the works it cites.
Domingo-Enrich, C., Drozdzal, M., Karrer, B., and Chen, R. T · 2024
Later among the works it cites.
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al · 2024
Later among the works it cites.
Reward-directed score-based diffusion models via q-learning
Gao, X., Zha, J., and Zhou, X. Y · 2024
Later among the works it cites.
Diffusion policy policy optimization
Ren, A. Z., Lidard, J., Ankile, L. L., Simeonov, A., Agrawal, P., Majumdar, A., Burchfiel, B., Dai, H., and Simchowitz, M · 2024
Later among the works it cites.
Fine-tuning of diffusion models via stochastic control: entropy regularization and beyond
Tang, W · 2024
Later among the works it cites.
Score-based diffusion models via stochastic differential equations–a technical tutorial
Tang, W. and Zhao, H · 2024
Later among the works it cites.
Fine-tuning of continuous-time diffusion models as entropy-regularized control
Uehara, M., Zhao, Y., Black, K., Hajiramezanali, E., Scalia, G., Diamant, N. L., Tseng, A. M., Biancalani, T., and Levine, S · 2024
Later among the works it cites.
Diffusion model alignment using direct preference optimization
Wallace, B., Dang, M., Rafailov, R., Zhou, L., Lou, A., Purushwalkam, S., Ermon, S., Xiong, C., Joty, S., and Naik, N · 2024
Later among the works it cites.
Preference tuning with human feedback on language, speech, and vision tasks: A survey
Winata, G. I., Zhao, H., Das, A., Tang, W., Yao, D. D., Zhang, S.-X., and Sahu, S · 2024
Later among the works it cites.
Imagereward: Learning and evaluating human preferences for text-to-image generation
Xu, J., Liu, X., Wu, Y., Tong, Y., Li, Q., Ding, M., Tang, J., and Dong, Y · 2024
Later among the works it cites.
Maximum entropy inverse reinforcement learning of diffusion models with energy-based models
Yoon, S., Hwang, H., Kwon, D., Noh, Y.-K., and Park, F. C · 2024
Later among the works it cites.
Self-play fine-tuning of diffusion models for text-to-image generation
Yuan, H., Chen, Z., Ji, K., and Gu, Q · 2024
Later among the works it cites.