Fetching the paper…
Reading the bibliography…
Preference optimization for diffusion models aims to align them with human preferences for images.
R. A. Bradley and M. E. Terry, “Rank analysis of incomplete block designs: I. the method of paired comparisons,” Biometrika , vol. 39, pp. 324–345, 1952
1952
Earlier work this paper cites.
D. P. Kingma, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114 , 2013
2013
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in MICCAI , 2015, pp. 234–241
2015
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in NeurIPS , vol. 33, 2020, pp. 6840–6851
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in ICML , vol. 139, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in ICLR , 2021
2021
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in CVPR , 2022, pp. 10 684–10 695
2022
Earlier work this paper cites.
J. Li, D. Li, C. Xiong, and S. Hoi, “BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in ICML , vol. 162, 2022, pp. 12 888–12 900
2022
Earlier work this paper cites.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” arXiv preprint arXiv:2207.12598 , 2022
2022
Earlier work this paper cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev, “LAION-5B: An open large-scale dataset for training next generation image-text models,” in NeurIPS , vol. 35, 2022, pp. 25 278–25 294
2022
Earlier work this paper cites.
E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in ICLR , 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” in NeurIPS , vol. 36, 2023, pp. 53 728–53 741
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in ICCV , 2023, pp. 4195–4205
2023
Earlier work this paper cites.
Y. Kirstain, A. Polyak, U. Singer, S. Matiana, J. Penna, and O. Levy, “Pick-a-Pic: An open dataset of user preferences for text-to-image generation,” in NeurIPS , vol. 36, 2023, pp. 36 652–36 663
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Xu, X. Liu, Y. Wu, Y. Tong, Q. Li, M. Ding, J. Tang, and Y. Dong, “ImageReward: Learning and evaluating human preferences for text-to-image generation,” in NeurIPS , vol. 36, 2023, pp. 15 903–15 935
2023
Cited alongside, same era.
Y. Fan, O. Watkins, Y. Du, H. Liu, M. Ryu, C. Boutilier, P. Abbeel, M. Ghavamzadeh, K. Lee, and K. Lee, “DPOK: Reinforcement learning for fine-tuning text-to-image diffusion models,” in NeurIPS , vol. 36, 2023, pp. 79 858–79 885
2023
Cited alongside, same era.
D. Ghosh, H. Hajishirzi, and L. Schmidt, “Geneval: An object-focused framework for evaluating text-to-image alignment,” in NeurIPS , vol. 36, 2023, pp. 52 132–52 152
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Later among the works it cites.
J. Zhang, J. Wu, Y. Ren, X. Xia, H. Kuang, P. Xie, J. Li, X. Xiao, W. Huang, S. Wen, L. Fu, and G. Li, “Unifl: Improve latent diffusion model via unified feedback learning,” in NeurIPS , vol. 37, 2024, pp. 67 355–67 382
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev, “Reproducible scaling laws for contrastive language-image learning,” in CVPR , 2023, pp. 2818–2829
2023
Cited alongside, same era.
Y. Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow matching for generative modeling,” in ICLR , 2023
2023
Cited alongside, same era.
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, D. Podell, T. Dockhorn, Z. English, and R. Rombach, “Scaling rectified flow transformers for high-resolution image synthesis,” in ICML , vol. 235, 2024, pp. 12 606–12 633
2024
Cited alongside, same era.
2024
Cited alongside, same era.
B. Wallace, M. Dang, R. Rafailov, L. Zhou, A. Lou, S. Purushwalkam, S. Ermon, C. Xiong, S. Joty, and N. Naik, “Diffusion model alignment using direct preference optimization,” in CVPR , 2024, pp. 8228–8238
2024
Cited alongside, same era.
K. Black, M. Janner, Y. Du, I. Kostrikov, and S. Levine, “Training diffusion models with reinforcement learning,” in ICLR , 2024
2024
Cited alongside, same era.
K. Yang, J. Tao, J. Lyu, C. Ge, J. Chen, W. Shen, X. Zhu, and X. Li, “Using human feedback to fine-tune diffusion models without any reward model,” in CVPR , 2024, pp. 8941–8951
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Later among the works it cites.
L. Eyring, S. Karthik, K. Roth, A. Dosovitskiy, and Z. Akata, “Reno: Enhancing one-step text-to-image models through reward-based noise optimization,” in NeurIPS , vol. 37, 2024, pp. 125 487–125 519
2024
Later among the works it cites.
B. F. Labs, “Flux,” https://github.com/black-forest-labs/flux , 2024
2024
Later among the works it cites.
“Learning to reason with llms,” OpenAI, Tech. Rep., 2024. [Online]. Available: https://openai.com/index/learning-to-reason-with-llms/
2024
Later among the works it cites.
2024
Later among the works it cites.
2025
Closest in time.
2025
Closest in time.
X. Wu, Y. Hao, M. Zhang, K. Sun, Z. Huang, G. Song, Y. Liu, and H. Li, “Deep reward supervisions for tuning text-to-image diffusion models,” in ECCV , 2025, pp. 108–124
2025
Closest in time.
X. Zhang, L. Yang, G. Li, Y. Cai, xie jiake, Y. Tang, Y. Yang, M. Wang, and B. CUI, “Itercomp: Iterative composition-aware feedback learning from model gallery for text-to-image generation,” in ICLR , 2025
2025
Closest in time.
S. Kim, M. Kim, and D. Park, “Test-time alignment of diffusion models without reward over-optimization,” in ICLR , 2025
2025
Closest in time.
Z. Lin, D. Pathak, B. Li, J. Li, X. Xia, G. Neubig, P. Zhang, and D. Ramanan, “Evaluating text-to-visual generation with image-to-text generation,” in ECCV , 2025, pp. 366–384
2025
Closest in time.
K. Huang, C. Duan, K. Sun, E. Xie, Z. Li, and X. Liu, “T2i-compbench++: An enhanced and comprehensive benchmark for compositional text-to-image generation,” IEEE TPAMI , vol. 47, pp. 3563–3579, 2025
2025
Closest in time.