Fetching the paper…
Reading the bibliography…
The practical applications of diffusion models have been limited by the misalignment between generated images and corresponding text prompts.
Denoising diffusion probabilistic models, 2020
Ho, J., Jain, A., and Abbeel, P · 2006
Earlier work this paper cites.
Improved techniques for training score-based generative models, 2020
Song, Y. and Ermon, S · 2006
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2010
Earlier work this paper cites.
Denoising diffusion implicit models, 2022
Song, J., Meng, C., and Ermon, S · 2010
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation, 2015
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics, 2015
Sohl-Dickstein, J., Weiss, E. A., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Visual relationship detection with language priors, 2016
Lu, C., Krishna, R., Bernstein, M., and Fei-Fei, L · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Diffusion models beat gans on image synthesis, 2021
Dhariwal, P. and Nichol, A · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Openclip, July 2021
Ilharco, G., Wortsman, M., Wightman, R., Gordon, C., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., and Schmidt, L · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision, 2021
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback, 2022
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., et al · 2022
Earlier work this paper cites.
Prompt-to-prompt image editing with cross attention control, 2022
Hertz, A., Mokady, R., Tenenbaum, J., Aberman, K., Pritch, Y., and Cohen-Or, D · 2022
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning, 2022
Hessel, J., Holtzman, A., Forbes, M., Bras, R. L., and Choi, Y · 2022
Earlier work this paper cites.
Classifier-free diffusion guidance, 2022
Ho, J. and Salimans, T · 2022
Earlier work this paper cites.
Null-text inversion for editing real images using guided diffusion models, 2022
Mokady, R., Hertz, A., Aberman, K., Pritch, Y., and Cohen-Or, D · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., et al · 2022
Earlier work this paper cites.
Dreamfusion: Text-to-3d using 2d diffusion, 2022
Poole, B., Jain, A., Barron, J. T., and Mildenhall, B · 2022
Earlier work this paper cites.
Diffusion autoencoders: Toward a meaningful and decodable representation, 2022
Preechakul, K., Chatthee, N., Wizadwongsa, S., and Suwajanakorn, S · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2022
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Plug-and-play diffusion features for text-driven image-to-image translation, 2022
Tumanyan, N., Geyer, M., Bagon, S., and Dekel, T · 2022
Cited alongside, same era.
Geodiff: a geometric diffusion model for molecular conformation generation, 2022
Xu, M., Yu, L., Song, Y., Shi, C., Ermon, S., and Tang, J · 2022
Cited alongside, same era.
Instructpix2pix: Learning to follow image editing instructions, 2023
Hallucination of multimodal large language models: A survey, 2024
Bai, Z., Wang, P., Xiao, T., He, T., Han, Z., Zhang, Z., and Shou, M. Z · 2024
Later among the works it cites.
Training diffusion models with reinforcement learning, 2024
Black, K., Janner, M., Du, Y., Kostrikov, I., and Levine, S · 2024
Later among the works it cites.
Controllable generation with text-to-image diffusion models: A survey, 2024
Cao, P., Zhou, F., Song, Q., and Yang, L · 2024
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion, 2024
Chi, C., Xu, Z., Feng, S., Cousineau, E., Du, Y., Burchfiel, B., Tedrake, R., and Song, S · 2024
Later among the works it cites.
Directly fine-tuning diffusion models on differentiable rewards, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brooks, T., Holynski, A., and Efros, A. A · 2023
Cited alongside, same era.
Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing, 2023
Cao, M., Wang, X., Qi, Z., Shan, Y., Qie, X., and Zheng, Y · 2023
Cited alongside, same era.
Dpok: Reinforcement learning for fine-tuning text-to-image diffusion models, 2023
Fan, Y., Watkins, O., Du, Y., Liu, H., Ryu, M., Boutilier, C., Abbeel, P., Ghavamzadeh, M., Lee, K., and Lee, K · 2023
Cited alongside, same era.
Imagic: Text-based real image editing with diffusion models, 2023
Kawar, B., Zada, S., Lang, O., Tov, O., Chang, H., Dekel, T., Mosseri, I., and Irani, M · 2023
Cited alongside, same era.
Aligning text-to-image models using human feedback, 2023
Lee, K., Liu, H., Ryu, M., Watkins, O., Du, Y., Boutilier, C., Abbeel, P., Ghavamzadeh, M., and Gu, S. S · 2023
Cited alongside, same era.
Scalable diffusion models with transformers, 2023
Peebles, W. and Xie, S · 2023
Cited alongside, same era.
Quantitatively measuring and contrastively exploring heterogeneity for domain generalization
Tong, Y., Yuan, J., Zhang, M., Zhu, D., Zhang, K., Wu, F., and Kuang, K · 2023
Cited alongside, same era.
How many unicorns are in this image? a safety evaluation benchmark for vision llms, 2023
Tu, H., Cui, C., Wang, Z., Zhou, Y., Zhao, B., Han, J., Zhou, W., Yao, H., and Xie, C · 2023
Cited alongside, same era.
Clark, K., Vicol, P., Swersky, K., and Fleet, D. J · 2024
Later among the works it cites.
A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., and Liu, T · 2024
Later among the works it cites.
Comat: Aligning text-to-image diffusion model with image-to-text concept matching, 2024
Jiang, D., Song, G., Wu, X., Zhang, R., Shen, D., Zong, Z., Liu, Y., and Li, H · 2024
Later among the works it cites.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment, 2024
Liu, Y., Yao, Y., Ton, J.-F., Zhang, X., Guo, R., Cheng, H., Klochkov, Y., Taufiq, M. F., and Li, H · 2024
Later among the works it cites.
Fast prompt alignment for text-to-image generation, 2024
Mrini, K., Lu, H., Yang, L., Huang, W., and Wang, H · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model, 2024
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2024
Later among the works it cites.
Dragdiffusion: Harnessing diffusion models for interactive point-based image editing, 2024
Shi, Y., Xue, C., Liew, J. H., Pan, J., Yan, H., Zhang, W., Tan, V. Y. F., and Bai, S · 2024
Later among the works it cites.
Attngcg: Enhancing jailbreaking attacks on llms with attention manipulation, 2024
Wang, Z., Tu, H., Mei, J., Zhao, B., Wang, Y., and Xie, C · 2024
Later among the works it cites.
Yu, T., Yao, Y., Zhang, H., He, T., Han, Y., Cui, G., Hu, J., Liu, Z., Zheng, H.-T., Sun, M., and Chua, T.-S · 2024
Later among the works it cites.
Layoutdiffusion: Controllable diffusion model for layout-to-image generation, 2024
Zheng, G., Zhou, X., Li, X., Qi, Z., Shan, Y., and Li, X · 2024
Later among the works it cites.
Hu, Z., Zhang, F., Chen, L., Kuang, K., Li, J., Gao, K., Xiao, J., Wang, X., and Zhu, W · 2025
Closest in time.
Sdpo: Segment-level direct preference optimization for social agents, 2025
Kong, A., Ma, W., Zhao, S., Li, Y., Wu, Y., Wang, K., Liu, X., Li, Q., Qin, Y., and Huang, F · 2025
Closest in time.
Yang, J., Jin, D., Tang, A., Shen, L., Zhu, D., Chen, Z., Wang, D., Cui, Q., Zhang, Z., Zhou, J., et al · 2025
Closest in time.
Remedy: Recipe merging dynamics in large vision-language models
Zhu, D., Song, Y., Shen, T., Zhao, Z., Yang, J., Zhang, M., and Wu, C · 2025
Closest in time.