Fetching the paper…
Reading the bibliography…
Text-to-image diffusion models are typically trained to optimize the log-likelihood objective, which presents challenges in meeting specific requirements for downstream tasks, such as image aesthetics and image-text alignment.
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., Ganguli, S.: Deep unsupervised learning using nonequilibrium thermodynamics. In: Proc. Int. Conf. Mach. Learn. (2015)
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
Brock, A., Donahue, J., Simonyan, K.: Large scale gan training for high fidelity natural image synthesis. Int. Conf. Learn. Represent. (2019)
2019
Earlier work this paper cites.
Karras, T., Laine, S., Aila, T.: A style-based generator architecture for generative adversarial networks. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 4401–4410 (2019)
2019
Earlier work this paper cites.
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., Bengio, Y.: Generative adversarial networks. Communications of the ACM 63
2020
Earlier work this paper cites.
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. In: Adv. Neural Inform. Process. Syst. (2020)
2020
Earlier work this paper cites.
Song, Y., Ermon, S.: Improved techniques for training score-based generative models. In: Adv. Neural Inform. Process. Syst. (2020)
2020
Earlier work this paper cites.
Alam, M.M., Islam, M.T., Rahman, S.M.M.: Unified learning approach for egocentric hand gesture recognition and fingertip detection. Pattern Recognition (2021). https://doi.org/10.1016/j.patcog.2021.108200
2021
Earlier work this paper cites.
Dhariwal, P., Nichol, A.: Diffusion models beat gans on image synthesis. In: Adv. Neural Inform. Process. Syst. (2021)
2021
Earlier work this paper cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision. In: Proc. Int. Conf. Mach. Learn. pp. 8748–8763 (2021)
2021
Earlier work this paper cites.
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., Sutskever, I.: Zero-shot text-to-image generation. In: Proc. Int. Conf. Mach. Learn. (2021)
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
Song, J., Meng, C., Ermon, S.: Denoising diffusion implicit models. Int. Conf. Learn. Represent. (2021)
2021
Earlier work this paper cites.
Song, Y., Durkan, C., Murray, I., Ermon, S.: Maximum likelihood training of score-based diffusion models. Adv. Neural Inform. Process. Syst. 34
2021
Earlier work this paper cites.
Chen, C., Mo, J.: IQA-PyTorch: Pytorch toolbox for image quality assessment. [Online]. Available: https://github.com/chaofengc/IQA-PyTorch (2022)
2022
Earlier work this paper cites.
Christoph, S., Romain, B.: Laion-aesthetics. (2022)
2022
Earlier work this paper cites.
Ho, J., Salimans, T.: Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598 (2022)
2022
Earlier work this paper cites.
Hu, E.J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W.: Lora: Low-rank adaptation of large language models. Int. Conf. Learn. Represent. (2022)
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al.: Training language models to follow instructions with human feedback. In: Adv. Neural Inform. Process. Syst. (2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: IEEE Conf. Comput. Vis. Pattern Recog. (2022)
2022
Cited alongside, same era.
2023
Closest in time.
Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., Aberman, K.: Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 22500–22510 (2023)
2023
Closest in time.
Su, S., Lin, H., Hosu, V., Wiedemann, O., Sun, J., Zhu, Y., Liu, H., Zhang, Y., Saupe, D.: Going the extra mile in face image quality assessment: A novel database and model. IEEE Trans. Multimedia (2023)
2023
Closest in time.
Wu, J.Z., Ge, Y., Wang, X., Lei, S.W., Gu, Y., Shi, Y., Hsu, W., Shan, Y., Qie, X., Shou, M.Z.: Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In: Int. Conf. Comput. Vis. pp. 7623–7633 (2023)
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S.K.S., Ayan, B.K., Mahdavi, S.S., Lopes, R.G., et al.: Photorealistic text-to-image diffusion models with deep language understanding. In: Adv. Neural Inform. Process. Syst. (2022)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A.H., Chechik, G., Cohen-Or, D.: An image is worth one word: Personalizing text-to-image generation using textual inversion. Int. Conf. Learn. Represent. (2023)
2023
Cited alongside, same era.
Hao, Y., Chi, Z., Dong, L., Wei, F.: Optimizing prompts for text-to-image generation. Adv. Neural Inform. Process. Syst. (2023)
2023
Cited alongside, same era.
Wu, X., Sun, K., Zhu, F., Zhao, R., Li, H.: Better aligning text-to-image models with human preference. Int. Conf. Comput. Vis. (2023)
2023
Closest in time.
Xu, J., Liu, X., Wu, Y., Tong, Y., Li, Q., Ding, M., Tang, J., Dong, Y.: Imagereward: Learning and evaluating human preferences for text-to-image generation. Adv. Neural Inform. Process. Syst. (2023)
2023
Closest in time.
Zhang, L., Rao, A., Agrawala, M.: Adding conditional control to text-to-image diffusion models. In: Int. Conf. Comput. Vis. (2023)
2023
Closest in time.
Black, K., Janner, M., Du, Y., Kostrikov, I., Levine, S.: Training diffusion models with reinforcement learning. Int. Conf. Learn. Represent. (2024)
2024
Closest in time.
Chen, C., Mo, J., Hou, J., Wu, H., Liao, L., Sun, W., Yan, Q., Lin, W.: Topiq: A top-down approach from semantics to distortions for image quality assessment. In: IEEE Trans. Image Process. (2024)
2024
Closest in time.
Clark, K., Vicol, P., Swersky, K., Fleet, D.J.: Directly fine-tuning diffusion models on differentiable rewards. Int. Conf. Learn. Represent. (2024)
2024
Closest in time.
Fan, Y., Watkins, O., Du, Y., Liu, H., Ryu, M., Boutilier, C., Abbeel, P., Ghavamzadeh, M., Lee, K., Lee, K.: Dpok: Reinforcement learning for fine-tuning text-to-image diffusion models. Adv. Neural Inform. Process. Syst. (2024)
2024
Closest in time.
Guo, Y., Yang, C., Rao, A., Wang, Y., Qiao, Y., Lin, D., Dai, B.: Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. Int. Conf. Learn. Represent. (2024)
2024
Closest in time.
Kirstain, Y., Polyak, A., Singer, U., Matiana, S., Penna, J., Levy, O.: Pick-a-pic: An open dataset of user preferences for text-to-image generation. Adv. Neural Inform. Process. Syst. (2024)
2024
Closest in time.
Li, X., Hou, X., Loy, C.C.: When stylegan meets stable diffusion: a 𝒲 + \mathcal{W}_{+} adapter for personalized image generation. IEEE Conf. Comput. Vis. Pattern Recog. (2024)
2024
Closest in time.
Li, Y., Liu, X., Kag, A., Hu, J., Idelbayev, Y., Sagar, D., Wang, Y., Tulyakov, S., Ren, J.: Textcraftor: Your text encoder can be image quality controller. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 7985–7995 (2024)
2024
Closest in time.
Mou, C., Wang, X., Xie, L., Wu, Y., Zhang, J., Qi, Z., Shan, Y., Qie, X.: T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. AAAI (2024)
2024
Closest in time.
Wu, H., Zhang, Z., Zhang, E., Chen, C., Liao, L., Wang, A., Li, C., Sun, W., Yan, Q., Zhai, G., et al.: Q-bench: A benchmark for general-purpose foundation models on low-level vision. Int. Conf. Learn. Represent. (2024)
2024
Closest in time.
Wu, H., Zhang, Z., Zhang, E., Chen, C., Liao, L., Wang, A., Xu, K., Li, C., Hou, J., Zhai, G., Xue, G., Sun, W., Yan, Q., Lin, W.: Q-instruct: Improving low-level visual abilities for multi-modality foundation models. IEEE Conf. Comput. Vis. Pattern Recog. (2024)
2024
Closest in time.