Fetching the paper…
Reading the bibliography…
Text-to-Image (T2I) diffusion models have achieved remarkable success in image generation.
Clipscore: A reference-free evaluation metric for image captioning
Hessel, J.; Holtzman, A.; Forbes, M.; Bras, R. L.; and Choi, Y. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Earlier work this paper cites.
Improving image generation with better captions
Betker, J.; Goh, G.; Jing, L.; Brooks, T.; Wang, J.; Li, L.; Ouyang, L.; Zhuang, J.; Lee, J.; Guo, Y.; et al. 2023 · 2023
Earlier work this paper cites.
Training diffusion models with reinforcement learning
Black, K.; Janner, M.; Du, Y.; Kostrikov, I.; and Levine, S. 2023 · 2023
Earlier work this paper cites.
Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis
Feng, W.; He, X.; Fu, T.-J.; Jampani, V.; Akula, A.; Narayana, P.; Basu, S.; Wang, X. E.; and Wang, W. Y. 2023 · 2023
Earlier work this paper cites.
Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering
Hu, Y.; Liu, B.; Kasai, J.; Wang, Y.; Ostendorf, M.; Krishna, R.; and Smith, N. A. 2023 · 2023
Earlier work this paper cites.
VILA: Learning Image Aesthetics from User Comments with Vision-Language Pretraining
Ke, J.; Ye, K.; Yu, J.; Wu, Y.; Milanfar, P.; and Yang, F. 2023 · 2023
Earlier work this paper cites.
BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Earlier work this paper cites.
Visual Instruction Tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023 · 2023
Cited alongside, same era.
Introducing ChatGPT
OpenAI. 2023 · 2023
Cited alongside, same era.
Introducing Gemini: our largest and most capable AI model
Pichai, S.; and Hassabis, D. 2023 · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; Müller, J.; Penna, J.; and Rombach, R. 2023 · 2023
Cited alongside, same era.
Mp5: A multi-modal open-ended embodied system in minecraft via active perception
Qin, Y.; Zhou, E.; Liu, Q.; Yin, Z.; Sheng, L.; Zhang, R.; Qiao, Y.; and Shao, J. 2023 · 2023
Cited alongside, same era.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Wu, X.; Hao, Y.; Sun, K.; Chen, Y.; Zhu, F.; Zhao, R.; and Li, H. 2023 · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q.; Xu, H.; Xu, G.; Ye, J.; Yan, M.; Zhou, Y.; Wang, J.; Hu, A.; Shi, P.; Shi, Y.; et al. 2023 · 2023
Later among the works it cites.
Text-to-image diffusion model in generative ai: A survey
Zhang, C.; Zhang, C.; Zhang, M.; and Kweon, I. S. 2023 · 2023
Later among the works it cites.
Adding conditional control to text-to-image diffusion models
Zhang, L.; Rao, A.; and Agrawala, M. 2023 · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rafailov, R.; Sharma, A.; Mitchell, E.; Ermon, S.; Manning, C. D.; and Finn, C. 2023 · 2023
Cited alongside, same era.
Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment
Rassin, R.; Hirsch, E.; Glickman, D.; Ravfogel, S.; Goldberg, Y.; and Chechik, G. 2023 · 2023
Cited alongside, same era.
A picture is worth a thousand words: Principled recaptioning improves image generation
Segalis, E.; Valevski, D.; Lumen, D.; Matias, Y.; and Leviathan, Y. 2023 · 2023
Cited alongside, same era.
DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback
Sun, J.; Fu, D.; Hu, Y.; Wang, S.; Rassin, R.; Juan, D.-C.; Alon, D.; Herrmann, C.; van Steenkiste, S.; Krishna, R.; and Rashtchian, C. 2023 · 2023
Cited alongside, same era.
Diffusion Model Alignment Using Direct Preference Optimization
Wallace, B.; Dang, M.; Rafailov, R.; Zhou, L.; Lou, A.; Purushwalkam, S.; Ermon, S.; Xiong, C.; Joty, S.; and Naik, N. 2023 · 2023
Cited alongside, same era.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023a
Cited in the paper.
Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack
Dai, X.; Hou, J.; Ma, C.-Y.; Tsai, S.; Wang, J.; Wang, R.; Zhang, P.; Vandenhende, S.; Wang, X.; Dubey, A.; Yu, M.; Kadian, A.; Radenovic, F.; Mahajan, D.; Li, K.; Zhao, Y.; Petrovic, V.; Singh, M. K.; Motwani, S.; Wen, Y.; Song, Y.; Sumbaly, R.; Ramanathan, V.; He, Z.; Vajda, P.; and Parikh, D. 2023b
Cited in the paper.
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Later among the works it cites.
Reinforcement learning for fine-tuning text-to-image diffusion models
Fan, Y.; Watkins, O.; Du, Y.; Liu, H.; Ryu, M.; Boutilier, C.; Abbeel, P.; Ghavamzadeh, M.; Lee, K.; and Lee, K. 2024 · 2024
Closest in time.
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Kirstain, Y.; Polyak, A.; Singer, U.; Matiana, S.; Penna, J.; and Levy, O. 2024 · 2024
Closest in time.
Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
Li, D.; Kamko, A.; Akhgari, E.; Sabet, A.; Xu, L.; and Doshi, S. 2024 · 2024
Closest in time.
WorldSimBench: Towards Video Generation Models as World Simulators
Qin, Y.; Shi, Z.; Yu, J.; Wang, X.; Zhou, E.; Li, L.; Yin, Z.; Liu, X.; Sheng, L.; Shao, J.; et al. 2024 · 2024
Closest in time.
Laion-aesthetics
Schuhmann, C. 2022 · 2024
Closest in time.