Fetching the paper…
Reading the bibliography…
Story visualization has gained increasing attention in artificial intelligence.
Multi-concept customization of text-to-image diffusion
Kumari, N.; Zhang, B.; Zhang, R.; Shechtman, E.; and Zhu, J.-Y. 2023 · 1941
Earlier work this paper cites.
Image retrieval using scene graphs
Johnson, J.; Krishna, R.; Stark, M.; Li, L.-J.; Shamma, D.; Bernstein, M.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Imagine this! scripts to compositions to videos
Gupta, T.; Schwenk, D.; Farhadi, A.; Hoiem, D.; and Kembhavi, A. 2018 · 2018
Earlier work this paper cites.
Storygan: A sequential conditional gan for story visualization
Li, Y.; Gan, Z.; Shen, Y.; Liu, J.; Cheng, Y.; Wu, Y.; Carin, L.; Carlson, D.; and Gao, J. 2019 · 2019
Earlier work this paper cites.
Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations
Wu, H.; Mao, J.; Zhang, Y.; Jiang, Y.; Li, L.; Sun, W.; and Ma, W.-Y. 2019 · 2019
Earlier work this paper cites.
ERNIE: Enhanced Language Representation with Informative Entities
Zhang, Z.; Han, X.; Liu, Z.; Jiang, X.; Sun, M.; and Liu, Q. 2019 · 2019
Earlier work this paper cites.
Improved-storygan for sequential images visualization
Li, C.; Kong, L.; and Zhou, Z. 2020 · 2020
Earlier work this paper cites.
Character-preserving coherent story visualization
Song, Y.-Z.; Rui Tam, Z.; Chen, H.-J.; Lu, H.-H.; and Shuai, H.-H. 2020 · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
Character-centric story visualization via visual planning and token alignment
Chen, H.; Han, R.; Wu, T.-L.; Nakayama, H.; and Peng, N. 2022 · 2022
Earlier work this paper cites.
Dreamartist: Towards controllable one-shot text-to-image generation via contrastive prompt-tuning
Dong, Z.; Wei, P.; and Lin, L. 2022 · 2022
Earlier work this paper cites.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A. H.; Chechik, G.; and Cohen-Or, D. 2022 · 2022
Earlier work this paper cites.
Clustering generative adversarial networks for story visualization
Li, B.; Torr, P. H.; and Lukasiewicz, T. 2022 · 2022
Earlier work this paper cites.
Storydall-e: Adapting pretrained text-to-image transformers for story continuation
Maharana, A.; Hannan, D.; and Bansal, M. 2022 · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Cited alongside, same era.
Modular StoryGAN with background and theme awareness for story visualization
Szűcs, G.; and Al-Shouha, M. 2022 · 2022
Cited alongside, same era.
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Bai, J.; Bai, S.; Yang, S.; Wang, S.; Tan, S.; Wang, P.; Lin, J.; Zhou, C.; and Zhou, J. 2023 · 2023
Cited alongside, same era.
Improved visual story generation with adaptive context modeling
Feng, Z.; Ren, Y.; Yu, X.; Feng, X.; Tang, D.; Shi, S.; and Qin, B. 2023 · 2023
Cited alongside, same era.
Talecrafter: Interactive story visualization with multiple characters
Gong, Y.; Pang, Y.; Cun, X.; Xia, M.; He, Y.; Chen, H.; Wang, L.; Zhang, Y.; Wang, X.; Shan, Y.; et al. 2023 · 2023
Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion
Xie, J.; Li, Y.; Huang, Y.; Liu, H.; Zhang, W.; Zheng, Y.; and Shou, M. Z. 2023 · 2023
Later among the works it cites.
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models
Ye, H.; Zhang, J.; Liu, S.; Han, X.; and Yang, W. 2023 · 2023
Later among the works it cites.
Customnet: Zero-shot object customization with variable-viewpoints in text-to-image diffusion models
Yuan, Z.; Cao, M.; Wang, X.; Qi, Z.; Yuan, C.; and Shan, Y. 2023 · 2023
Later among the works it cites.
The Chosen One: Consistent Characters in Text-to-Image Diffusion Models
Avrahami, O.; Hertz, A.; Vinker, Y.; Arar, M.; Fruchter, S.; Fried, O.; Cohen-Or, D.; and Lischinski, D. 2024 · 2024
Closest in time.
Subject-driven text-to-image generation via apprenticeship learning
Chen, W.; Hu, H.; Li, Y.; Ruiz, N.; Jia, X.; Chang, M.-W.; and Cohen, W. W. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Svdiff: Compact parameter space for diffusion fine-tuning
Han, L.; Li, Y.; Zhang, H.; Milanfar, P.; Metaxas, D.; and Yang, F. 2023 · 2023
Cited alongside, same era.
Taming encoder for zero fine-tuning image customization with text-to-image diffusion models
Jia, X.; Zhao, Y.; Chan, K. C.; Li, Y.; Zhang, H.; Gong, B.; Hou, T.; Wang, H.; and Su, Y.-C. 2023 · 2023
Cited alongside, same era.
Unified multi-modal latent diffusion for joint subject and text conditional image generation
Ma, Y.; Yang, H.; Wang, W.; Fu, J.; and Liu, J. 2023 · 2023
Cited alongside, same era.
DINOv2: Learning Robust Visual Features without Supervision
Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H. V.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; Howes, R.; Huang, P.-Y.; Xu, H.; Sharma, V.; Li, S.-W.; Galuba, W.; Rabbat, M.; Assran, M.; Ballas, N.; Synnaeve, G.; Misra, I.; Jegou, H.; Mairal, J.; Labatut, P.; Joulin, A.; and Bojanowski, P. 2023 · 2023
Cited alongside, same era.
Make-a-story: Visual memory conditioned consistent story generation
Rahman, T.; Lee, H.-Y.; Ren, J.; Tulyakov, S.; Mahajan, S.; and Sigal, L. 2023 · 2023
Cited alongside, same era.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; and Aberman, K. 2023 · 2023
Cited alongside, same era.
Make-A-Storyboard: A General Framework for Storyboard with Disentangled and Merged Control
Su, S.; Guo, L.; Gao, L.; Shen, H. T.; and Song, J. 2023 · 2023
Cited alongside, same era.
Closest in time.
AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation
Cheng, J.; Lu, X.; Li, H.; Zai, K. L.; Yin, B.; Cheng, Y.; Yan, Y.; and Liang, X. 2024 · 2024
Closest in time.
Mix-of-show: Decentralized low-rank adaptation for multi-concept customization of diffusion models
Gu, Y.; Wang, X.; Wu, J. Z.; Shi, Y.; Chen, Y.; Fan, Z.; Xiao, W.; Zhao, R.; Chang, S.; Wu, W.; et al. 2024 · 2024
Closest in time.
Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing
Li, D.; Li, J.; and Hoi, S. 2024 · 2024
Closest in time.
Intelligent Grimm-Open-ended Visual Storytelling via Latent Diffusion Models
Liu, C.; Wu, H.; Zhong, Y.; Zhang, X.; Wang, Y.; and Xie, W. 2024 · 2024
Closest in time.
AttenCraft: Attention-guided Disentanglement of Multiple Concepts for Text-to-Image Customization
Shentu, J.; Watson, M.; and Moubayed, N. A. 2024 · 2024
Closest in time.
Training-free consistent text-to-image generation
Tewel, Y.; Kaduri, O.; Gal, R.; Kasten, Y.; Wolf, L.; Chechik, G.; and Atzmon, Y. 2024 · 2024
Closest in time.
Instantid: Zero-shot identity-preserving generation in seconds
Wang, Q.; Bai, X.; Wang, H.; Qin, Z.; and Chen, A. 2024 · 2024
Closest in time.
Yang, Y.; Wang, W.; Peng, L.; Song, C.; Chen, Y.; Li, H.; Yang, X.; Lu, Q.; Cai, D.; Wu, B.; et al. 2024 · 2024
Closest in time.
Attention Calibration for Disentangled Text-to-Image Personalization
Zhang, Y.; Yang, M.; Zhou, Q.; and Wang, Z. 2024 · 2024
Closest in time.