Fetching the paper…
Reading the bibliography…
Story visualization aims to create visually compelling images or videos corresponding to textual narratives.
N. Otsu et al. , “A threshold selection method from gray-level histograms,” Automatica , vol. 11, no. 285-296, pp. 23–27, 1975
1975
Earlier work this paper cites.
K. Carter, “The place of story in the study of teaching and teacher education,” Educational researcher , vol. 22, no. 1, pp. 5–18, 1993
1993
Earlier work this paper cites.
C. Klimmt, C. Roth, I. Vermeulen, P. Vorderer, and F. S. Roth, “Forecasting the experience of future entertainment technology: “interactive storytelling” and media enjoyment,” Games and Culture , vol. 7, no. 3, pp. 187–208, 2012
2012
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in ICLR , Y. Bengio and Y. LeCun, Eds., 2014
2014
Earlier work this paper cites.
D. Kostons and B. B. de Koning, “Does visualization affect monitoring accuracy, restudy choice, and comprehension scores of students in primary education?” Contemporary Educational Psychology , vol. 51, pp. 1–10, 2017
2017
Earlier work this paper cites.
L. Metz, B. Poole, D. Pfau, and J. Sohl-Dickstein, “Unrolled generative adversarial networks,” in ICLR , 2017
2017
Earlier work this paper cites.
M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein generative adversarial networks,” in ICML , 2017, pp. 214–223
2017
Earlier work this paper cites.
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. C. Courville, “Improved training of Wasserstein GANs,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
T. Gupta, D. Schwenk, A. Farhadi, D. Hoiem, and A. Kembhavi, “Imagine this! scripts to compositions to videos,” in ECCV , 2018, pp. 598–613
2018
Earlier work this paper cites.
Z. Cao, F. Wei, W. Li, and S. Li, “Faithful to the original: Fact aware neural abstractive summarization,” in AAAI , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
S. Wang, M. Yu, X. Guo, Z. Wang, T. Klinger, W. Zhang, S. Chang, G. Tesauro, B. Zhou, and J. Jiang, “ R 3 R^{3} : Reinforced ranker-reader for open-domain question answering,” in AAAI , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
S. Min, V. Zhong, R. Socher, and C. Xiong, “Efficient and robust question answering from minimal context over documents,” in ACL , 2018, pp. 1725–1735
2018
Earlier work this paper cites.
Y. Li, Z. Gan, Y. Shen, J. Liu, Y. Cheng, Y. Wu, L. Carin, D. Carlson, and J. Gao, “StoryGAN: A sequential conditional gan for story visualization,” in CVPR , June 2019
2019
Earlier work this paper cites.
M. Gui, J. Tian, R. Wang, and Z. Yang, “Attention optimization for abstractive document summarization,” in EMNLP-IJCNLP 2019 , 2019, pp. 1222–1228
2019
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in ICLR , 2020
2020
Earlier work this paper cites.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in ICLR , 2020
2020
Earlier work this paper cites.
Y.-Z. Song, Z. Rui Tam, H.-J. Chen, H.-H. Lu, and H.-H. Shuai, “Character-preserving coherent story visualization,” in ECCV . Springer, 2020, pp. 18–33
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” NeurIPS , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
C. Meister, S. Lazov, I. Augenstein, and R. Cotterell, “Is sparse attention more interpretable?” in ACL/IJCNLP . Association for Computational Linguistics, 2021, pp. 122–129
2021
Earlier work this paper cites.
A. Q. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” in ICML , 2021, pp. 8162–8171
2021
Earlier work this paper cites.
P. Dhariwal and A. Nichol, “Diffusion models beat GANs on image synthesis,” NeurIPS , vol. 34, pp. 8780–8794, 2021
2021
Earlier work this paper cites.
A. Maharana and M. Bansal, “Integrating visuospatial, linguistic, and commonsense structure into story visualization,” in EMNLP , 2021, pp. 6772–6786
2021
Earlier work this paper cites.
A. Maharana, D. Hannan, and M. Bansal, “Improving generation and evaluation of visual stories via semantic consistency,” in NAACL HLT , 2021, pp. 2427–2442
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021, pp. 8748–8763
2021
Earlier work this paper cites.
E. J. Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al. , “LoRA: Low-rank adaptation of large language models,” in ICLR , 2021
2021
Earlier work this paper cites.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” in NeurIPS 2021 Workshop , 2021
2021
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in CVPR , 2022, pp. 10 684–10 695
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” NeurIPS , vol. 35, pp. 24 824–24 837, 2022
2022
Earlier work this paper cites.
B. Li, P. H. Torr, and T. Lukasiewicz, “Clustering generative adversarial networks for story visualization,” in ACM MM , 2022, pp. 769–778
2022
Earlier work this paper cites.
Y. Ma, H. Yang, B. Liu, J. Fu, and J. Liu, “AI illustrator: Translating raw descriptions into images by prompt-based cross-modal generation,” in ACM MM , 2022, pp. 4282–4290
2022
Earlier work this paper cites.
C. Saharia, W. Chan, H. Chang, C. Lee, J. Ho, T. Salimans, D. Fleet, and M. Norouzi, “Palette: Image-to-image diffusion models,” in ACM SIGGRAPH , 2022, pp. 1–10
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
H. Chen, R. Han, T.-L. Wu, H. Nakayama, and N. Peng, “Character-centric story visualization via visual planning and token alignment,” in EMNLP , 2022, pp. 8259–8272
2022
Earlier work this paper cites.
A. Maharana, D. Hannan, and M. Bansal, “StoryDALL-E: Adapting pretrained text-to-image transformers for story continuation,” in ECCV . Springer, 2022, pp. 70–87
2022
Cited alongside, same era.
J. Li, D. Li, C. Xiong, and S. Hoi, “BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,” in ICML . PMLR, 2022, pp. 12 888–12 900
2022
Cited alongside, same era.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” NeurIPS , vol. 35, pp. 27 730–27 744, 2022
2022
Cited alongside, same era.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” NeurIPS , vol. 35, pp. 24 824–24 837, 2022
2022
Cited alongside, same era.
2024
Closest in time.
G. Sun, W. Liang, J. Dong, J. Li, Z. Ding, and Y. Cong, “Create your world: Lifelong text-to-image diffusion,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024
2024
Closest in time.
M. Zhang, Z. Cai, L. Pan, F. Hong, X. Guo, L. Yang, and Z. Liu, “MotionDiffuse: Text-driven human motion generation with diffusion model,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 46, no. 6, pp. 4115–4128, 2024
2024
Closest in time.
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma et al. , “Scaling instruction-finetuned language models,” Journal of Machine Learning Research , vol. 25, no. 70, pp. 1–53, 2024
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
F.-A. Croitoru, V. Hondru, R. T. Ionescu, and M. Shah, “Diffusion models in vision: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2023
2023
Cited alongside, same era.
Z. Liu, P. Dai, R. Li, X. Qi, and C.-W. Fu, “DreamStone: Image as a stepping stone for text-guided 3d shape generation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 12, pp. 14 385–14 403, 2023
2023
Cited alongside, same era.
F. Zhan, Y. Yu, R. Wu, J. Zhang, S. Lu, L. Liu, A. Kortylewski, C. Theobalt, and E. Xing, “Multimodal image synthesis and editing: The generative AI era,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 12, pp. 15 098–15 119, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
N. Ruiz, Y. Li, V. Jampani, Y. Pritch, M. Rubinstein, and K. Aberman, “DreamBooth: Fine tuning text-to-image diffusion models for subject-driven generation,” in CVPR , 2023, pp. 22 500–22 510
2023
Cited alongside, same era.
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan, “Tree of thoughts: Deliberate problem solving with large language models,” NeurIPS , vol. 36, 2024
2024
Closest in time.
O. Avrahami, A. Hertz, Y. Vinker, M. Arar, S. Fruchter, O. Fried, D. Cohen-Or, and D. Lischinski, “The chosen one: Consistent characters in text-to-image diffusion models,” in ACM SIGGRAPH 2024 conference papers , 2024, pp. 1–12
2024
Closest in time.
S. Jang, J. Jo, K. Lee, and S. J. Hwang, “Identity decoupling for multi-subject personalization of text-to-image models,” in NeurIPS , 2024
2024
Closest in time.
C. Liu, H. Wu, Y. Zhong, X. Zhang, Y. Wang, and W. Xie, “Intelligent grimm-open-ended visual storytelling via latent diffusion models,” in CVPR , 2024, pp. 6190–6200
2024
Closest in time.
Y. Tewel, O. Kaduri, R. Gal, Y. Kasten, L. Wolf, G. Chechik, and Y. Atzmon, “Training-free consistent text-to-image generation,” TOG , vol. 43, no. 4, pp. 1–18, 2024
2024
Closest in time.
Y. Zhou, D. Zhou, M.-M. Cheng, J. Feng, and Q. Hou, “StoryDiffusion: Consistent self-attention for long-range image and video generation,” NeurIPS , vol. 37, pp. 110 315–110 340, 2024
2024
Closest in time.
2024
Closest in time.
X. Pan, P. Qin, Y. Li, H. Xue, and W. Chen, “Synthesizing coherent story with auto-regressive latent diffusion models,” in WACV , 2024, pp. 2920–2930
2024
Closest in time.
X. Peng, J. Zhu, B. Jiang, Y. Tai, D. Luo, J. Zhang, W. Lin, T. Jin, C. Wang, and R. Ji, “PortraitBooth: A versatile portrait model for fast identity-preserved personalization,” in CVPR , June 2024, pp. 27 080–27 090
2024
Closest in time.
2024
Closest in time.
S. Cui, J. Guo, X. An, J. Deng, Y. Zhao, X. Wei, and Z. Feng, “IDAdapter: Learning mixed features for tuning-free personalization of text-to-image models,” in CVPR Workshops , June 2024, pp. 950–959
2024
Closest in time.
X. Chen, L. Huang, Y. Liu, Y. Shen, D. Zhao, and H. Zhao, “AnyDoor: Zero-shot object-level image customization,” in CVPR , June 2024, pp. 6593–6602
2024
Closest in time.
Z. Li, M. Cao, X. Wang, Z. Qi, M.-M. Cheng, and Y. Shan, “PhotoMaker: Customizing realistic human photos via stacked id embedding,” in CVPR , June 2024, pp. 8640–8650
2024
Closest in time.
D. Li, J. Li, and S. Hoi, “BLIP-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing,” NeurIPS , vol. 36, 2024
2024
Closest in time.
Y. Hao, Z. Chi, L. Dong, and F. Wei, “Optimizing prompts for text-to-image generation,” NeurIPS , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
L. Yang, Z. Yu, C. Meng, M. Xu, S. Ermon, and C. Bin, “Mastering text-to-image diffusion: Recaptioning, planning, and generating with multimodal LLMs,” in ICML , 2024
2024
Closest in time.
2024
Closest in time.
S. Zhuang, K. Li, X. Chen, Y. Wang, Z. Liu, Y. Qiao, and Y. Wang, “Vlogger: Make your dream a vlog,” in CVPR . IEEE, 2024, pp. 8806–8817
2024
Closest in time.
W. Ren, H. Yang, G. Zhang, C. Wei, X. Du, W. Huang, and W. Chen, “ConsistI2V: Enhancing visual consistency for image-to-video generation,” TMLR , 2024
2024
Closest in time.
G. Feng, B. Zhang, Y. Gu, H. Ye, D. He, and L. Wang, “Towards revealing the mystery behind chain of thought: a theoretical perspective,” NeurIPS , vol. 36, 2024
2024
Closest in time.
Y. Alaluf, D. Garibi, O. Patashnik, H. Averbuch-Elor, and D. Cohen-Or, “Cross-image attention for zero-shot appearance transfer,” in ACM SIGGRAPH 2024 Conference Papers , 2024, pp. 1–12
2024
Closest in time.
D. Shen, G. Song, Z. Xue, F.-Y. Wang, and Y. Liu, “Rethinking the spatial inconsistency in classifier-free diffusion guidance,” in CVPR , 2024, pp. 9370–9379
2024
Closest in time.
M. Chen, I. Laina, and A. Vedaldi, “Training-free layout control with cross-attention guidance,” in WACV , 2024, pp. 5343–5353
2024
Closest in time.
2024
Closest in time.
S. Fu, N. Tamir, S. Sundaram, L. Chai, R. Zhang, T. Dekel, and P. Isola, “DreamSim: Learning new dimensions of human visual similarity using synthetic data,” NeurIPS , vol. 36, 2024
2024
Closest in time.
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, Q. Jiang, C. Li, J. Yang, H. Su et al. , “Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection,” in ECCV . Springer, 2024, pp. 38–55
2024
Closest in time.
2024
Closest in time.
C. Zhu, K. Li, Y. Ma, C. He, and X. Li, “MultiBooth: Towards generating all your concepts in an image from text,” in AAAI , vol. 39, no. 10, 2025, pp. 10 923–10 931
2025
Closest in time.