Fetching the paper…
Reading the bibliography…
Story visualization aims to generate coherent image sequences that faithfully represent a narrative and match given character references.
Visual storytelling
Ting-Hao Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, et al · 2016
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Imagine this! scripts to compositions to videos
Tanmay Gupta, Dustin Schwenk, Ali Farhadi, Derek Hoiem, and Aniruddha Kembhavi · 2018
Earlier work this paper cites.
Arcface: Additive angular margin loss for deep face recognition
Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou · 2019
Earlier work this paper cites.
Storygan: A sequential conditional gan for story visualization
Yitong Li, Zhe Gan, Yelong Shen, Jingjing Liu, Yu Cheng, Yuexin Wu, Lawrence Carin, David Carlson, and Jianfeng Gao · 2019
Earlier work this paper cites.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Earlier work this paper cites.
Integrating visuospatial, linguistic and commonsense structure into story visualization
Adyasha Maharana and Mohit Bansal · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Character-centric story visualization via visual planning and token alignment
Hong Chen, Rujun Han, Te-Lin Wu, Hideki Nakayama, and Nanyun Peng · 2022
Earlier work this paper cites.
Generalizable neural performer: Learning robust radiance fields for human novel view synthesis
Wei Cheng, Su Xu, Jingtan Piao, Chen Qian, Wayne Wu, Kwan-Yee Lin, and Hongsheng Li · 2022
Earlier work this paper cites.
Word-level fine-grained story visualization
Bowen Li and Thomas Lukasiewicz · 2022
Earlier work this paper cites.
Collaborative neural rendering using anime character sheets
Zuzeng Lin, Ailin Huang, and Zhewei Huang · 2022
Earlier work this paper cites.
Dino: Detr with improved denoising anchor boxes for end-to-end object detection, 2022
Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M. Ni, and Heung-Yeung Shum · 2022
Earlier work this paper cites.
Storybench: A multifaceted benchmark for continuous story visualization
Emanuele Bugliarello, H Hernan Moraldo, Ruben Villegas, Mohammad Babaeizadeh, Mohammad Taghi Saffar, Han Zhang, Dumitru Erhan, Vittorio Ferrari, Pieter-Jan Kindermans, and Paul Voigtlaender · 2023
Earlier work this paper cites.
Dna-rendering: A diverse neural actor repository for high-fidelity human-centric rendering
Wei Cheng, Ruixiang Chen, Siming Fan, Wanqi Yin, Keyu Chen, Zhongang Cai, Jingbo Wang, Yang Gao, Zhengming Yu, Zhengyu Lin, et al · 2023
Earlier work this paper cites.
Improved visual story generation with adaptive context modeling
Zhangyin Feng, Yuchen Ren, Xinmiao Yu, Xiaocheng Feng, Duyu Tang, Shuming Shi, and Bing Qin · 2023
Earlier work this paper cites.
Dreamsim: Learning new dimensions of human visual similarity using synthetic data
Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola · 2023
Earlier work this paper cites.
Zero-shot generation of coherent storybook from plain text story using diffusion models
Hyeonho Jeong, Gihyun Kwon, and Jong Chul Ye · 2023
Earlier work this paper cites.
Renderme-360: a large digital asset library and benchmarks towards high-fidelity head avatars
Dongwei Pan, Long Zhuo, Jingtan Piao, Huiwen Luo, Wei Cheng, Yuxin Wang, Siming Fan, Shengqi Liu, Lei Yang, Bo Dai, et al · 2023
Earlier work this paper cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach · 2023
Earlier work this paper cites.
Exploring clip for assessing the look and feel of images
Jianyi Wang, Kelvin CK Chan, and Chen Change Loy · 2023
Earlier work this paper cites.
A survey on video diffusion models
Zhen Xing, Qijun Feng, Haoran Chen, Qi Dai, Han Hu, Hang Xu, Zuxuan Wu, and Yu-Gang Jiang · 2023
Earlier work this paper cites.
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Earlier work this paper cites.
Manga generation via layout-controllable diffusion
Siyu Chen, Dengjie Li, Zenghao Bao, Yao Zhou, Lingfeng Tan, Yujie Zhong, and Zheng Zhao · 2024
Earlier work this paper cites.
Autostudio: Crafting consistent subjects in multi-turn interactive image generation
Junhao Cheng, Xi Lu, Hanhui Li, Khun Loun Zai, Baiqiao Yin, Yuhao Cheng, Yiqiang Yan, and Xiaodan Liang · 2024
Earlier work this paper cites.
Theatergen: Character management with llm for consistent multi-turn image generation
Junhao Cheng, Baiqiao Yin, Kaixin Cai, Minbin Huang, Hanhui Li, Yuxin He, Xi Lu, Yue Li, Yifei Li, Yuhao Cheng, et al · 2024
Earlier work this paper cites.
Scaling rectified flow transformers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al · 2024
Earlier work this paper cites.
Short film dataset (sfd): A benchmark for story-level video understanding
Ridouane Ghermi, Xi Wang, Vicky Kalogeiton, and Ivan Laptev · 2024
Earlier work this paper cites.
A survey on quality metrics for text-to-image generation
Sebastian Hartwig, Dominik Engel, Leon Sick, Hannah Kniesel, Tristan Payer, Poonam Poonam, Michael Glöckler, Alex Bäuerle, and Timo Ropinski · 2024
Earlier work this paper cites.
Dreamstory: Open-domain story visualization by llm-guided multi-subject consistent diffusion
Huiguo He, Huan Yang, Zixi Tuo, Yuan Zhou, Qiuyue Wang, Yuhang Zhang, Zeyu Liu, Wenhao Huang, Hongyang Chao, and Jian Yin · 2024
Earlier work this paper cites.
Storyagent: Customized storytelling video generation via multi-agent collaboration
Panwen Hu, Jin Jiang, Jianqi Chen, Mingfei Han, Shengcai Liao, Xiaojun Chang, and Xiaodan Liang · 2024
Cited alongside, same era.
Story3d-agent: Exploring 3d storytelling visualization with large language models
Yuzhou Huang, Yiran Qin, Shunlin Lu, Xintao Wang, Rui Huang, Ying Shan, and Ruimao Zhang · 2024
Cited alongside, same era.
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al · 2024
Cited alongside, same era.
Visagent: Narrative-preserving story visualization framework
Seungkwon Kim, GyuTae Park, Sangyeon Kim, and Seung-Hun Nam · 2024
Cited alongside, same era.
Vast 1.0: A unified framework for controllable and consistent video generation
Chi Zhang, Yuanzhi Liang, Xi Qiu, Fangqiu Yi, and Xuelong Li · 2024
Later among the works it cites.
Dialogue director: Bridging the gap in dialogue visualization for multimodal storytelling
Min Zhang, Zilin Wang, Liyan Chen, Kunhong Liu, and Juncong Lin · 2024
Later among the works it cites.
Moviedreamer: Hierarchical generation for coherent long visual sequence, 2024
Canyu Zhao, Mingyu Liu, Wen Wang, Jianlong Yuan, Hao Chen, Bo Zhang, and Chunhua Shen · 2024
Later among the works it cites.
Storydiffusion: Consistent self-attention for long-range image and video generation
Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng, and Qibin Hou · 2024
Later among the works it cites.
Long prompt weighted stable diffusion embedding
Shudong Zhu(Andrew Zhu) · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anim-director: A large multimodal model powered agent for controllable animation video generation
Yunxin Li, Haoyuan Shi, Baotian Hu, Longyue Wang, Jiashun Zhu, Jinyi Xu, Zhen Zhao, and Min Zhang · 2024
Cited alongside, same era.
Intelligent grimm - open-ended visual storytelling via latent diffusion models
Chang Liu, Haoning Wu, Yujie Zhong, Xiaoyun Zhang, Yanfeng Wang, and Weidi Xie · 2024
Cited alongside, same era.
Intelligent grimm-open-ended visual storytelling via latent diffusion models
Chang Liu, Haoning Wu, Yujie Zhong, Xiaoyun Zhang, Yanfeng Wang, and Weidi Xie · 2024
Cited alongside, same era.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al · 2024
Cited alongside, same era.
Sora: A review on background, technology, limitations, and opportunities of large vision models
Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, et al · 2024
Cited alongside, same era.
Object isolated attention for consistent story visualization
Xiangyang Luo, Junhao Cheng, Yifan Xie, Xin Zhang, Tao Feng, Zhou Liu, Fei Ma, and Fei Yu · 2024
Cited alongside, same era.
Story-Adapter: A Training-free Iterative Framework for Long Story Visualization, 2024
Jiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang, Mude Hui, Bingjie Xu, and Yuyin Zhou · 2024
Cited alongside, same era.
Story-adapter: A training-free iterative framework for long story visualization
Jiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang, Mude Hui, Bingjie Xu, and Yuyin Zhou · 2024
Cited alongside, same era.
Vlogger: Make your dream a vlog
Shaobin Zhuang, Kunchang Li, Xinyuan Chen, Yaohui Wang, Ziwei Liu, Yu Qiao, and Yali Wang · 2024
Later among the works it cites.
Doubao ai assistant
ByteDance Inc · 2025
Closest in time.
Gemini 2.0 flash: Native image generation in google ai studio
Google DeepMind · 2025
Closest in time.
Google gemini2. experiment with gemini 2.0 flash native image generation, 2025
Google DeepMind · 2025
Closest in time.
Aesthetic predictor v2.5
discus0434 · 2025
Closest in time.
Vinabench: Benchmark for faithful and consistent visual narratives
Silin Gao, Sheryl Mathew, Li Mi, Sepideh Mamooler, Mengjie Zhao, Hiromi Wakaki, Yuki Mitsufuji, Syrielle Montariol, and Antoine Bosselut · 2025
Closest in time.
Seedream 2.0: A native chinese-english bilingual image generation foundation model
Lixue Gong, Xiaoxia Hou, Fanshi Li, Liang Li, Xiaochen Lian, Fei Liu, Liyang Liu, Wei Liu, Wei Lu, Yichun Shi, et al · 2025
Closest in time.
Long context tuning for video generation
Yuwei Guo, Ceyuan Yang, Ziyan Yang, Zhibei Ma, Zhijie Lin, Zhenheng Yang, Dahua Lin, and Lu Jiang · 2025
Closest in time.
Step-video-ti2v technical report: A state-of-the-art text-driven image-to-video generation model
Haoyang Huang, Guoqing Ma, Nan Duan, Xing Chen, Changyi Wan, Ranchen Ming, Tianyu Wang, Bo Wang, Zhiying Lu, Aojie Li, et al · 2025
Closest in time.
Text2story: Advancing video storytelling with text guidance
Taewon Kang, Divya Kothandaraman, and Ming C. Lin · 2025
Closest in time.
Brmgo: Ai-powered tool for story script generation
MagicLight AI · 2025
Closest in time.
Shenbi: Ai-powered scriptwriting tool by maoyan
Maoyan Entertainment · 2025
Closest in time.
Introducing morphic studio
Morphic, Inc · 2025
Closest in time.
Gpt-4.1: Advanced large language model for natural language understanding and generation
OpenAI · 2025
Closest in time.
Favor-bench: A comprehensive benchmark for fine-grained video motion understanding
Chongjun Tu, Lin Zhang, Pengtao Chen, Peng Ye, Xianfang Zeng, Wei Cheng, Gang Yu, and Tao Chen · 2025
Closest in time.
Typemovie: Text-to-video storytelling with style and rhythm
TypeMovie Team · 2025
Closest in time.
Automated movie generation via multi-agent cot planning, 2025
Mike Zheng Shou Weijia Wu, Zeyu Zhu · 2025
Closest in time.
Less-to-more generalization: Unlocking more controllability by in-context generation
Shaojin Wu, Mengqi Huang, Wenxu Wu, Yufeng Cheng, Fei Ding, and Qian He · 2025
Closest in time.
Less-to-more generalization: Unlocking more controllability by in-context generation
Shaojin Wu, Mengqi Huang, Wenxu Wu, Yufeng Cheng, Fei Ding, and Qian He · 2025
Closest in time.
Automated movie generation via multi-agent cot planning
Weijia Wu, Zeyu Zhu, and Mike Zheng Shou · 2025
Closest in time.
Xuenan Xu, Jiahao Mei, Chenliang Li, Yuning Wu, Ming Yan, Shaopeng Lai, Ji Zhang, and Mengyue Wu · 2025
Closest in time.
Storyweaver: A unified world model for knowledge-enhanced story character customization
Jinlu Zhang, Jiji Tang, Rongsheng Zhang, Tangjie Lv, and Xiaoshuai Sun · 2025
Closest in time.
Dropletvideo: A dataset and approach to explore integral spatio-temporal consistent video generation
Runze Zhang, Guoguang Du, Xiaochuan Li, Qi Jia, Liang Jin, Lu Liu, Jingjing Wang, Cong Xu, Zhenhua Guo, Yaqian Zhao, et al · 2025
Closest in time.
Envisioning beyond the pixels: Benchmarking reasoning-informed visual editing
Xiangyu Zhao, Peiyuan Zhang, Kexian Tang, Hao Li, Zicheng Zhang, Guangtao Zhai, Junchi Yan, Hua Yang, Xue Yang, and Haodong Duan · 2025
Closest in time.
Styleme3d: Stylization with disentangled priors by multiple encoders on 3d gaussians
Cailin Zhuang, Yaoqi Hu, Xuanyang Zhang, Wei Cheng, Jiacheng Bao, Shengqi Liu, Yiying Yang, Xianfang Zeng, Gang Yu, and Ming Li · 2025
Closest in time.