Fetching the paper…
Reading the bibliography…
Open-ended story visualization is a challenging task that involves generating coherent image sequences from a given storyline.
N. Otsu et al. , “A threshold selection method from gray-level histograms,” Automatica , vol. 11, no. 285-296, pp. 23–27, 1975
1975
Earlier work this paper cites.
K. Madej, “Towards digital narrative for children: from education to entertainment, a historical perspective,” Computers in Entertainment (CIE) , vol. 1, no. 1, 2003
2003
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” Advances in neural information processing systems , vol. 27, 2014
2014
Earlier work this paper cites.
F. Turan and I. Ulutas, “Using storybooks as a character education tools.” Journal of education and practice , vol. 7, no. 15, pp. 169–176, 2016
2016
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
Y. Li, Z. Gan, Y. Shen, J. Liu, Y. Cheng, Y. Wu, L. Carin, D. Carlson, and J. Gao, “Storygan: A sequential conditional gan for story visualization,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 6329–6338
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 10 684–10 695
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 10 684–10 695
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” arXiv preprint arXiv:2207.12598 , 2022
2022
Earlier work this paper cites.
T. Rahman, H.-Y. Lee, J. Ren, S. Tulyakov, S. Mahajan, and L. Sigal, “Make-a-story: Visual memory conditioned consistent story generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 2493–2502
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 4195–4205
2023
Cited alongside, same era.
G. Ding, C. Zhao, W. Wang, Z. Yang, Z. Liu, H. Chen, and C. Shen, “Freecustom: Tuning-free customized image generation for multi-concept composition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 9089–9098
2024
Later among the works it cites.
Y. Tewel, O. Kaduri, R. Gal, Y. Kasten, L. Wolf, G. Chechik, and Y. Atzmon, “Training-free consistent text-to-image generation,” ACM Transactions on Graphics (TOG) , vol. 43, no. 4, pp. 1–18, 2024
2024
Later among the works it cites.
“Storydiffusion: Consistent self-attention for long-range image and video generation,” NeurIPS , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Li, C. Li, W. Zhu, B. Yu, Y. Zhao, C. Wan, H. You, H. Shi, and Y. Lin, “Instant-3d: Instant neural radiance field training towards on-device ar/vr 3d reconstruction,” in Proceedings of the 50th Annual International Symposium on Computer Architecture , 2023, pp. 1–13
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
X. Pan, P. Qin, Y. Li, H. Xue, and W. Chen, “Synthesizing coherent story with auto-regressive latent diffusion models,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 2920–2930
2024
Cited alongside, same era.
X. Luo, X. Zhang, Y. Xie, X. Tong, W. Yu, H. Chang, F. Ma, and F. R. Yu, “Codeswap: Symmetrically face swapping based on prior codebook,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 6910–6919
2024
Later among the works it cites.
Y. Zeng, V. M. Patel, H. Wang, X. Huang, T.-C. Wang, M.-Y. Liu, and Y. Balaji, “Jedi: Joint-image diffusion models for finetuning-free personalized text-to-image generation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 6786–6795
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Xue, X. Luo, Z. Hu, X. Zhang, X. Xiang, Y. Dai, J. Liu, Z. Zhang, M. Li, J. Yang et al. , “Human motion video generation: A survey,” Authorea Preprints , 2024
2024
Later among the works it cites.
OpenAI, “Chatgpt: Gpt-4,” https://openai.com/chatgpt, 2024
2024
Later among the works it cites.
Y. Gu, X. Wang, J. Z. Wu, Y. Shi, Y. Chen, Z. Fan, W. Xiao, R. Zhao, S. Chang, W. Wu et al. , “Mix-of-show: Decentralized low-rank adaptation for multi-concept customization of diffusion models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
D. Zhou, Y. Li, F. Ma, X. Zhang, and Y. Yang, “Migc: Multi-instance generation controller for text-to-image synthesis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 6818–6828
2024
Later among the works it cites.
2024
Later among the works it cites.
C. Liu, H. Wu, Y. Zhong, X. Zhang, Y. Wang, and W. Xie, “Intelligent grimm-open-ended visual storytelling via latent diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 6190–6200
2024
Later among the works it cites.
C. Schuhmann, “Improved aesthetic predictor,” https://github.com/christophschuhmann/improved-aesthetic-predictor, 2024
2024
Later among the works it cites.