Fetching the paper…
Reading the bibliography…
The excellent text-to-image synthesis capability of diffusion models has driven progress in synthesizing coherent visual stories.
D. Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv: Machine Learning,arXiv: Machine Learning
2013
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18
2015
Earlier work this paper cites.
A. Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning.,” Cornell University - arXiv,Cornell University - arXiv
2017
Earlier work this paper cites.
T. Gupta, D. Schwenk, A. Farhadi, D. Hoiem, and A. Kembhavi, “Imagine this! scripts to compositions to videos,” arXiv: Computer Vision and Pattern Recognition,arXiv: Computer Vision and Pattern Recognition
2018
Earlier work this paper cites.
Y. Li, Z. Gan, Y. Shen, J. Liu, Y. Cheng, Y. Wu, L. Carin, D. Carlson, and J. Gao, “Storygan: A sequential conditional gan for story visualization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2019
Earlier work this paper cites.
G. Zeng, Z. Li, and Y. Zhang, “Pororogan: An improved story visualization model on pororo-sv dataset,” in Proceedings of the 2019 3rd International Conference on Computer Science and Artificial Intelligence
2019
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Communications of the ACM
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems
2020
Earlier work this paper cites.
Y.-Z. Song, Z. Rui Tam, H.-J. Chen, H.-H. Lu, and H.-H. Shuai, “Character-preserving coherent story visualization,” in European Conference on Computer Vision
2020
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” International Conference on Machine Learning,International Conference on Machine Learning
2021
Cited alongside, same era.
2022
Later among the works it cites.
B. Li and T. Lukasiewicz, “Word-level fine-grained story visualization,” arXiv e-prints
2022
Later among the works it cites.
2022
Later among the works it cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in International Conference on Machine Learning
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al
2021
Cited alongside, same era.
A. Q. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” in International Conference on Machine Learning
2021
Cited alongside, same era.
A. Maharana, D. Hannan, and M. Bansal, “Storydall-e: Adapting pretrained text-to-image transformers for story continuation,” in European Conference on Computer Vision
2022
Cited alongside, same era.
2022
Later among the works it cites.
J. Li, D. Li, C. Xiong, and S. Hoi, “Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in International Conference on Machine Learning
2022
Later among the works it cites.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” arXiv preprint arXiv:2207.12598
2022
Later among the works it cites.