Fetching the paper…
Reading the bibliography…
Large-scale text-to-video diffusion models have demonstrated an exceptional ability to synthesize diverse videos.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling. 2013 · 2013
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation. In Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18 . Springer, 234–241
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics. In International Conference on Machine Learning (ICML) . 2256–2265
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015 · 2015
Earlier work this paper cites.
Image style transfer using convolutional neural networks. In IEEE/CVF Conferences on Computer Vision and Pattern Recognition (CVPR) . 2414–2423
Leon A Gatys, Alexander S Ecker, and Matthias Bethge. 2016 · 2016
Earlier work this paper cites.
Artistic style transfer for videos. In Pattern Recognition: 38th German Conference, GCPR 2016, Hannover, Germany, September 12-15, 2016, Proceedings 38 . Springer, 26–36
Manuel Ruder, Alexey Dosovitskiy, and Thomas Brox. 2016 · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis. 2016 · 2016
Earlier work this paper cites.
Coherent online video style transfer. In Proceedings of the IEEE International Conference on Computer Vision . 1105–1114
Dongdong Chen, Jing Liao, Lu Yuan, Nenghai Yu, and Gang Hua. 2017 · 2017
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization. In IEEE International Conference on Computer Vision (ICCV) . 1501–1510
Xun Huang and Serge Belongie. 2017 · 2017
Earlier work this paper cites.
Universal style transfer via feature transforms
Yijun Li, Chen Fang, Jimei Yang, Zhaowen Wang, Xin Lu, and Ming-Hsuan Yang. 2017 · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Fast video multi-style transfer. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 3222–3230
Wei Gao, Yijun Li, Yihang Yin, and Ming-Hsuan Yang. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Manigan: Text-guided image manipulation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 7880–7889
Bowen Li, Xiaojuan Qi, Thomas Lukasiewicz, and Philip HS Torr. 2020 · 2020
Earlier work this paper cites.
Arbitrary video style transfer via multi-channel correlation. In in Proc. 35th AAAI Conf. Artif. Intell. (AAAI) , Vol. 35. 1210–1217
Yingying Deng, Fan Tang, Weiming Dong, Haibin Huang, Chongyang Ma, and Changsheng Xu. 2021 · 2021
Earlier work this paper cites.
Layered neural atlases for consistent video editing
Yoni Kasten, Dolev Ofri, Oliver Wang, and Tali Dekel. 2021 · 2021
Earlier work this paper cites.
GLIDE: Towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021 · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision. In Proc. ICML . PMLR, 8748–8763
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation. In International Conference on Machine Learning . PMLR, 8821–8831
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Cited alongside, same era.
Godiva: Generating open-domain videos from natural descriptions
Chenfei Wu, Lun Huang, Qianxi Zhang, Binyang Li, Lei Ji, Fan Yang, Guillermo Sapiro, and Nan Duan. 2021 · 2021
Cited alongside, same era.
ediffi: Text-to-image diffusion models with an ensemble of expert denoisers
Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, Bryan Catanzaro, et al · 2022
Cited alongside, same era.
Text2live: Text-driven layered image and video editing. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XV . Springer, 707–723
Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten, and Tali Dekel. 2022 · 2022
Cited alongside, same era.
SinFusion: Training Diffusion Models on a Single Image or Video
Yaniv Nikankin, Niv Haim, and Michal Irani. 2022 · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022 · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR) . 10684–10695
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Later among the works it cites.
Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S Sara Mahdavi, Rapha Gontijo Lopes, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tim Brooks, Aleksander Holynski, and Alexei A Efros. 2022 · 2022
Cited alongside, same era.
Stytr2: Image style transfer with transformers. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR) . 11326–11336
Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Xingjia Pan, Lei Wang, and Changsheng Xu. 2022 · 2022
Cited alongside, same era.
Cogview2: Faster and better text-to-image generation via hierarchical transformers
Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. 2022 · 2022
Cited alongside, same era.
Language-Driven Artistic Style Transfer. In European Conference on Computer Vision (ECCV)
Tsu-Jui Fu, Xin Eric Wang, and William Yang Wang. 2022 · 2022
Cited alongside, same era.
StyleGAN-NADA: CLIP-guided domain adaptation of image generators
Rinon Gal, Or Patashnik, Haggai Maron, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2022 · 2022
Cited alongside, same era.
Prompt-to-prompt image editing with cross attention control
Amir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman, Yael Pritch, and Daniel Cohen-Or. 2022 · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. 2022 · 2022
Cited alongside, same era.
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al · 2022
Later among the works it cites.
Splicing vit features for semantic appearance transfer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10748–10757
Narek Tumanyan, Omer Bar-Tal, Shai Bagon, and Tali Dekel. 2022 · 2022
Later among the works it cites.
Nüwa: Visual synthesis pre-training for neural visual world creation. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XVI . Springer, 720–736
Chenfei Wu, Jian Liang, Lei Ji, Fan Yang, Yuejian Fang, Daxin Jiang, and Nan Duan. 2022b · 2022
Later among the works it cites.
Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation
Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Weixian Lei, Yuchao Gu, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. 2022a · 2022
Later among the works it cites.
Inversion-Based Creativity Transfer with Diffusion Models
Yuxin Zhang, Nisha Huang, Fan Tang, Haibin Huang, Chongyang Ma, Weiming Dong, and Changsheng Xu. 2022a · 2022
Later among the works it cites.
Magicvideo: Efficient video generation with latent diffusion models
Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. 2022 · 2022
Later among the works it cites.
Structure and Content-Guided Video Synthesis with Diffusion Models
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis. 2023 · 2023
Closest in time.
Deforum Stable Diffusion V0.7
Eddy Hu. 2023 · 2023
Closest in time.
Exploring the Temporal Consistency of Arbitrary Style Transfer: A Channelwise Perspective
Xiaoyu Kong, Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Yongyong Chen, Zhenyu He, and Changsheng Xu. 2023 · 2023
Closest in time.
Blind Video Deflickering by Neural Filtering with a Flawed Atlas
Chenyang Lei, Xuanchi Ren, Zhaoxiang Zhang, and Qifeng Chen. 2023 · 2023
Closest in time.
FateZero: Fusing Attentions for Zero-shot Text-based Video Editing
Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. 2023 · 2023
Closest in time.
VideoCrafter: A Toolkit for Text-to-Video Generation and Editing
Yong Zhang, Yingqing He, Haoxin Chen, and Menghan Xia. 2023 · 2023
Closest in time.
StyleCLIP: Text-Driven Manipulation of StyleGAN Imagery. In IEEE/CVF International Conference on Computer Vision (ICCV) . 2085–2094
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. 2021 · 2094
Closest in time.