Fetching the paper…
Reading the bibliography…
Text-to-video (T2V) models have shown remarkable capabilities in generating diverse videos.
Multi-concept customization of text-to-image diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1931–1941
Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. 2023 · 1941
Earlier work this paper cites.
Separating style and content with bilinear models
Joshua B Tenenbaum and William T Freeman. 2000 · 2000
Earlier work this paper cites.
Image analogies. In Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques . 327–340
Aaron Hertzmann, Charles E Jacobs, Nuria Oliver, Brian Curless, and David H Salesin. 2001 · 2001
Earlier work this paper cites.
Efficient example-based painting and synthesis of 2d directional texture
Bin Wang, Wenping Wang, Huaiping Yang, and Jiaguang Sun. 2004 · 2004
Earlier work this paper cites.
Video watercolorization using bidirectional texture advection
Adrien Bousseau, Fabrice Neyret, Joëlle Thollot, and David Salesin. 2007 · 2007
Earlier work this paper cites.
Wiki Art Gallery, Inc.: A case for critical thinking
Fred Phillips and Brandy Mackintosh. 2011 · 2011
Earlier work this paper cites.
Style transfer via image component analysis
Wei Zhang, Chen Cao, Shifeng Chen, Jianzhuang Liu, and Xiaoou Tang. 2013 · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Image style transfer using convolutional neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2414–2423
Leon A Gatys, Alexander S Ecker, and Matthias Bethge. 2016 · 2016
Earlier work this paper cites.
Artistic style transfer for videos. In Pattern Recognition: 38th German Conference, GCPR 2016, Hannover, Germany, September 12-15, 2016, Proceedings 38 . Springer, 26–36
Manuel Ruder, Alexey Dosovitskiy, and Thomas Brox. 2016 · 2016
Earlier work this paper cites.
Coherent online video style transfer. In Proceedings of the IEEE International Conference on Computer Vision . 1105–1114
Dongdong Chen, Jing Liao, Lu Yuan, Nenghai Yu, and Gang Hua. 2017 · 2017
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization. In Proceedings of the IEEE international conference on computer vision . 1501–1510
Xun Huang and Serge Belongie. 2017 · 2017
Earlier work this paper cites.
Universal style transfer via feature transforms
Yijun Li, Chen Fang, Jimei Yang, Zhaowen Wang, Xin Lu, and Ming-Hsuan Yang. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Stylizing video by example
Ondřej Jamriška, Šárka Sochorová, Ondřej Texler, Michal Lukáč, Jakub Fišer, Jingwan Lu, Eli Shechtman, and Daniel S‘ykora. 2019 · 2019
Earlier work this paper cites.
Arbitrary style transfer with style-attentional networks. In proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5880–5888
Dae Young Park and Kwang Hee Lee. 2019 · 2019
Earlier work this paper cites.
Fast video multi-style transfer. In Proceedings of the IEEE/CVF winter conference on applications of computer vision . 3222–3230
Wei Gao, Yijun Li, Yihang Yin, and Ming-Hsuan Yang. 2020 · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems (NeurIPS) , Vol. 33. 6840–6851
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020 · 2020
Earlier work this paper cites.
Raft: Recurrent all-pairs field transforms for optical flow. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16 . Springer, 402–419
Zachary Teed and Jia Deng. 2020 · 2020
Earlier work this paper cites.
Arbitrary style transfer using neurally-guided patch-based synthesis
Ondřej Texler, David Futschik, Jakub Fišer, Michal Lukáč, Jingwan Lu, Eli Shechtman, and Daniel Sỳkora. 2020a · 2020
Earlier work this paper cites.
Interactive video stylization using few-shot patch-based training
Ondřej Texler, David Futschik, Michal Kučera, Ondřej Jamriška, Šárka Sochorová, Menclei Chai, Sergey Tulyakov, and Daniel Sỳkora. 2020b · 2020
Earlier work this paper cites.
Artflow: Unbiased image style transfer via reversible neural flows. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 862–871
Jie An, Siyu Huang, Yibing Song, Dejing Dou, Wei Liu, and Jiebo Luo. 2021 · 2021
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 1728–1738
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman. 2021 · 2021
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision . 9650–9660
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. 2021 · 2021
Earlier work this paper cites.
Arbitrary video style transfer via multi-channel correlation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 1210–1217
Yingying Deng, Fan Tang, Weiming Dong, Haibin Huang, Chongyang Ma, and Changsheng Xu. 2021 · 2021
Earlier work this paper cites.
Manifold alignment for semantically aligned style transfer. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 14861–14869
Jing Huo, Shiyin Jin, Wenbin Li, Jing Wu, Yu-Kun Lai, Yinghuan Shi, and Yang Gao. 2021 · 2021
Cited alongside, same era.
Adaattn: Revisit attention mechanism in arbitrary neural style transfer. In Proceedings of the IEEE/CVF international conference on computer vision . 6649–6658
Songhua Liu, Tianwei Lin, Dongliang He, Fu Li, Meiling Wang, Xin Li, Zhengxing Sun, Qian Li, and Errui Ding. 2021 · 2021
Cited alongside, same era.
Improved denoising diffusion probabilistic models. In International Conference on Machine Learning (ICML) . PMLR, 8162–8171
Alexander Quinn Nichol and Prafulla Dhariwal. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision. In International Conference on Machine Learning (ICML)
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Cited alongside, same era.
Text2video-zero: Text-to-image diffusion models are zero-shot video generators
Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. 2023 · 2023
Closest in time.
Evalcrafter: Benchmarking and evaluating large video generation models
Yaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang, Yong Zhang, Haoxin Chen, Yang Liu, Tieyong Zeng, Raymond Chan, and Ying Shan. 2023 · 2023
Closest in time.
Text-guided synthesis of eulerian cinemagraphs
Aniruddha Mahapatra, Aliaksandr Siarohin, Hsin-Ying Lee, Sergey Tulyakov, and Jun-Yan Zhu. 2023 · 2023
Closest in time.
GPT-4V(ision) System Card
OpenAI. 2023 · 2023
Closest in time.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023a · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stytr2: Image style transfer with transformers. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 11326–11336
Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Xingjia Pan, Lei Wang, and Changsheng Xu. 2022 · 2022
Cited alongside, same era.
Cogview2: Faster and better text-to-image generation via hierarchical transformers
Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. 2022 · 2022
Cited alongside, same era.
An image is worth one word: Personalizing text-to-image generation using textual inversion
Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2022 · 2022
Cited alongside, same era.
Latent video diffusion models for high-fidelity video generation with arbitrary lengths
Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. 2022 · 2022
Cited alongside, same era.
Imagen video: High definition video generation with diffusion models
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al · 2022
Cited alongside, same era.
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. 2022b · 2022
Cited alongside, same era.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang. 2022 · 2022
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models. In International Conference on Learning Representations (ICLR)
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023b · 2023
Closest in time.
Fatezero: Fusing attentions for zero-shot text-based video editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision
Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen. 2023 · 2023
Closest in time.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 22500–22510
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2023 · 2023
Closest in time.
Instantbooth: Personalized text-to-image generation without test-time finetuning
Jing Shi, Wei Xiong, Zhe Lin, and Hyun Joon Jung. 2023 · 2023
Closest in time.
StyleDrop: Text-to-Image Generation in Any Style. In Advances in Neural Information Processing Systems (NeurIPS)
Kihyuk Sohn, Nataniel Ruiz, Kimin Lee, Daniel Castro Chin, Irina Blok, Huiwen Chang, Jarred Barber, Lu Jiang, Glenn Entis, Yuanzhen Li, et al · 2023
Closest in time.
Modelscope text-to-video technical report
Jiuniu Wang, Hangjie Yuan, Dayou Chen, Yingya Zhang, Xiang Wang, and Shiwei Zhang. 2023c · 2023
Closest in time.
LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models
Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, et al · 2023
Closest in time.
StyleAdapter: A Single-Pass LoRA-Free Model for Stylized Image Generation
Zhouxia Wang, Xintao Wang, Liangbin Xie, Zhongang Qi, Ying Shan, Wenping Wang, and Ping Luo. 2023b · 2023
Closest in time.
Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 15943–15953
Yuxiang Wei, Yabo Zhang, Zhilong Ji, Jinfeng Bai, Lei Zhang, and Wangmeng Zuo. 2023 · 2023
Closest in time.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 7623–7633
Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei, Yuchao Gu, Yufei Shi, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou. 2023 · 2023
Closest in time.
Dynamicrafter: Animating open-domain images with video diffusion priors
Jinbo Xing, Menghan Xia, Yong Zhang, Haoxin Chen, Xintao Wang, Tien-Tsin Wong, and Ying Shan. 2023 · 2023
Closest in time.
Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation. In ACM SIGGRAPH Asia 2023 Conference Proceedings
Shuai Yang, Yifan Zhou, Ziwei Liu, , and Chen Change Loy. 2023 · 2023
Closest in time.
IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. 2023 · 2023
Closest in time.
Inversion-based style transfer with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10146–10156
Yuxin Zhang, Nisha Huang, Fan Tang, Haibin Huang, Chongyang Ma, Weiming Dong, and Changsheng Xu. 2023 · 2023
Closest in time.
Videocrafter2: Overcoming data limitations for high-quality video diffusion models
Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. 2024 · 2024
Closest in time.
Tokenflow: Consistent diffusion features for consistent video editing. In In International Conference on Learning Representations (ICLR)
Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. 2024 · 2024
Closest in time.
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Yuwei Guo, Ceyuan Yang, Anyi Rao, Yaohui Wang, Yu Qiao, Dahua Lin, and Bo Dai. 2024 · 2024
Closest in time.
Videocomposer: Compositional video synthesis with motion controllability
Xiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen, Jiuniu Wang, Yingya Zhang, Yujun Shen, Deli Zhao, and Jingren Zhou. 2024 · 2024
Closest in time.
ToonCrafter: Generative Cartoon Interpolation
Jinbo Xing, Hanyuan Liu, Menghan Xia, Yong Zhang, Xintao Wang, Ying Shan, and Tien-Tsin Wong. 2024 · 2024
Closest in time.
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8703–8712
Shuai Yang, Yifan Zhou, Ziwei Liu, and Chen Change Loy. 2024 · 2024
Closest in time.