Fetching the paper…
Reading the bibliography…
Customized text-to-video generation aims to generate high-quality videos guided by text prompts and subject references.
N. Kumari, B. Zhang, R. Zhang, E. Shechtman, and J.-Y. Zhu, “Multi-concept customization of text-to-image diffusion,” in IEEE/CVF Computer Vision and Pattern Recognition Conference , 2023, pp. 1931–1941
1941
Earlier work this paper cites.
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” in Advances in neural information processing systems , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in Advances in neural information processing systems , 2017, pp. 6309–6318
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Y. Balaji, M. R. Min, B. Bai, R. Chellappa, and H. P. Graf, “Conditional gan with discriminative filter generation for text-to-video synthesis,” in International Joint Conference on Artificial Intelligence , vol. 1, no. 2019, 2019, p. 2
2019
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
J. Ho and T. Salimans, “Classifier-free diffusion guidance,” in NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications , 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International Conference on Machine Learning , 2021, pp. 8748–8763
2021
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in International Conference on Learning Representations , 2021
2021
Earlier work this paper cites.
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in International Conference on Computer Vision , 2021, pp. 9650–9660
2021
Earlier work this paper cites.
C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu, “Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps,” Advances in neural information processing systems , vol. 35, pp. 5775–5787, 2022
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in IEEE/CVF Computer Vision and Pattern Recognition Conference , 2022, pp. 10 684–10 695
2022
Earlier work this paper cites.
I. Skorokhodov, S. Tulyakov, and M. Elhoseiny, “Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2,” in IEEE/CVF Computer Vision and Pattern Recognition Conference , 2022, pp. 3626–3636
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
P. von Platen, S. Patil, A. Lozhkov, P. Cuenca, N. Lambert, K. Rasul, M. Davaadorj, and T. Wolf, “Diffusers: State-of-the-art diffusion models,” https://github.com/huggingface/diffusers , 2022
2022
Earlier work this paper cites.
S. Gugger, L. Debut, T. Wolf, P. Schmid, Z. Mueller, S. Mangrulkar, M. Sun, and B. Bossan, “Accelerate: Training and inference at scale made simple, efficient and adaptable,” https://github.com/huggingface/accelerate , 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick, “Segment anything,” International Conference on Computer Vision , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W. Hong, M. Ding, W. Zheng, X. Liu, and J. Tang, “Cogvideo: Large-scale pretraining for text-to-video generation via transformers,” International Conference on Learning Representations , 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
S. Sterling, “Zeroscope,” https://huggingface.co/cerspense/zeroscope_v2_576w , 2023
2023
Later among the works it cites.
Y. Gu, X. Wang, J. Z. Wu, Y. Shi, Y. Chen, Z. Fan, W. Xiao, R. Zhao, S. Chang, W. Wu et al. , “Mix-of-show: Decentralized low-rank adaptation for multi-concept customization of diffusion models,” Advances in Neural Information Processing Systems , vol. 36, 2023
2023
Later among the works it cites.
S. Sterling, “Zeroscope xl,” https://huggingface.co/cerspense/zeroscope_v2_XL , 2023
2023
Later among the works it cites.
P. Esser, J. Chiu, P. Atighehchian, J. Granskog, and A. Germanidis, “Structure and content-guided video synthesis with diffusion models,” in International Conference on Computer Vision , 2023, pp. 7346–7356
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Villegas, M. Babaeizadeh, P.-J. Kindermans, H. Moraldo, H. Zhang, M. T. Saffar, S. Castro, J. Kunze, and D. Erhan, “Phenaki: Variable length video generation from open domain textual descriptions,” in International Conference on Learning Representations , 2023
2023
Cited alongside, same era.
J. Zhu, H. Ma, J. Chen, and J. Yuan, “Motionvideogan: A novel video generator based on the motion space learned from image pairs,” IEEE Transactions on Multimedia , vol. 25, pp. 9370–9382, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
A. Köksal, K. E. Ak, Y. Sun, D. Rajan, and J. H. Lim, “Controllable video generation with text-based instructions,” IEEE transactions on multimedia , vol. 26, pp. 190–201, 2023
2023
Cited alongside, same era.
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, D. Parikh, S. Gupta, and Y. Taigman, “Make-a-video: Text-to-video generation without text-video data,” in International Conference on Learning Representations , 2023
2023
Cited alongside, same era.
A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis, “Align your latents: High-resolution video synthesis with latent diffusion models,” in IEEE/CVF Computer Vision and Pattern Recognition Conference , 2023, pp. 22 563–22 575
2023
Cited alongside, same era.
2023
Later among the works it cites.
Y. Zeng, G. Wei, J. Zheng, J. Zou, Y. Wei, Y. Zhang, and H. Li, “Make pixels dance: High-dynamic video generation,” IEEE/CVF Computer Vision and Pattern Recognition Conference , pp. 8850–8860, 2024
2024
Closest in time.
D. Kondratyuk, L. Yu, X. Gu, J. Lezama, J. Huang, R. Hornung, H. Adam, H. Akbari, Y. Alon, V. Birodkar et al. , “Videopoet: A large language model for zero-shot video generation,” International Conference on Machine Learning , 2024
2024
Closest in time.
H. Chen, Y. Zhang, X. Cun, M. Xia, X. Wang, C. Weng, and Y. Shan, “Videocrafter2: Overcoming data limitations for high-quality video diffusion models,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 7310–7320
2024
Closest in time.
Y. Jiang, T. Wu, S. Yang, C. Si, D. Lin, Y. Qiao, C. C. Loy, and Z. Liu, “Videobooth: Diffusion-based video generation with image prompts,” IEEE/CVF Computer Vision and Pattern Recognition Conference , 2024
2024
Closest in time.
Y. Wei, S. Zhang, Z. Qing, H. Yuan, Z. Liu, Y. Liu, Y. Zhang, J. Zhou, and H. Shan, “Dreamvideo: Composing your dream videos with customized subject and motion,” in IEEE/CVF Computer Vision and Pattern Recognition Conference , 2024
2024
Closest in time.
Y. Zhang, M. Yang, Q. Zhou, and Z. Wang, “Attention calibration for disentangled text-to-image personalization,” IEEE/CVF Computer Vision and Pattern Recognition Conference , 2024
2024
Closest in time.
Y. Hu, C. Luo, and Z. Chen, “A benchmark for controllable text -image-to-video generation,” IEEE Transactions on Multimedia , vol. 26, pp. 1706–1719, 2024
2024
Closest in time.
M. Zhao, W. Wang, T. Chen, R. Zhang, and R. Li, “Ta2v: Text-audio guided video generation,” IEEE Transactions on Multimedia , vol. 26, pp. 7250–7264, 2024
2024
Closest in time.
Y. Guo, C. Yang, A. Rao, Y. Wang, Y. Qiao, D. Lin, and B. Dai, “Animatediff: Animate your personalized text-to-image diffusion models without specific tuning,” International Conference on Learning Representations , 2024
2024
Closest in time.
X. Wang, H. Yuan, S. Zhang, D. Chen, J. Wang, Y. Zhang, Y. Shen, D. Zhao, and J. Zhou, “Videocomposer: Compositional video synthesis with motion controllability,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
M. Geyer, O. Bar-Tal, S. Bagon, and T. Dekel, “Tokenflow: Consistent diffusion features for consistent video editing,” International Conference on Learning Representations , 2024
2024
Closest in time.
J. Ma, J. Liang, C. Chen, and H. Lu, “Subject-diffusion: Open domain personalized text-to-image generation without test-time fine-tuning,” in ACM SIGGRAPH 2024 Conference Papers , 2024, pp. 1–12
2024
Closest in time.
Y. Jiang, Q. Liu, D. Chen, L. Yuan, and Y. Fu, “Animediff: Customized image generation of anime characters using diffusion model,” IEEE Transactions on Multimedia , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Anthropic, “Introducing the next generation of claude,” https://www.anthropic.com/news/claude-3-family , 2024
2024
Closest in time.