Fetching the paper…
Reading the bibliography…
Significant advancements have been achieved in the realm of large-scale pre-trained text-to-video Diffusion Models (VDMs).
Soomro K, Zamir AR, Shah M (2012) Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:12120402
2012
Earlier work this paper cites.
Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y (2014) Generative adversarial networks. NIPS
2014
Earlier work this paper cites.
Ronneberger O, Fischer P, Brox T (2015) U-net: Convolutional networks for biomedical image segmentation. Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18 pp 234–241
2015
Earlier work this paper cites.
Srivastava N, Mansimov E, Salakhudinov R (2015) Unsupervised learning of video representations using lstms. In: ICML
2015
Earlier work this paper cites.
Reed S, Akata Z, Yan X, Logeswaran L, Schiele B, Lee H (2016) Generative adversarial text to image synthesis. In: ICML , PMLR, pp 1060–1069
2016
Earlier work this paper cites.
Vondrick C, Pirsiavash H, Torralba A (2016) Generating videos with scene dynamics. NIPS
2016
Earlier work this paper cites.
Xu J, Mei T, Yao T, Rui Y (2016) Msr-vtt: A large video description dataset for bridging video and language. In: CVPR
2016
Earlier work this paper cites.
Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S (2017) Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30
2017
Earlier work this paper cites.
Saito M, Matsumoto E, Saito S (2017) Temporal generative adversarial nets with singular value clipping. In: ICCV
2017
Earlier work this paper cites.
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser Ł, Polosukhin I (2017) Attention is all you need. Advances in neural information processing systems 30
2017
Earlier work this paper cites.
Zhang H, Xu T, Li H, Zhang S, Wang X, Huang X, Metaxas DN (2017) Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks. In: ICCV , pp 5907–5915
2017
Earlier work this paper cites.
Hong S, Yang D, Choi J, Lee H (2018) Inferring semantic layout for hierarchical text-to-image synthesis. In: CVPR , pp 7986–7994
2018
Earlier work this paper cites.
Tulyakov S, Liu MY, Yang X, Kautz J (2018) MoCoGAN: Decomposing Motion and Content for Video Generation. In: CVPR
2018
Earlier work this paper cites.
Unterthiner T, Van Steenkiste S, Kurach K, Marinier R, Michalski M, Gelly S (2018) Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:181201717
2018
Earlier work this paper cites.
Xu T, Zhang P, Huang Q, Zhang H, Gan Z, Huang X, He X (2018) Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In: CVPR , pp 1316–1324
2018
Earlier work this paper cites.
Ho J, Jain A, Abbeel P (2020) Denoising diffusion probabilistic models. NeurIPS 33:6840–6851
2020
Earlier work this paper cites.
Raffel C, Shazeer N, Roberts A, Lee K, Narang S, Matena M, Zhou Y, Li W, Liu PJ (2020) Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21(140):1–67, URL http://jmlr.org/papers/v21/20-074.html
2020
Earlier work this paper cites.
Bain M, Nagrani A, Varol G, Zisserman A (2021) Frozen in time: A joint video and image encoder for end-to-end retrieval. In: Proceedings of the IEEE/CVF International Conference on Computer Vision , pp 1728–1738
2021
Earlier work this paper cites.
Le Moing G, Ponce J, Schmid C (2021) Ccvs: Context-aware controllable video synthesis. NeurIPS
2021
Earlier work this paper cites.
Meng C, He Y, Song Y, Song J, Wu J, Zhu JY, Ermon S (2021) Sdedit: Guided image synthesis and editing with stochastic differential equations. In: International Conference on Learning Representations
2021
Cited alongside, same era.
Nichol A, Dhariwal P, Ramesh A, Shyam P, Mishkin P, McGrew B, Sutskever I, Chen M (2021) Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:211210741
2021
Cited alongside, same era.
Skorokhodov I, Tulyakov S, Elhoseiny M (2021) Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2. arXiv preprint arXiv:211214683
2021
Cited alongside, same era.
Tian Y, Ren J, Chai M, Olszewski K, Peng X, Metaxas DN, Tulyakov S (2021) A Good Image Generator Is What You Need for High-Resolution Video Synthesis. In: ICLR
2021
Cited alongside, same era.
Schuhmann C, Beaumont R, Vencu R, Gordon C, Wightman R, Cherti M, Coombes T, Katta A, Mullis C, Wortsman M, et al. (2022) Laion-5b: An open large-scale dataset for training next generation image-text models. arXiv preprint arXiv:221008402
2022
Later among the works it cites.
Singer U, Polyak A, Hayes T, Yin X, An J, Zhang S, Hu Q, Yang H, Ashual O, Gafni O, et al. (2022) Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:220914792
2022
Later among the works it cites.
Voleti V, Jolicoeur-Martineau A, Pal C (2022) Masked conditional video diffusion for prediction, generation, and interpolation. arXiv preprint arXiv:220509853
2022
Later among the works it cites.
Yang R, Srivastava P, Mandt S (2022) Diffusion probabilistic modeling for video generation. arXiv preprint arXiv:220309481
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yan W, Zhang Y, Abbeel P, Srinivas A (2021) Videogpt: Video generation using vq-vae and transformers. arXiv preprint arXiv:210410157
2021
Cited alongside, same era.
Yu S, Tack J, Mo S, Kim H, Kim J, Ha JW, Shin J (2021) Generating Videos with Dynamics-aware Implicit Generative Adversarial Networks. In: ICLR
2021
Cited alongside, same era.
Zhang H, Koh JY, Baldridge J, Lee H, Yang Y (2021) Cross-modal contrastive learning for text-to-image generation. In: CVPR , pp 833–842
2021
Cited alongside, same era.
Balaji Y, Nah S, Huang X, Vahdat A, Song J, Zhang Q, Kreis K, Aittala M, Aila T, Laine S, Catanzaro B, et al. (2022) eDiff-I: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:221101324
2022
Cited alongside, same era.
Ge S, Hayes T, Yang H, Yin X, Pang G, Jacobs D, Huang JB, Parikh D (2022) Long video generation with time-agnostic vqgan and time-sensitive transformer. arXiv preprint arXiv:220403638
2022
Cited alongside, same era.
Gu S, Chen D, Bao J, Wen F, Zhang B, Chen D, Yuan L, Guo B (2022) Vector quantized diffusion model for text-to-image synthesis. In: CVPR , pp 10696–10706
2022
Cited alongside, same era.
Harvey W, Naderiparizi S, Masrani V, Weilbach C, Wood F (2022) Flexible diffusion modeling of long videos. arXiv preprint arXiv:220511495
2022
Cited alongside, same era.
He Y, Yang T, Zhang Y, Shan Y, Chen Q (2022) Latent video diffusion models for high-fidelity video generation with arbitrary lengths. arXiv preprint arXiv:221113221
2022
Cited alongside, same era.
2022
Later among the works it cites.
An J, Zhang S, Yang H, Gupta S, Huang JB, Luo J, Yin X (2023) Latent-shift: Latent diffusion with temporal shift for efficient text-to-video generation. arXiv preprint arXiv:230408477
2023
Closest in time.
Chen H, Xia M, He Y, Zhang Y, Cun X, Yang S, Xing J, Liu Y, Chen Q, Wang X, et al. (2023) Videocrafter1: Open diffusion models for high-quality video generation. arXiv preprint arXiv:231019512
2023
Closest in time.
Esser P, Chiu J, Atighehchian P, Granskog J, Germanidis A (2023) Structure and content-guided video synthesis with diffusion models. arXiv preprint arXiv:230203011
2023
Closest in time.
Ge S, Nah S, Liu G, Poon T, Tao A, Catanzaro B, Jacobs D, Huang JB, Liu MY, Balaji Y (2023) Preserve your own correlation: A noise prior for video diffusion models. arXiv preprint arXiv:230510474
2023
Closest in time.
Huang Z, He Y, Yu J, Zhang F, Si C, Jiang Y, Zhang Y, Wu T, Jin Q, Chanpaisit N, et al. (2023) Vbench: Comprehensive benchmark suite for video generative models. arXiv preprint arXiv:231117982
2023
Closest in time.
Jeong H, Park GY, Ye JC (2023) Vmc: Video motion customization using temporal attention adaption for text-to-video diffusion models. arXiv preprint arXiv:231200845
2023
Closest in time.
Khachatryan L, Movsisyan A, Tadevosyan V, Henschel R, Wang Z, Navasardyan S, Shi H (2023) Text2video-zero: Text-to-image diffusion models are zero-shot video generators. arXiv preprint arXiv:230313439
2023
Closest in time.
Kondratyuk D, Yu L, Gu X, Lezama J, Huang J, Hornung R, Adam H, Akbari H, Alon Y, Birodkar V, et al. (2023) Videopoet: A large language model for zero-shot video generation. arXiv preprint arXiv:231214125
2023
Closest in time.
Luo Z, Chen D, Zhang Y, Huang Y, Wang L, Shen Y, Zhao D, Zhou J, Tan T (2023) VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation. In: CVPR
2023
Closest in time.
Shen X, Li X, Elhoseiny M (2023) MoStGAN-V: Video Generation With Temporal Motion Styles. In: CVPR
2023
Closest in time.
Yin S, Wu C, Yang H, Wang J, Wang X, Ni M, Yang Z, Li L, Liu S, Yang F, et al. (2023) Nuwa-xl: Diffusion over diffusion for extremely long video generation. arXiv preprint arXiv:230312346
2023
Closest in time.
Zhao R, Gu Y, Wu JZ, Zhang DJ, Liu J, Wu W, Keppo J, Shou MZ (2023) Motiondirector: Motion customization of text-to-video diffusion models. arXiv preprint arXiv:231008465
2023
Closest in time.
Bar-Tal O, Chefer H, Tov O, Herrmann C, Paiss R, Zada S, Ephrat A, Hur J, Li Y, Michaeli T, et al. (2024) Lumiere: A space-time diffusion model for video generation. arXiv preprint arXiv:240112945
2024
Closest in time.
Huang Z, He Y, Yu J, Zhang F, Si C, Jiang Y, Zhang Y, Wu T, Jin Q, Chanpaisit N, Wang Y, Chen X, Wang L, Lin D, Qiao Y, Liu Z (2024) VBench: Comprehensive Benchmark Suite for Video Generative Models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2024
Closest in time.