Fetching the paper…
Reading the bibliography…
With the emerging diffusion models, recently, text-to-video generation has aroused increasing attention.
Adversarial Video Generation on Complex Datasets
Clark, A.; Donahue, J.; ; and Simonyan, K. 2019 · 1907
Earlier work this paper cites.
STM: SpatioTemporal and Motion Encoding for Action Recognition
Jiang, B.; Wang, M.; Gan, W.; Wu, W.; and Yan, J. 2019 · 2009
Earlier work this paper cites.
A Short Note on the Kinetics-700-2020 Human Action Dataset
Smaira, L.; Carreira, J.; Noland, E.; Clancy, E.; Wu, A.; and Zisserman, A. 2020 · 2010
Earlier work this paper cites.
Denoising Diffusion Implicit Models
Song, J.; Meng, C.; and Ermon, S. 2022 · 2010
Earlier work this paper cites.
Generating Videos with Scene Dynamics
Vondrick, C.; Pirsiavash, H.; ; and Torralba, A. 2016 · 2016
Earlier work this paper cites.
Video Pixel Networks
Kalchbrenner, N.; Oord, A.; Simonyan, K.; Danihelka, I.; Vinyals, O.; Graves, A.; ; and Kavukcuoglu, K. 2017 · 2017
Earlier work this paper cites.
Neural Discrete Representation Learning
van den Oord, A.; Vinyals, O.; and Kavukcuoglu, K. 2018 · 2018
Earlier work this paper cites.
Generative adversarial networks
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2020 · 2020
Earlier work this paper cites.
Denoising Diffusion Probabilistic Models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; and Narang, S. 2020 · 2020
Earlier work this paper cites.
Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
Bain, M.; Nagrani, A.; Varol, G.; and Zisserman, A. 2021 · 2021
Cited alongside, same era.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Cited alongside, same era.
Glide: Towards photorealistic image generation and editing with text-guided diffusion models
Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; and Chen, M. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; and Agarwal, S. 2021 · 2021
Cited alongside, same era.
Flexible Diffusion Modeling of Long Videos
Harvey, W.; Naderiparizi, S.; Masrani, V.; Weilbach, C.; ; and Wood, F. 2022 · 2022
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; and Ghasemipour, K. 2022 · 2022
Later among the works it cites.
LAION-5B: A new era of open large-scale multi-modal datasets
Schuhmann, C.; Vencu, R.; Beaumont, R.; Coombes, T.; Gordon, C.; Katta, A.; Kaczmarczyk, R.; ; and Jitsev, J. 2022 · 2022
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Singer, U.; Polyak, A.; Hayes, T.; Yin, X.; An, J.; Zhang, S.; Hu, Q.; Yang, H.; Ashual, O.; and Gafni, O. 2022 · 2022
Later among the works it cites.
Advancing high-resolution videolanguage representation with large-scale video transcriptions
Xue, H.; Hang, T.; Zeng, Y.; Sun, Y.; Liu, B.; Yang, H.; Fu, J.; ; and Guo, B. 2022 · 2022
Later among the works it cites.
Align Your Latents: High-Resolution Video Synthesis With Latent Diffusion Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ho, J.; Salimans, T.; Gritsenko, A.; Chan, W.; Norouzi, M.; and Fleet, D. J. 2022 · 2022
Cited alongside, same era.
CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Hong, W.; Ding, M.; Zheng, W.; Liu, X.; and Tang, J. 2022 · 2022
Cited alongside, same era.
Hierarchical Text-Conditional Image Generation with CLIP Latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer., B. 2022 · 2022
Cited alongside, same era.
Blattmann, A.; Rombach, R.; Ling, H.; Dockhorn, T.; Kim, S. W.; Fidler, S.; and Kreis, K. 2023 · 2023
Closest in time.
AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Guo, Y.; Yang, C.; Rao, A.; Wang, Y.; Qiao, Y.; Lin, D.; and Dai, B. 2023 · 2023
Closest in time.
Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators
Khachatryan, L.; Movsisyan, A.; Tadevosyan, V.; Henschel, R.; Wang, Z.; Navasardyan, S.; and Shi, H. 2023 · 2023
Closest in time.
Conditional Image-to-Video Generation With Latent Flow Diffusion Models
Ni, H.; Shi, C.; Li, K.; Huang, S. X.; and Min, M. R. 2023 · 2023
Closest in time.
Video Probabilistic Diffusion Models in Projected Latent Space
Yu, S.; Sohn, K.; Kim, S.; and Shin, J. 2023 · 2023
Closest in time.