Fetching the paper…
Reading the bibliography…
Recent works have successfully extended large-scale text-to-image models to the video domain, producing promising results but at a high computational cost and requiring a large amount of video data.
Conditional generative adversarial nets
Mirza, M.; and Osindero, S. 2014 · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O.; Fischer, P.; and Brox, T. 2015 · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; and Ganguli, S. 2015 · 2015
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A.; Vinyals, O.; et al. 2017 · 2017
Earlier work this paper cites.
Video-to-Video Synthesis
Wang, T.-C.; Liu, M.-Y.; Zhu, J.-Y.; Liu, G.; Tao, A.; Kautz, J.; and Catanzaro, B. 2018 · 2018
Earlier work this paper cites.
Everybody dance now
Chan, C.; Ginosar, S.; Zhou, T.; and Efros, A. A. 2019 · 2019
Earlier work this paper cites.
Liquid warping gan: A unified framework for human motion imitation, appearance transfer and novel view synthesis
Liu, W.; Piao, Z.; Min, J.; Luo, W.; Ma, L.; and Gao, S. 2019 · 2019
Earlier work this paper cites.
First order motion model for image animation
Siarohin, A.; Lathuilière, S.; Tulyakov, S.; Ricci, E.; and Sebe, N. 2019 · 2019
Earlier work this paper cites.
Few-shot Video-to-Video Synthesis
Wang, T.-C.; Liu, M.-Y.; Tao, A.; Liu, G.; Kautz, J.; and Catanzaro, B. 2019 · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
G3AN: Disentangling Appearance and Motion for Video Generation
Wang, Y.; Bilinski, P.; Bremond, F.; and Dantcheva, A. 2020 · 2020
Earlier work this paper cites.
ImaGINator: Conditional Spatio-Temporal GAN for Video Generation
WANG, Y.; Bilinski, P.; Bremond, F.; and Dantcheva, A. 2020 · 2020
Earlier work this paper cites.
Is space-time attention all you need for video understanding?
Bertasius, G.; Wang, H.; and Torresani, L. 2021 · 2021
Earlier work this paper cites.
Diffusion models beat gans on image synthesis
Dhariwal, P.; and Nichol, A. 2021 · 2021
Earlier work this paper cites.
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Hessel, J.; Holtzman, A.; Forbes, M.; Le Bras, R.; and Choi, Y. 2021 · 2021
Cited alongside, same era.
Variational diffusion models
Kingma, D.; Salimans, T.; Poole, B.; and Ho, J. 2021 · 2021
Cited alongside, same era.
Benchmark for compositional text-to-image synthesis
Park, D. H.; Azadi, S.; Liu, X.; Darrell, T.; and Rohrbach, A. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Denoising Diffusion Implicit Models
Song, J.; Meng, C.; and Ermon, S. 2021 · 2021
Cited alongside, same era.
Score-Based Generative Modeling through Stochastic Differential Equations
Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2021 · 2021
Pix2Video: Video Editing using Image Diffusion
Ceylan, D.; Huang, C. P.; and Mitra, N. J. 2023 · 2023
Closest in time.
Control-A-Video: Controllable Text-to-Video Generation with Diffusion Models
Chen, W.; Wu, J.; Xie, P.; Wu, H.; Li, J.; Xia, X.; Xiao, X.; and Lin, L. 2023 · 2023
Closest in time.
Structure and content-guided video synthesis with diffusion models
Esser, P.; Chiu, J.; Atighehchian, P.; Granskog, J.; and Germanidis, A. 2023 · 2023
Closest in time.
CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Hong, W.; Ding, M.; Zheng, W.; Liu, X.; and Tang, J. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Godiva: Generating open-domain videos from natural descriptions
Wu, C.; Huang, L.; Zhang, Q.; Li, B.; Ji, L.; Yang, F.; Sapiro, G.; and Duan, N. 2021 · 2021
Cited alongside, same era.
Make-a-scene: Scene-based text-to-image generation with human priors
Gafni, O.; Polyak, A.; Ashual, O.; Sheynin, S.; Parikh, D.; and Taigman, Y. 2022 · 2022
Cited alongside, same era.
Hierarchical text-conditional image generation with clip latents
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022 · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022 · 2022
Cited alongside, same era.
Latent Image Animator: Learning to Animate Images via Latent Space Navigation
Wang, Y.; Yang, D.; Bremond, F.; and Dantcheva, A. 2022 · 2022
Cited alongside, same era.
Hu, Z.; and Xu, D. 2023 · 2023
Closest in time.
Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators
Khachatryan, L.; Movsisyan, A.; Tadevosyan, V.; Henschel, R.; Wang, Z.; Navasardyan, S.; and Shi, H. 2023 · 2023
Closest in time.
Video-P2P: Video Editing with Cross-attention Control
Liu, S.; Zhang, Y.; Li, W.; Lin, Z.; and Jia, J. 2023 · 2023
Closest in time.
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
Ma, Y.; He, Y.; Cun, X.; Wang, X.; Shan, Y.; Li, X.; and Chen, Q. 2023 · 2023
Closest in time.
Dreamix: Video diffusion models are general video editors
Molad, E.; Horwitz, E.; Valevski, D.; Acha, A. R.; Matias, Y.; Pritch, Y.; Leviathan, Y.; and Hoshen, Y. 2023 · 2023
Closest in time.
Mou, C.; Wang, X.; Xie, L.; Zhang, J.; Qi, Z.; Shan, Y.; and Qie, X. 2023 · 2023
Closest in time.
Fatezero: Fusing attentions for zero-shot text-based video editing
Qi, C.; Cun, X.; Zhang, Y.; Lei, C.; Wang, X.; Shan, Y.; and Chen, Q. 2023 · 2023
Closest in time.
Make-A-Video: Text-to-Video Generation without Text-Video Data
Singer, U.; Polyak, A.; Hayes, T.; Yin, X.; An, J.; Zhang, S.; Hu, Q.; Yang, H.; Ashual, O.; Gafni, O.; Parikh, D.; Gupta, S.; and Taigman, Y. 2023 · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
Zhang, L.; and Agrawala, M. 2023 · 2023
Closest in time.
ControlVideo: Training-free Controllable Text-to-Video Generation
Zhang, Y.; Wei, Y.; Jiang, D.; Zhang, X.; Zuo, W.; and Tian, Q. 2023 · 2023
Closest in time.