D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114 , 2013
Original
2013
Earlier work this paper cites.
D. Rezende and S. Mohamed, “Variational inference with normalizing flows,” in ICML , 2015, pp. 1530–1538
2015
Earlier work this paper cites.
L. Dinh, J. Sohl-Dickstein, and S. Bengio, “Density estimation using real nvp,” ICLR , May 2016
2016
Earlier work this paper cites.
X. Jun, T. Mei, T. Yao, and Y. Rui, “MSR-VTT: A large video description dataset for bridging video and language,” ACM MM , Jan 2016
2016
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals et al. , “Neural discrete representation learning,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a local nash equilibrium,” NeurIPS , Jan 2017
2017
Earlier work this paper cites.
W. Xiong, W. Luo, L. Ma, W. Liu, and J. Luo, “Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks,” in CVPR , 2018, pp. 2364–2373
2018
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in ACL , Jan 2018
2018
Earlier work this paper cites.
S. Nam, C. Ma, M. Chai, W. Brendel, N. Xu, and S. J. Kim, “End-to-end time-lapse video synthesis from a single outdoor image,” in CVPR , 2019, pp. 1409–1418
2019
Earlier work this paper cites.
Y. Endo, Y. Kanamori, and S. Kuriyama, “Animating landscape: self-supervised learning of decoupled motion and appearance for single-image video synthesis,” ACM Transactions on Graphics (TOG) , vol. 38, no. 6, pp. 1–19, 2019
2019
Earlier work this paper cites.
T. Unterthiner, S. Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly, “Fvd: A new metric for video generation,” ICLR , Mar 2019
2019
Earlier work this paper cites.
J. Zhang, C. Xu, L. Liu, M. Wang, X. Wu, Y. Liu, and Y. Jiang, “Dtvnet: Dynamic time-lapse video generation via single still image,” in ECCV , 2020, pp. 300–315
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” NeurIPS , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” NeurIPS , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
H. Xue, B. Liu, H. Yang, J. Fu, H. Li, and J. Luo, “Learning fine-grained motion embedding for landscape animation,” in ACM MM , 2021, pp. 291–299
2021
Earlier work this paper cites.
W. Xia, Y. Yang, J.-H. Xue, and B. Wu, “Tedigan: Text-guided diverse face image generation and manipulation,” in CVPR , 2021, pp. 2256–2265
2021
Earlier work this paper cites.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in ICML , 2021, pp. 8821–8831
2021
Earlier work this paper cites.
M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” in ICCV , 2021
2021
Earlier work this paper cites.
H. Xue, B. Liu, H. Yang, J. Fu, H. Li, and J. Luo, “Learning fine-grained motion embedding for landscape animation,” in ACM MM , 2021, pp. 291–299
2021
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in ICLR , 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021, pp. 8748–8763
2021
Earlier work this paper cites.
C. Wu, L. Huang, Q. Zhang, B. Li, L. Ji, F. Yang, G. Sapiro, and N. Duan, “Godiva: Generating open-domain videos from natural descriptions,” arXiv preprint arXiv:2104.14806 , 2021
Original
2021
Earlier work this paper cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al. , “Laion-5b: An open large-scale dataset for training next generation image-text models,” NeurIPS , vol. 35, pp. 25 278–25 294, 2022
2022
Earlier work this paper cites.
A. Sauer, K. Schwarz, and A. Geiger, “Stylegan-xl: Scaling stylegan to large diverse datasets,” in SIGGRAPH , 2022, pp. 1–10
2022
Earlier work this paper cites.
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans et al. , “Photorealistic text-to-image diffusion models with deep language understanding,” NeurIPS , vol. 35, pp. 36 479–36 494, 2022
2022
Earlier work this paper cites.
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen, “Hierarchical text-conditional image generation with clip latents,” arXiv preprint arXiv:2204.06125 , 2022
Original
2022
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in CVPR , Jun 2022
2022
Earlier work this paper cites.
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang, “Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,” in ICML , 2022, pp. 23 318–23 340
2022
Earlier work this paper cites.
J. Wang, H. Yuan, D. Chen, Y. Zhang, X. Wang, and S. Zhang, “Modelscope text-to-video technical report,” arXiv preprint arXiv:2308.06571 , 2023
Original
2023
Earlier work this paper cites.
A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis, “Align your latents: High-resolution video synthesis with latent diffusion models,” in CVPR , 2023, pp. 22 563–22 575
2023
Earlier work this paper cites.