Fetching the paper…
Reading the bibliography…
The video generation field has witnessed rapid improvements with the introduction of recent diffusion models.
Wallace GK (1991) The jpeg still picture compression standard. Communications of the ACM 34(4):30–44
1991
Earlier work this paper cites.
Aifanti N, Papachristou C, Delopoulos A (2010) The mug facial expression database. In: WIAMIS, pp 1–4
2010
Earlier work this paper cites.
Song Y, Demirdjian D, Davis R (2011) Tracking body and hands for gesture recognition: Natops aircraft handling signals database. In: IEEE FG, pp 500–506
2011
Earlier work this paper cites.
Kingma DP, Welling M (2014) Auto-encoding variational bayes. In: ICLR
2014
Earlier work this paper cites.
Ronneberger O, Fischer P, Brox T (2015) U-net: Convolutional networks for biomedical image segmentation. In: MICCAI, Springer, pp 234–241
2015
Earlier work this paper cites.
Sohl-Dickstein J, Weiss E, Maheswaranathan N, Ganguli S (2015) Deep unsupervised learning using nonequilibrium thermodynamics. In: ICML, pp 2256–2265
2015
Earlier work this paper cites.
Çiçek Ö, Abdulkadir A, Lienkamp SS, Brox T, Ronneberger O (2016) 3d u-net: learning dense volumetric segmentation from sparse annotation. In: MICCAI, Springer, pp 424–432
2016
Earlier work this paper cites.
Alemi AA, Fischer I, Dillon JV, Murphy K (2017) Deep variational information bottleneck. In: ICLR
2017
Earlier work this paper cites.
Ebert F, Finn C, Lee AX, Levine S (2017) Self-supervised visual planning with temporal skip connections. In: CoRL, pp 344–356
2017
Earlier work this paper cites.
Heusel M, Ramsauer H, Unterthiner T, Nessler B, Hochreiter S (2017) Gans trained by a two time-scale update rule converge to a local nash equilibrium. In: NeurIPS, vol 30
2017
Earlier work this paper cites.
van den Oord A, Vinyals O, kavukcuoglu k (2017) Neural discrete representation learning. In: Guyon I, Luxburg UV, Bengio S, Wallach H, Fergus R, Vishwanathan S, Garnett R (eds) NeurIPS, vol 30
2017
Earlier work this paper cites.
Babaeizadeh M, Finn C, Erhan D, Campbell R, Levine S (2018) Stochastic variational video prediction. In: ICLR
2018
Earlier work this paper cites.
Li Y, Fang C, Yang J, Wang Z, Lu X, Yang MH (2018) Flow-grounded spatial-temporal video prediction from still images. In: ECCV, p 609–625
2018
Earlier work this paper cites.
Tulyakov S, Liu MY, Yang X, Kautz J (2018) Mocogan: Decomposing motion and content for video generation. In: CVPR, pp 1526–1535
2018
Earlier work this paper cites.
Unterthiner T, van Steenkiste S, Kurach K, Marinier R, Michalski M, Gelly S (2018) Towards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:181201717
2018
Earlier work this paper cites.
Xiong W, Luo W, Ma L, Liu W, Luo J (2018) Learning to generate time-lapse videos using multi-stage dynamic generative adversarial networks. In: CVPR, pp 2364–2373
2018
Earlier work this paper cites.
Clark A, Donahue J, Simonyan K (2019) Adversarial video generation on complex datasets. arXiv preprint arXiv:190706571
2019
Earlier work this paper cites.
Endo Y, Kanamori Y, Kuriyama S (2019) Animating landscape: self-supervised learning of decoupled motion and appearance for single-image video synthesis. ACM TOG 38(6):1–19
2019
Earlier work this paper cites.
Razavi A, van den Oord A, Vinyals O (2019) Generating diverse high-fidelity images with vq-vae-2. In: NeurIPS
2019
Earlier work this paper cites.
Weissenborn D, Täckström O, Uszkoreit J (2019) Scaling autoregressive video models. In: ICLR
2019
Earlier work this paper cites.
Chen M, Radford A, Child R, Wu J, Jun H, Luan D, Sutskever I (2020) Generative pretraining from pixels. In: ICML, PMLR, pp 1691–1703
2020
Earlier work this paper cites.
Goodfellow I, Pouget-Abadie J, Mirza M, Xu B, Warde-Farley D, Ozair S, Courville A, Bengio Y (2020) Generative adversarial networks. Communications of the ACM 63(11):139–144
2020
Earlier work this paper cites.
Ho J, Jain A, Abbeel P (2020) Denoising diffusion probabilistic models. In: NeurIPS, vol 33, pp 6840–6851
2020
Earlier work this paper cites.
Logacheva E, Suvorov R, Khomenko O, Mashikhin A, Lempitsky V (2020) Deeplandscape: Adversarial modeling of landscape videos. In: ECCV, Springer, pp 256–272
2020
Earlier work this paper cites.
Saito M, Saito S, Koyama M, Kobayashi S (2020) Train sparsely, generate densely: Memory-efficient unsupervised training of high-resolution temporal gan. IJCV 128(10-11):2586–2606
2020
Earlier work this paper cites.
Sheng L, Pan J, Guo J, Shao J, Loy CC (2020) High-quality video generation from static structural annotations. IJCV 128:2552–2569
2020
Cited alongside, same era.
Wang Y, Bilinski P, Bremond F, Dantcheva A (2020) G3an: Disentangling appearance and motion for video generation. In: CVPR, pp 5264–5273
2020
Cited alongside, same era.
Zhang J, Xu C, Liu L, Wang M, Wu X, Liu Y, Jiang Y (2020) Dtvnet: Dynamic time-lapse video generation via single still image. In: ECCV, Springer, pp 300–315
2020
Cited alongside, same era.
Babaeizadeh M, Saffar MT, Nair S, Levine S, Finn C, Erhan D (2021) Fitvid: Overfitting in pixel-level video prediction. arXiv preprint arXiv:210613195
2021
Cited alongside, same era.
Blattmann A, Milbich T, Dorkenwald M, Ommer B (2021) Understanding object dynamics for interactive image-to-video synthesis. In: CVPR, pp 5171–5181
2021
Zheng C, Vuong LT, Cai J, Phung D (2022) MoVQ: Modulating quantized vectors for high-fidelity image generation. In: NeurIPS
2022
Later among the works it cites.
Zhou D, Wang W, Yan H, Lv W, Zhu Y, Feng J (2022) Magicvideo: Efficient video generation with latent diffusion models. arXiv preprint arXiv:221111018
2022
Later among the works it cites.
Chen H, Xia M, He Y, Zhang Y, Cun X, Yang S, Xing J, Liu Y, Chen Q, Wang X, et al (2023) Videocrafter1: Open diffusion models for high-quality video generation. arXiv preprint arXiv:231019512
2023
Closest in time.
Fu TJ, Yu L, Zhang N, Fu CY, Su JC, Wang WY, Bell S (2023) Tell me what happened: Unifying text-guided video completion via multimodal masked video generation. In: CVPR, pp 10681–10692
2023
Closest in time.
Han L, Li Y, Zhang H, Milanfar P, Metaxas D, Yang F (2023) SVDiff: Compact parameter space for diffusion fine-tuning. In: ICCV, pp 7323–7334
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dhariwal P, Nichol A (2021) Diffusion models beat gans on image synthesis. In: NeurIPS, vol 34, pp 8780–8794
2021
Cited alongside, same era.
Dorkenwald M, Milbich T, Blattmann A, Rombach R, Derpanis KG, Ommer B (2021) Stochastic image-to-video synthesis using cinns. In: CVPR, pp 3741–3752
2021
Cited alongside, same era.
Esser P, Rombach R, Ommer B (2021) Taming transformers for high-resolution image synthesis. In: CVPR, pp 12873–12883
2021
Cited alongside, same era.
Menapace W, Lathuilière S, Tulyakov S, Siarohin A, Ricci E (2021) Playable video generation. In: CVPR, pp 10061–10070
2021
Cited alongside, same era.
Nichol AQ, Dhariwal P (2021) Improved denoising diffusion probabilistic models. In: ICML, pp 8162–8171
2021
Cited alongside, same era.
Rakhimov R, Volkhonskiy D, Artemov A, Zorin D, Burnaev E (2021) Latent video transformer. In: VISIGRAPP, pp 101–112
2021
Cited alongside, same era.
Shrivastava G, Shrivastava A (2021) Diverse video generation using a gaussian process trigger. In: ICLR
2021
Cited alongside, same era.
2023
Closest in time.
Hong W, Ding M, Zheng W, Liu X, Tang J (2023) Cogvideo: Large-scale pretraining for text-to-video generation via transformers. In: ICLR
2023
Closest in time.
Hu Y, Luo C, Chen Z (2023) A benchmark for controllable text-image-to-video generation. IEEE TMM pp 1–14
2023
Closest in time.
Ni H, Shi C, Li K, Huang SX, Min MR (2023) Conditional image-to-video generation with latent flow diffusion models. In: CVPR, pp 18444–18455
2023
Closest in time.
Shen S, Zhao W, Meng Z, Li W, Zhu Z, Zhou J, Lu J (2023) Difftalk: Crafting diffusion models for generalized talking head synthesis. arXiv preprint arXiv:230103786
2023
Closest in time.
Singer U, Polyak A, Hayes T, Yin X, An J, Zhang S, Hu Q, Yang H, Ashual O, Gafni O, Parikh D, Gupta S, Taigman Y (2023) Make-a-video: Text-to-video generation without text-video data. In: ICLR
2023
Closest in time.
Sun M, Wang W, Zhu X, Liu J (2023) Moso: Decomposing motion, scene and object for video prediction. In: CVPR, pp 18727–18737
2023
Closest in time.
Villegas R, Babaeizadeh M, Kindermans PJ, Moraldo H, Zhang H, Saffar MT, Castro S, Kunze J, Erhan D (2023) Phenaki: Variable length video generation from open domain textual descriptions. In: ICLR
2023
Closest in time.
Xu X, Wang Y, Wang L, Yu B, Jia J (2023) Conditional temporal variational autoencoder for action video prediction. IJCV pp 1–24
2023
Closest in time.
Zhang S, Wang J, Zhang Y, Zhao K, Yuan H, Qin Z, Wang X, Zhao D, Zhou J (2023) I2VGen-XL: High-quality image-to-video synthesis via cascaded diffusion models. arXiv preprint arXiv:231104145
2023
Closest in time.
Zhu J, Yang H, He H, Wang W, Tuo Z, Cheng WH, Gao L, Song J, Fu J (2023) MovieFactory: Automatic movie creation from text using large generative models for language and images. In: ACM MM, p 9313–9319
2023
Closest in time.
Gal R, Vinker Y, Alaluf Y, Bermano AH, Cohen-Or D, Shamir A, Chechik G (2024) Breathing life into sketches using text-to-video priors. In: CVPR
2024
Closest in time.
Guo Y, Yang C, Rao A, Liang Z, Wang Y, Qiao Y, Agrawala M, Lin D, Dai B (2024) AnimateDiff: Animate your personalized text-to-image diffusion models without specific tuning. In: ICLR
2024
Closest in time.
Gupta A, Yu L, Sohn K, Gu X, Hahn M, Li FF, Essa I, Jiang L, Lezama J (2024) Photorealistic video generation with diffusion models. In: Leonardis A, Ricci E, Roth S, Russakovsky O, Sattler T, Varol G (eds) ECCV, pp 393–411
2024
Closest in time.
Hu L (2024) Animate anyone: Consistent and controllable image-to-video synthesis for character animation. In: CVPR, pp 8153–8163
2024
Closest in time.
Huang Z, He Y, Yu J, Zhang F, Si C, Jiang Y, Zhang Y, Wu T, Jin Q, Chanpaisit N, Wang Y, Chen X, Wang L, Lin D, Qiao Y, Liu Z (2024) Vbench: Comprehensive benchmark suite for video generative models. In: CVPR, pp 21807–21818
2024
Closest in time.
Ni H, Egger B, Lohit S, Cherian A, Wang Y, Koike-Akino T, Huang SX, Marks TK (2024) Ti2v-zero: Zero-shot image conditioning for text-to-video diffusion models. In: CVPR, pp 9015–9025
2024
Closest in time.
Shi X, Huang Z, Wang FY, Bian W, Li D, Zhang Y, Zhang M, Cheung KC, See S, Qin H, Dai J, Li H (2024) Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling. In: ACM SIGGRAPH, New York, NY, USA, SIGGRAPH ’24, DOI
2024
Closest in time.
Xing J, Xia M, Zhang Y, Chen H, Yu W, Liu H, Liu G, Wang X, Shan Y, Wong TT (2024) Dynamicrafter: Animating open-domain images with video diffusion priors. In: Leonardis A, Ricci E, Roth S, Russakovsky O, Sattler T, Varol G (eds) ECCV, pp 399–417
2024
Closest in time.
Zeng Y, Wei G, Zheng J, Zou J, Wei Y, Zhang Y, Li H (2024) Make pixels dance: High-dynamic video generation. In: CVPR, pp 8850–8860
2024
Closest in time.