Fetching the paper…
Reading the bibliography…
The first-in-first-out (FIFO) video diffusion, built on a pre-trained text-to-video model, has recently emerged as an effective approach for tuning-free long video generation.
Gaussian Temporal Awareness Networks for Action Localization
Long, F.; Yao, T.; Qiu, Z.; Tian, X.; Luo, J.; and Mei, T. 2019 · 2019
Earlier work this paper cites.
Learning Transferable Visual Models from Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Earlier work this paper cites.
SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations
Meng, C.; He, Y.; Song, Y.; Song, J.; Wu, J.; Zhu, J.-Y.; and Ermon, S. 2022 · 2022
Earlier work this paper cites.
MCVD-Masked Conditional Video Diffusion for Prediction, Generation, and Interpolation
Voleti, V.; Jolicoeur-Martineau, A.; and Pal, C. 2022 · 2022
Earlier work this paper cites.
Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation
An, J.; Zhang, S.; Yang, H.; Gupta, S.; Huang, J.-B.; Luo, J.; and Yin, X. 2023 · 2023
Earlier work this paper cites.
Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models
Blattmann, A.; Rombach, R.; Ling, H.; Dockhorn, T.; Kim, S. W.; Fidler, S.; and Kreis, K. 2023 · 2023
Earlier work this paper cites.
VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
Chen, H.; Xia, M.; He, Y.; Zhang, Y.; Cun, X.; Yang, S.; Xing, J.; Liu, Y.; Chen, Q.; Wang, X.; et al. 2023 · 2023
Earlier work this paper cites.
Diffusion Self-Guidance for Controllable Image Generation
Epstein, D.; Jabri, A.; Poole, B.; Efros, A. A.; and Holynski, A. 2023 · 2023
Earlier work this paper cites.
CogVideo: Large-Scale Pretraining for Text-to-Video Generation via Transformers
Hong, W.; Ding, M.; Zheng, W.; Liu, X.; and Tang, J. 2023 · 2023
Earlier work this paper cites.
Bi-calibration Networks for Weakly-Supervised Video Representation Learning
Long, F.; Yao, T.; Qiu, Z.; Tian, X.; Luo, J.; and Mei, T. 2023 · 2023
Earlier work this paper cites.
OpenAI. 2023 · 2023
Earlier work this paper cites.
Scalable Diffusion Models with Transformers
Peebles, W.; and Xie, S. 2023 · 2023
Earlier work this paper cites.
Make-A-Video: Text-to-Video Generation without Text-Video Data
Singer, U.; Polyak, A.; Hayes, T.; Yin, X.; An, J.; Zhang, S.; Hu, Q.; Yang, H.; Ashual, O.; Gafni, O.; et al. 2023 · 2023
Cited alongside, same era.
Phenaki: Variable Length Video Generation from Open Domain Textual Description
Villegas, R.; Babaeizadeh, M.; Kindermans, P.-J.; Moraldo, H.; Zhang, H.; Saffar, M. T.; Castro, S.; Kunze, J.; and Erhan, D. 2023 · 2023
Cited alongside, same era.
Freedom: Training-Free Energy-Guided Conditional Diffusion Model
Yu, J.; Wang, Y.; Zhao, C.; Ghanem, B.; and Zhang, J. 2023 · 2023
Cited alongside, same era.
DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L. M.; and Shum, H.-Y. 2023 · 2023
Cited alongside, same era.
OpenLEAF: Open-Domain Interleaved Image-Text Generation and Evaluation
An, J.; Yang, Z.; Li, L.; Wang, J.; Lin, K.; Liu, Z.; Wang, L.; and Luo, J. 2024 · 2024
Cited alongside, same era.
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
Jin, Y.; Sun, Z.; Xu, K.; Chen, L.; Jiang, H.; Huang, Q.; Song, C.; Liu, Y.; Zhang, D.; Song, Y.; Gai, K.; and Mu, Y. 2024 · 2024
Later among the works it cites.
FIFO-Diffusion: Generating Infinite Videos from Text without Training
Kim, J.; Kang, J.; Choi, J.; and Han, B. 2024 · 2024
Later among the works it cites.
VideoStudio: Generating Consistent-Content and Multi-Scene Videos
Long, F.; Qiu, Z.; Yao, T.; and Mei, T. 2024 · 2024
Later among the works it cites.
MEVG: Multi-event Video Generation with Text-to-Video Models
Oh, G.; Jeong, J.; Kim, S.; Byeon, W.; Kim, J.; Kim, S.; Kwon, H.; and Kim, S. 2024 · 2024
Later among the works it cites.
ConditionVideo: Training-Free Condition-Guided Video Generation
Peng, B.; Chen, X.; Wang, Y.; Lu, C.; and Qiao, Y. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training-Free Layout Control with Cross-Attention Guidance
Chen, M.; Laina, I.; and Vedaldi, A. 2024 · 2024
Cited alongside, same era.
Exploiting the Signal-Leak Bias in Diffusion Models
Everaert, M. N.; Fitsios, A.; Bocchio, M.; Arpa, S.; Süsstrunk, S.; and Achanta, R. 2024 · 2024
Cited alongside, same era.
AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Guo, Y.; Yang, C.; Rao, A.; Liang, Z.; Wang, Y.; Qiao, Y.; Agrawala, M.; Lin, D.; and Dai, B. 2024 · 2024
Cited alongside, same era.
StreamingT2V: Consistent, Dynamic, and Extendable Long Video Generation from Text
Henschel, R.; Khachatryan, L.; Hayrapetyan, D.; Poghosyan, H.; Tadevosyan, V.; Wang, Z.; Navasardyan, S.; and Shi, H. 2024 · 2024
Cited alongside, same era.
VBench: Comprehensive Benchmark Suite for Video Generative Models
Huang, Z.; He, Y.; Yu, J.; Zhang, F.; Si, C.; Jiang, Y.; Zhang, Y.; Wu, T.; Jin, Q.; Chanpaisit, N.; et al. 2024 · 2024
Cited alongside, same era.
PEEKABOO: Interactive Video Generation via Masked-Diffusion
Jain, Y.; Nasery, A.; Vineet, V.; and Behl, H. 2024 · 2024
Cited alongside, same era.
VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models
Chen, H.; Zhang, Y.; Cun, X.; Xia, M.; Wang, X.; Weng, C.; and Shan, Y. 2024a
Cited in the paper.
FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling
Qiu, H.; Xia, M.; Zhang, Y.; He, Y.; Wang, X.; Shan, Y.; and Liu, Z. 2024 · 2024
Later among the works it cites.
Video-Infinity: Distributed Long Video Generation
Tan, Z.; Yang, X.; Liu, S.; and Wang, X. 2024 · 2024
Later among the works it cites.
Training-free Consistent Text-to-Image Generation
Tewel, Y.; Kaduri, O.; Gal, R.; Kasten, Y.; Wolf, L.; Chechik, G.; and Atzmon, Y. 2024 · 2024
Later among the works it cites.
VideoTetris: Towards Compositional Text-to-Video Generation
Tian, Y.; Yang, L.; Yang, H.; Gao, Y.; Deng, Y.; Chen, J.; Wang, X.; Yu, Z.; Tao, X.; Wan, P.; et al. 2024 · 2024
Later among the works it cites.
EasyAnimate: A High-Performance Long Video Generation Method based on Transformer Architecture
Xu, J.; Zou, X.; Huang, K.; Chen, Y.; Liu, B.; Cheng, M.; Shi, X.; and Huang, J. 2024 · 2024
Later among the works it cites.
TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion Models
Zhang, Z.; Long, F.; Pan, Y.; Qiu, Z.; Yao, T.; Cao, Y.; and Mei, T. 2024 · 2024
Later among the works it cites.