Fetching the paper…
Reading the bibliography…
Significant advancements have been made in the field of video generation, with the open-source community contributing a wealth of research papers and tools for training high-quality models.
Image quality assessment: from error visibility to structural similarity
Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli · 2004
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma · 2013
Earlier work this paper cites.
FiLM: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang · 2018
Earlier work this paper cites.
Character region awareness for text detection
Y. Baek, B. Lee, D. Han, S. Yun, and H. Lee · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
mT5: A massively multilingual pre-trained text-to-text transformer
L. Xue · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
M. Bain, A. Nagrani, G. Varol, and A. Zisserman · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Walt: Watch and learn 2d amodal representation from time-lapse imagery
N. Dinesh Reddy, R. Tamburo, and S. G. Narasimhan · 2022
Earlier work this paper cites.
Video diffusion models
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet · 2022
Earlier work this paper cites.
xFormers: A modular and hackable transformer modelling library
B. Lefaudeux, F. Massa, D. Liskovich, W. Xiong, V. Caggiano, S. Naren, M. Xu, J. Hu, M. Tintore, S. Zhang, P. Labatut, D. Haziza, L. Wehrstedt, J. Reizenstein, and G. Sizov · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Earlier work this paper cites.
LAION-5B: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Earlier work this paper cites.
Make-A-Video: Text-to-video generation without text-video data
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni, et al · 2022
Earlier work this paper cites.
MCVD: Masked conditional video diffusion for prediction, generation, and interpolation
V. Voleti, A. Jolicoeur-Martineau, and C. Pal · 2022
Cited alongside, same era.
Watermark-Detection
Watermark-Detection Contributors · 2022
Cited alongside, same era.
Advancing high-resolution video-language representation with large-scale video transcriptions
H. Xue, T. Hang, Y. Zeng, Y. Sun, B. Liu, H. Yang, J. Fu, and B. Guo · 2022
Cited alongside, same era.
Align your latents: High-resolution video synthesis with latent diffusion models
A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis · 2023
Cited alongside, same era.
PixArt- α \alpha : Fast training of diffusion transformer for photorealistic text-to-image synthesis
J. Chen, J. Yu, C. Ge, L. Yao, E. Xie, Y. Wu, Z. Wang, J. Kwok, P. Luo, H. Lu, et al · 2023
Cited alongside, same era.
Scaling rectified flow transformers for high-resolution image synthesis
P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, et al · 2024
Closest in time.
DreamStory: Open-domain story visualization by llm-guided multi-subject consistent diffusion
H. He, H. Yang, Z. Tuo, Y. Zhou, Q. Wang, Y. Zhang, Z. Liu, W. Huang, H. Chao, and J. Yin · 2024
Closest in time.
VBench: Comprehensive benchmark suite for video generative models
Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit, et al · 2024
Closest in time.
Kling AI
Kuaishou · 2024
Closest in time.
Dream Machine
Luma Lab · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Dao · 2023
Cited alongside, same era.
Tag2Text: Guiding vision-language model via image tagging
X. Huang, Y. Zhang, J. Ma, W. Tian, R. Feng, Y. Zhang, Y. Li, Y. Guo, and L. Zhang · 2023
Cited alongside, same era.
S. A. Jacobs, M. Tanaka, C. Zhang, M. Zhang, S. L. Song, S. Rajbhandari, and Y. He · 2023
Cited alongside, same era.
Ring attention with blockwise transformers for near-infinite context
H. Liu, M. Zaharia, and P. Abbeel · 2023
Cited alongside, same era.
VDT: General-purpose video diffusion transformers via mask modeling
H. Lu, G. Yang, N. Fei, Y. Huo, Z. Lu, P. Luo, and M. Ding · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Cited alongside, same era.
MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation
L. Ruan, Y. Ma, H. Yang, H. He, B. Liu, J. Fu, N. J. Yuan, Q. Jin, and B. Guo · 2023
Cited alongside, same era.
X. Ma, Y. Wang, G. Jia, X. Chen, Z. Liu, Y.-F. Li, C. Chen, and Y. Qiao · 2024
Closest in time.
Hailuo AI
MiniMax · 2024
Closest in time.
OpenVid-1M: A large-scale high-quality dataset for text-to-video generation
K. Nan, R. Xie, P. Zhou, T. Fan, Z. Yang, Z. Chen, X. Li, J. Yang, and Y. Tai · 2024
Closest in time.
Open-Sora-Plan
PKU-Yuan Lab and Tuzhan AI etc · 2024
Closest in time.
SDXL: Improving latent diffusion models for high-resolution image synthesis
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockhorn, J. Müller, J. Penna, and R. Rombach · 2024
Closest in time.
Movie gen: A cast of media foundation models, 2024
A. Polyak, A. Zohar, A. Brown, A. Tjandra, A. Sinha, A. Lee, A. Vyas, B. Shi, C.-Y. Ma, C.-Y. Chuang, D. Yan, D. Choudhary, D. Wang, G. Sethi, G. Pang, H. Ma, I. Misra, J. Hou, J. Wang, K. Jagadeesh, K. Li, L. Zhang, M. Singh, M. Williamson, M. Le, M. Yu, M. K. Singh, P. Zhang, P. Vajda, Q. Duval, R. Girdhar, R. Sumbaly, S. S. Rambhatla, S. Tsai, S. Azadi, S. Datta, S. Chen, S. Bell, S. Ramaswamy, S. Sheynin, S. Bhattacharya, S. Motwani, T. Xu, T. Li, T. Hou, W.-N. Hsu, X. Yin, X. Dai, Y. Taigman, Y. Luo, Y.-C. Liu, Y.-C. Wu, Y. Zhao, Y. Kirstain, Z. He, Z. He, A. Pumarola, A. Thabet, A. Sanakoyeu, A. Mallya, B. Guo, B. Araya, B. Kerr, C. Wood, C. Liu, C. Peng, D. Vengertsev, E. Schonfeld, E. Blanchard, F. Juefei-Xu, F. Nord, J. Liang, J. Hoffman, J. Kohler, K. Fire, K. Sivakumar, L. Chen, L. Yu, L. Gao, M. Georgopoulos, R. Moritz, S. K. Sampson, S. Li, S. Parmeggiani, S. Fine, T. Fowler, V. Petrovic, and Y. Du · 2024
Closest in time.
PySceneDetect
PySceneDetect Contributors · 2024
Closest in time.
Gen-3 Alpha
RunwayML · 2024
Closest in time.
T2V-CompBench: A comprehensive benchmark for compositional text-to-video generation
K. Sun, K. Huang, X. Liu, Y. Wu, Z. Xu, Z. Li, and X. Liu · 2024
Closest in time.
VideoComposer: Compositional video synthesis with motion controllability
X. Wang, H. Yuan, S. Zhang, D. Chen, J. Wang, Y. Zhang, Y. Shen, D. Zhao, and J. Zhou · 2024
Closest in time.
CogVideoX: Text-to-video diffusion models with an expert transformer
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng, et al · 2024
Closest in time.
Open-Sora: Democratizing efficient video production for all
Z. Zheng, X. Peng, T. Yang, C. Shen, S. Li, H. Liu, Y. Zhou, T. Li, and Y. You · 2024
Closest in time.