Fetching the paper…
Reading the bibliography…
We present Vchitect-2.0, a parallel transformer architecture designed to scale up video diffusion models for large-scale text-to-video generation.
1909
Earlier work this paper cites.
1910
Earlier work this paper cites.
2006
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,”
2020
Earlier work this paper cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,”
2020
Earlier work this paper cites.
J. L. Chien-Chin Huang, Gu Jin, “Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping,”
2020
Earlier work this paper cites.
Z. Teed and J. Deng, “Raft: Recurrent all-pairs field transforms for optical flow,” in
2020
Earlier work this paper cites.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in
2021
Earlier work this paper cites.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in
2021
Earlier work this paper cites.
M. Ding, Z. Yang, W. Hong, W. Zheng, C. Zhou, D. Yin, J. Lin, X. Zou, Z. Shao, H. Yang
2021
Earlier work this paper cites.
J. Lin, R. Men, A. Yang, C. Zhou, M. Ding, Y. Zhang, P. Wang, A. Wang, L. Jiang, X. Jia
2021
Earlier work this paper cites.
P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,”
2021
Earlier work this paper cites.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” in
2021
Earlier work this paper cites.
B. Koonce and B. Koonce, “Efficientnet,”
2021
Earlier work this paper cites.
J. Yu, Y. Xu, J. Y. Koh, T. Luong, G. Baid, Z. Wang, V. Vasudevan, A. Ku, Y. Yang, B. K. Ayan
2022
Cited alongside, same era.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in
2022
Cited alongside, same era.
U. Singer, A. Polyak, T. Hayes, X. Yin, J. An, S. Zhang, Q. Hu, H. Yang, O. Ashual, O. Gafni
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Y. Zhao, A. Gu, R. Varma, L. Luo, C. chin Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, B. Nguyen, G. Chauhan, Y. Hao, and S. Li, “Pytorch fsdp: Experiences on scaling fully sharded data parallel,” pp. 3848–3860, 2023. [Online]. Available:
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Wang, Y. He, Y. Li, K. Li, J. Yu, X. Ma, X. Li, G. Chen, X. Chen, Y. Wang
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, and T. Salimans, “Imagen video: High definition video generation with diffusion models,” 2022
2022
Cited alongside, same era.
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, “Video diffusion models,” 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in
2023
Cited alongside, same era.
J. Chen, J. Yu, C. Ge, L. Yao, E. Xie, Y. Wu, Z. Wang, J. Kwok, P. Luo, H. Lu
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
T.-S. Chen, A. Siarohin, W. Menapace, E. Deyneka, H. wei Chao, B. E. Jeon, Y. Fang, H.-Y. Lee, J. Ren, M.-H. Yang, and S. Tulyakov, “Panda-70m: Captioning 70m videos with multiple cross-modality teachers,” 2024
2024
Later among the works it cites.
H. Chen, Y. Zhang, X. Cun, M. Xia, X. Wang, C. Weng, and Y. Shan, “Videocrafter2: Overcoming data limitations for high-quality video diffusion models,” 2024
2024
Later among the works it cites.
NVIDIA, “Context parallelism overview,” 2024, accessed: 2024-11-15. [Online]. Available:
2024
Later among the works it cites.
J. Fang and S. Zhao, “Usp: A unified sequence parallelism approach for long context generative ai,”
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Huang, Y. He, J. Yu, F. Zhang, C. Si, Y. Jiang, Y. Zhang, T. Wu, Q. Jin, N. Chanpaisit
2024
Later among the works it cites.
Y. Zhang, B. Li, h. Liu, Y. j. Lee, L. Gui, D. Fu, J. Feng, Z. Liu, and C. Li, “Llava-next: A strong zero-shot video understanding model,” April 2024. [Online]. Available:
2024
Later among the works it cites.
K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Luo
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
K. AI, “Kling ai official website,” 2024, accessed: 2024-12-10. [Online]. Available:
2024
Later among the works it cites.
Z. Yang, J. Teng, W. Zheng, M. Ding, S. Huang, J. Xu, Y. Yang, W. Hong, X. Zhang, G. Feng
2024
Later among the works it cites.
P.-Y. Lab and T. A. etc., “Open-sora-plan,” Apr. 2024. [Online]. Available:
2024
Later among the works it cites.