Fetching the paper…
Reading the bibliography…
This paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance.
Overview of the h. 264/avc video coding standard
Wiegand, T., Sullivan, G. J., Bjontegaard, G., and Luthra, A. (2003) · 2003
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A. R., and Shah, M. (2012) · 2012
Earlier work this paper cites.
Image hash
Contributors, I. H. (2013) · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. (2013) · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C. (2016) · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J. and Zisserman, A. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Ha, D. and Schmidhuber, J. (2018) · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
Unterthiner, T., Van Steenkiste, S., Kurach, K., Marinier, R., Michalski, M., and Gelly, S. (2018) · 2018
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R. (2019) · 2019
Earlier work this paper cites.
An advert creation system for 3d product placements
Bacher, I., Javidnia, H., Dev, S., Agrahari, R., Hossari, M., Nicholson, M., Conran, C., Tang, J., Song, P., Corrigan, D., et al. (2021) · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P. (2020) · 2020
Earlier work this paper cites.
Raft: Recurrent all-pairs field transforms for optical flow
Teed, Z. and Deng, J. (2020) · 2020
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Esser, P., Rombach, R., and Ommer, B. (2021) · 2021
Earlier work this paper cites.
Sequence parallelism: Long sequence training from system perspective
Li, S., Xue, F., Baranwal, C., Li, Y., and You, Y. (2021) · 2021
Earlier work this paper cites.
Nvidia h100 tensor core gpu architecture
Corporation, N. (2022) · 2022
Earlier work this paper cites.
Latent video diffusion models for high-fidelity video generation with arbitrary lengths
He, Y., Yang, T., Zhang, Y., Shan, Y., and Chen, Q. (2022) · 2022
Earlier work this paper cites.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J. (2022) · 2022
Earlier work this paper cites.
Dall· e 2 preview-risks and limitations
Mishkin, P., Ahmad, L., Brundage, M., Krueger, G., and Sastry, G. (2022) · 2022
Earlier work this paper cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. (2022) · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022) · 2022
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al. (2022) · 2022
Earlier work this paper cites.
Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2
Skorokhodov, I., Tulyakov, S., and Elhoseiny, M. (2022) · 2022
Earlier work this paper cites.
Magicvideo: Efficient video generation with latent diffusion models
Zhou, D., Wang, W., Yan, H., Lv, W., Zhu, Y., and Feng, J. (2022) · 2022
Earlier work this paper cites.
Building normalizing flows with stochastic interpolants
Albergo, M. S. and Vanden-Eijnden, E. (2023) · 2023
Earlier work this paper cites.
Improving image generation with better captions
Betker, J., Goh, G., Jing, L., Brooks, T., Wang, J., Li, L., Ouyang, L., Zhuang, J., Lee, J., Guo, Y., et al. (2023) · 2023
Cited alongside, same era.
Pixart-alphaalpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis
Chen, J., Yu, J., Ge, C., Yao, L., Xie, E., Wu, Y., Wang, Z., Kwok, J., Luo, P., Lu, H., et al. (2023) · 2023
Cited alongside, same era.
Nvidia announces dgx gh200 ai supercomputer
Corporation, N. (2023) · 2023
Cited alongside, same era.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A. P., Caron, M., Geirhos, R., Alabdulmohsin, I., et al. (2023) · 2023
Cited alongside, same era.
Preserve your own correlation: A noise prior for video diffusion models
Ge, S., Nah, S., Liu, G., Poon, T., Tao, A., Catanzaro, B., Jacobs, D., Huang, J.-B., Liu, M.-Y., and Balaji, Y. (2023) · 2023
Cited alongside, same era.
Venhancer: Generative space-time enhancement for video generation
He, J., Xue, T., Liu, D., Lin, X., Gao, P., Lin, D., Qiao, Y., Ouyang, W., and Liu, Z. (2024) · 2024
Later among the works it cites.
Ella: Equip diffusion models with llm for enhanced semantic alignment
Hu, X., Wang, R., Fang, Y., Fu, B., Cheng, P., and Yu, G. (2024) · 2024
Later among the works it cites.
Vbench: Comprehensive benchmark suite for video generative models
Huang, Z., He, Y., Yu, J., Zhang, F., Si, C., Jiang, Y., Zhang, Y., Wu, T., Jin, Q., Chanpaisit, N., et al. (2024) · 2024
Later among the works it cites.
Prompt-a-video: Prompt your video diffusion model via preference-aligned llm
Ji, Y., Zhang, J., Wu, J., Zhang, S., Chen, S., GE, C., Sun, P., Chen, W., Shao, W., Xiao, X., et al. (2024) · 2024
Later among the works it cites.
Megascale: Scaling large language model training to more than 10,000 gpus
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Girdhar, R., Singh, M., Brown, A., Duval, Q., Azadi, S., Rambhatla, S. S., Shah, A., Yin, X., Parikh, D., and Misra, I. (2023) · 2023
Cited alongside, same era.
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation
Huang, K., Sun, K., Xie, E., Li, Z., and Liu, X. (2023) · 2023
Cited alongside, same era.
Jacobs, S. A., Tanaka, M., Zhang, C., Zhang, M., Song, S. L., Rajbhandari, S., and He, Y. (2023) · 2023
Cited alongside, same era.
Reducing activation recomputation in large transformer models
Korthikanti, V. A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B. (2023) · 2023
Cited alongside, same era.
Flow matching for generative modeling
Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M. (2023) · 2023
Cited alongside, same era.
Flow straight and fast: Learning to generate and transfer data with rectified flow
Liu, X., Gong, C., and qiang liu (2023) · 2023
Cited alongside, same era.
Dinov2: Learning robust visual features without supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al. (2023) · 2023
Cited alongside, same era.
Jiang, Z., Lin, H., Zhong, Y., Huang, Q., Chen, Y., Zhang, Z., Peng, Y., Li, X., Xie, C., Nong, S., et al. (2024) · 2024
Later among the works it cites.
Pyramidal flow matching for efficient video generative modeling
Jin, Y., Sun, Z., Li, N., Xu, K., Jiang, H., Zhuang, N., Huang, Q., Song, Y., Mu, Y., and Lin, Z. (2024) · 2024
Later among the works it cites.
Miradata: A large-scale video dataset with long durations and structured captions
Ju, X., Gao, Y., Zhang, Z., Yuan, Z., Wang, X., Zeng, A., Xiong, Y., Xu, Q., and Shan, Y. (2024) · 2024
Later among the works it cites.
Hunyuanvideo: A systematic framework for large video generative models
Kong, W., Tian, Q., Zhang, Z., Min, R., Dai, Z., Zhou, J., Xiong, J., Li, X., Wu, B., Zhang, J., et al. (2024) · 2024
Later among the works it cites.
Kling ai
Kuaishou (2024) · 2024
Later among the works it cites.
Open-sora-plan
Lab, P.-Y. and etc., T. A. (2024) · 2024
Later among the works it cites.
Chameleon: Plug-and-play compositional reasoning with large language models
Lu, P., Peng, B., Cheng, H., Galley, M., Chang, K.-W., Wu, Y. N., Zhu, S.-C., and Gao, J. (2024) · 2024
Later among the works it cites.
Openvid-1m: A large-scale high-quality dataset for text-to-video generation
Nan, K., Xie, R., Zhou, P., Fan, T., Yang, Z., Chen, Z., Li, X., Yang, J., and Tai, Y. (2024) · 2024
Later among the works it cites.
Movie gen: A cast of media foundation models
Polyak, A., Zohar, A., Brown, A., Tjandra, A., Sinha, A., Lee, A., Vyas, A., Shi, B., Ma, C.-Y., Chuang, C.-Y., et al. (2024) · 2024
Later among the works it cites.
Oasis: A universe in a transformer
Quevedo, J., McIntyre, Q., Campbell, S., and Wachen, R. (2024) · 2024
Later among the works it cites.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T. (2024) · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. (2024) · 2024
Later among the works it cites.
Autoregressive model beats diffusion: Llama for scalable image generation
Sun, P., Jiang, Y., Chen, S., Zhang, S., Peng, B., Luo, P., and Yuan, Z. (2024) · 2024
Later among the works it cites.
Vchitect-2.0: Parallel transformer for scaling up video diffusion models
Team, V. (2024) · 2024
Later among the works it cites.
Diffusion models are real-time game engines
Valevski, D., Leviathan, Y., Arar, M., and Fruchter, S. (2024) · 2024
Later among the works it cites.
Bytecheckpoint: A unified checkpointing system for llm development
Wan, B., Han, M., Sheng, Y., Lai, Z., Zhang, M., Zhang, J., Peng, Y., Lin, H., Liu, X., and Wu, C. (2024) · 2024
Later among the works it cites.
Show-o: One single transformer to unify multimodal understanding and generation
Xie, J., Mao, W., Bai, Z., Zhang, D. J., Wang, W., Lin, K. Q., Gu, Y., Chen, Z., Yang, Z., and Shou, M. Z. (2024) · 2024
Later among the works it cites.
Make pixels dance: High-dynamic video generation
Zeng, Y., Wei, G., Zheng, J., Zou, J., Wei, Y., Zhang, Y., and Li, H. (2024) · 2024
Later among the works it cites.
Virbo: Multimodal multilingual avatar video generation in digital marketing
Zhang, J., Chen, J., Wang, C., Yu, Z., Qi, T., Liu, C., and Wu, D. (2024) · 2024
Later among the works it cites.
Open-sora: Democratizing efficient video production for all
Zheng, Z., Peng, X., Yang, T., Shen, C., Li, S., Liu, H., Zhou, Y., Li, T., and You, Y. (2024) · 2024
Later among the works it cites.
Transfusion: Predict the next token and diffuse images with one multi-modal model
Zhou, C., Yu, L., Babu, A., Tirumala, K., Yasunaga, M., Shamis, L., Kahn, J., Ma, X., Zettlemoyer, L., and Levy, O. (2024) · 2024
Later among the works it cites.
Cosmos world foundation model platform for physical ai
Agarwal, N., Ali, A., Bala, M., Balaji, Y., Barker, E., Cai, T., Chattopadhyay, P., Chen, Y., Cui, Y., Ding, Y., et al. (2025) · 2025
Closest in time.
Yuan, L., Wang, J., Sun, H., Zhang, Y., and Lin, Y. (2025) · 2025
Closest in time.