Fetching the paper…
Reading the bibliography…
Recent advancements have established Diffusion Transformers (DiTs) as a dominant framework in generative modeling.
The laplacian pyramid as a compact image code
Burt, P. J. and Adelson, E. H · 1983
Earlier work this paper cites.
Pyramid methods in image processing
Adelson, E. H., Burt, P. J., Anderson, C. H., Ogden, J. M., and Bergen, J. R · 1984
Earlier work this paper cites.
The structure of images
Koenderink, J. J · 2004
Earlier work this paper cites.
Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture
Eigen, D. and Fergus, R · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
A unified multi-scale deep convolutional neural network for fast object detection
Cai, Z., Fan, Q., Feris, R. S., and Vasconcelos, N · 2016
Earlier work this paper cites.
Attention to scale: Scale-aware semantic image segmentation
Chen, L., Yang, Y., Wang, J., Xu, W., and Yuille, A. L · 2016
Earlier work this paper cites.
Feature pyramid networks for object detection
Lin, T., Dollár, P., Girshick, R. B., He, K., Hariharan, B., and Belongie, S. J · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Temporal generative adversarial nets with singular value clipping
Saito, M., Matsumoto, E., and Saito, S · 2017
Earlier work this paper cites.
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
Chen, L., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A. L · 2018
Earlier work this paper cites.
Mocogan: Decomposing motion and content for video generation
Tulyakov, S., Liu, M., Yang, X., and Kautz, J · 2018
Earlier work this paper cites.
End-to-end object detection with transformers
Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S · 2020
Earlier work this paper cites.
Vggsound: A large-scale audio-visual dataset
Chen, H., Xie, W., Vedaldi, A., and Zisserman, A · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Kong, J., Kim, J., and Bae, J · 2020
Earlier work this paper cites.
Multi-scale interactive network for salient object detection
Pang, Y., Zhao, X., Zhang, L., and Lu, H · 2020
Earlier work this paper cites.
Feature pyramid transformer
Zhang, D., Zhang, H., Tang, J., Wang, M., Hua, X., and Sun, Q · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Earlier work this paper cites.
Multiscale vision transformers
Fan, H., Xiong, B., Mangalam, K., Li, Y., Yan, Z., Malik, J., and Feichtenhofer, C · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2021
Earlier work this paper cites.
Multi-scale high-resolution vision transformer for semantic segmentation
Gu, J., Kwon, H., Wang, D., Ye, W., Li, M., Chen, Y., Lai, L., Chandra, V., and Pan, D. Z · 2022
Earlier work this paper cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R. B · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2022
Cited alongside, same era.
Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2
Skorokhodov, I., Tulyakov, S., and Elhoseiny, M · 2022
Cited alongside, same era.
All are worth words: A vit backbone for diffusion models
Bao, F., Nie, S., Xue, K., Cao, Y., Li, C., Su, H., and Zhu, J · 2023
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., English, Z., Voleti, V., Letts, A., et al · 2023
Cited alongside, same era.
Kling ai, 2024
Kuaishou · 2024
Later among the works it cites.
Open-sora-plan, April 2024
Lab, P.-Y. and etc., T. A · 2024
Later among the works it cites.
Pika, 2024
Labs, P · 2024
Later among the works it cites.
Video-foley: Two-stage video-to-sound generation via temporal event condition for foley sound
Lee, J., Im, J., Kim, D., and Nam, J · 2024
Later among the works it cites.
Lipman, Y., Havasi, M., Holderrieth, P., Shaul, N., Le, M., Karrer, B., Chen, R. T., Lopez-Paz, D., Ben-Hamu, H., and Gat, I · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fang, A., Jose, A. M., Jain, A., Schmidt, L., Toshev, A., and Shankar, V · 2023
Cited alongside, same era.
Preserve your own correlation: A noise prior for video diffusion models
Ge, S., Nah, S., Liu, G., Poon, T., Tao, A., Catanzaro, B., Jacobs, D., Huang, J., Liu, M., and Balaji, Y · 2023
Cited alongside, same era.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Hong, W., Ding, M., Zheng, W., Liu, X., and Tang, J · 2023
Cited alongside, same era.
Flow matching for generative modeling
Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Cited alongside, same era.
Phenaki: Variable length video generation from open domain textual descriptions
Villegas, R., Babaeizadeh, M., Kindermans, P., Moraldo, H., Zhang, H., Saffar, M. T., Castro, S., Kunze, J., and Erhan, D · 2023
Cited alongside, same era.
Revisiting weak-to-strong consistency in semi-supervised semantic segmentation
Yang, L., Qi, L., Feng, L., Zhang, W., and Shi, Y · 2023
Cited alongside, same era.
Liu, D., Zhao, S., Zhuo, L., Lin, W., Qiao, Y., Li, H., and Gao, P · 2024
Later among the works it cites.
Dreammachine, 2024
LumaLab · 2024
Later among the works it cites.
Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models
Luo, S., Yan, C., Hu, C., and Zhao, H · 2024
Later among the works it cites.
Openvid-1m: A large-scale high-quality dataset for text-to-video generation
Nan, K., Xie, R., Zhou, P., Fan, T., Yang, Z., Chen, Z., Li, X., Yang, J., and Tai, Y · 2024
Later among the works it cites.
Sora, 2024
OpenAI · 2024
Later among the works it cites.
Hierarchical patch diffusion models for high-resolution video generation
Skorokhodov, I., Menapace, W., Siarohin, A., and Tulyakov, S · 2024
Later among the works it cites.
Journeydb: A benchmark for generative image understanding
Sun, K., Pan, J., Ge, Y., Li, H., Duan, H., Wu, X., Zhang, R., Zhou, A., Qin, Z., Wang, Y., et al · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., et al · 2024
Later among the works it cites.
V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models
Wang, H., Ma, J., Pascual, S., Cartwright, R., and Cai, W · 2024
Later among the works it cites.
Sonicvisionlm: Playing sound with vision language models
Xie, Z., Yu, S., He, Q., and Li, M · 2024
Later among the works it cites.
Cogvideox: Text-to-video diffusion models with an expert transformer
Yang, Z., Teng, J., Zheng, W., Ding, M., Huang, S., Xu, J., Yang, Y., Hong, W., Zhang, X., Feng, G., et al · 2024
Later among the works it cites.
Diverse and aligned audio-to-video generation via text-to-video model adaptation
Yariv, G., Gat, I., Benaim, S., Wolf, L., Schwartz, I., and Adi, Y · 2024
Later among the works it cites.
Language model beats diffusion - tokenizer is key to visual generation
Yu, L., Lezama, J., Gundavarapu, N. B., Versari, L., Sohn, K., Minnen, D., Cheng, Y., Gupta, A., Gu, X., Hauptmann, A. G., Gong, B., Yang, M., Essa, I., Ross, D. A., and Jiang, L · 2024
Later among the works it cites.
Allegro: Open the black box of commercial-level video generation model
Zhou, Y., Wang, Q., Cai, Y., and Yang, H · 2024
Later among the works it cites.
Lumina-next: Making lumina-t2x stronger and faster with next-dit
Zhuo, L., Du, R., Xiao, H., Li, Y., Liu, D., Huang, R., Liu, W., Zhao, L., Wang, F.-Y., Ma, Z., et al · 2024
Later among the works it cites.
Vchitect-2.0: Parallel transformer for scaling up video diffusion models
Fan, W., Si, C., Song, J., Yang, Z., He, Y., Zhuo, L., Huang, Z., Dong, Z., He, J., Pan, D., et al · 2025
Closest in time.
Gen-3 alpha
Runway Research · 2025
Closest in time.
Vidu: Online video learning platform
VIDU · 2025
Closest in time.