Fetching the paper…
Reading the bibliography…
In this paper, we introduce PixArt-\Sigma, a Diffusion Transformer model~(DiT) capable of directly generating images at 4K resolution.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft COCO: Common objects in context. In: ECCV (2014)
2014
Earlier work this paper cites.
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., Hochreiter, S.: GANs trained by a two time-scale update rule converge to a local nash equilibrium. In: NeurIPS (2017)
2017
Earlier work this paper cites.
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. In: arXiv (2017)
2017
Earlier work this paper cites.
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al.: Improving language understanding by generative pre-training. OpenAI blog (2018)
2018
Earlier work this paper cites.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al.: Language models are unsupervised multitask learners. OpenAI blog (2019)
2019
Earlier work this paper cites.
Wang, S., Li, B.Z., Khabsa, M., Fang, H., Ma, H.: Linformer: Self-attention with linear complexity. In: arXiv (2020)
2020
Earlier work this paper cites.
Chen, T., Cheng, Y., Gan, Z., Yuan, L., Zhang, L., Wang, Z.: Chasing sparsity in vision transformers: An end-to-end exploration. In: NeurIPS (2021)
2021
Earlier work this paper cites.
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al.: Rethinking attention with performers. In: ICLR (2021)
2021
Earlier work this paper cites.
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B.: Swin transformer: Hierarchical vision transformer using shifted windows. In: ICCV (2021)
2021
Earlier work this paper cites.
Lu, J., Yao, J., Zhang, J., Zhu, X., Xu, H., Gao, W., Xu, C., Xiang, T., Zhang, L.: Soft: Softmax-free transformer with linear complexity. In: NeurIPS (2021)
2021
Earlier work this paper cites.
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., Jégou, H.: Training data-efficient image transformers & distillation through attention. In: ICML (2021)
2021
Earlier work this paper cites.
Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., Shao, L.: Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. In: ICCV (2021)
2021
Earlier work this paper cites.
Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J.M., Luo, P.: SegFormer: Simple and efficient design for semantic segmentation with transformers. In: NeurIPS (2021)
2021
Earlier work this paper cites.
Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z.H., Tay, F.E., Feng, J., Yan, S.: Tokens-to-token ViT: Training vision transformers from scratch on imagenet. In: ICCV (2021)
2021
Earlier work this paper cites.
Zheng, S., Lu, J., Zhao, H., Zhu, X., Luo, Z., Wang, Y., Fu, Y., Feng, J., Xiang, T., Torr, P.H., et al.: Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In: CVPR (2021)
2021
Earlier work this paper cites.
Zhu, X., Su, W., Lu, L., Li, B., Wang, X., Dai, J.: Deformable DETR: Deformable transformers for end-to-end object detection. ICLR (2021)
2021
Earlier work this paper cites.
Chung, H.W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al.: Scaling instruction-finetuned language models. In: arXiv (2022)
2022
Earlier work this paper cites.
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: CVPR (2022)
2022
Earlier work this paper cites.
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E.L., Ghasemipour, K., Gontijo Lopes, R., Karagol Ayan, B., Salimans, T., et al.: Photorealistic text-to-image diffusion models with deep language understanding. In: NeurIPS (2022)
2022
Earlier work this paper cites.
Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., Shao, L.: Pvt v2: Improved baselines with pyramid vision transformer. Computational Visual Media (2022)
2022
Cited alongside, same era.
Xia, Z., Pan, X., Song, S., Li, L.E., Huang, G.: Vision transformer with deformable attention. In: CVPR (2022)
2022
Cited alongside, same era.
Aesthetic predictor (2023), https://github.com/christophschuhmann/improved-aesthetic-predictor
2023
Cited alongside, same era.
Bao, F., Nie, S., Xue, K., Cao, Y., Li, C., Su, H., Zhu, J.: All are worth words: A vit backbone for diffusion models. In: CVPR (2023)
2023
Cited alongside, same era.
Chandra, A., Tünnermann, L., Löfstedt, T., Gratz, R.: Transformer-based deep learning for predicting protein properties in the life sciences. eLife Sciences Publications, Ltd (2023)
2023
OpenAI: Gpt-4v(ision) system card. In: OpenAI (2023)
2023
Later among the works it cites.
Peebles, W., Xie, S.: Scalable diffusion models with transformers. In: ICCV (2023)
2023
Later among the works it cites.
Pernias, P., Rampas, D., Richter, M.L., Pal, C., Aubreville, M.: Würstchen: An efficient architecture for large-scale text-to-image diffusion models. In: ICLR (2023)
2023
Later among the works it cites.
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., Rombach, R.: SDXL: Improving latent diffusion models for high-resolution image synthesis. In: arXiv (2023)
2023
Later among the works it cites.
Sauer, A., Lorenz, D., Blattmann, A., Rombach, R.: Adversarial diffusion distillation. In: arXiv (2023)
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Chen, L., Li, J., Dong, X., Zhang, P., He, C., Wang, J., Zhao, F., Lin, D.: ShareGPT4V: Improving large multi-modal models with better captions. In: arXiv (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Chen, X., Liu, Z., Tang, H., Yi, L., Zhao, H., Han, S.: SparseViT: Revisiting activation sparsity for efficient high-resolution vision transformer. In: CVPR (2023)
2023
Cited alongside, same era.
Gao, S., Zhou, P., Cheng, M.M., Yan, S.: Masked diffusion transformer is a strong image synthesizer. In: ICCV (2023)
2023
Cited alongside, same era.
Ge, C., Ding, X., Tong, Z., Yuan, L., Wang, J., Song, Y., Luo, P.: Advancing vision transformers with group-mix attention. In: arXiv (2023)
2023
Cited alongside, same era.
von Glehn, I., Spencer, J.S., Pfau, D.: A self-attention ansatz for ab-initio quantum chemistry. In: ICLR (2023)
2023
Cited alongside, same era.
Hatamizadeh, A., Song, J., Liu, G., Kautz, J., Vahdat, A.: Diffit: Diffusion vision transformers for image generation. In: arXiv (2023)
2023
Cited alongside, same era.
Xie, E., Yao, L., Shi, H., Liu, Z., Zhou, D., Liu, Z., Li, J., Li, Z.: DiffFit: Unlocking transferability of large diffusion models via simple parameter-efficient fine-tuning. In: ICCV (2023)
2023
Later among the works it cites.
Xue, Z., Song, G., Guo, Q., Liu, B., Zong, Z., Liu, Y., Luo, P.: Raphael: Text-to-image generation via large mixture of diffusion paths. In: NeurIPS (2023)
2023
Later among the works it cites.
Zheng, H., Nie, W., Vahdat, A., Anandkumar, A.: Fast training of diffusion models with masked transformers. In: arXiv (2023)
2023
Later among the works it cites.
Chen, J., Wu, Y., Luo, S., Xie, E., Paul, S., Luo, P., Zhao, H., Li, Z.: Pixart- δ \delta : Fast and controllable image generation with latent consistency models. In: arXiv (2024)
2024
Closest in time.
Chen, J., Yu, J., Ge, C., Yao, L., Xie, E., Wu, Y., Wang, Z., Kwok, J., Luo, P., Lu, H., Li, Z.: PixArt- α \alpha : Fast training of diffusion transformer for photorealistic text-to-image synthesis. In: ICLR (2024)
2024
Closest in time.
Du, R., Chang, D., Hospedales, T., Song, Y.Z., Ma, Z.: Demofusion: Democratising high-resolution image generation with no $$$. In: CVPR (2024)
2024
Closest in time.
He, Y., Yang, S., Chen, H., Cun, X., Xia, M., Zhang, Y., Wang, X., He, R., Chen, Q., Shan, Y.: ScaleCrafter: Tuning-free higher-resolution visual generation with diffusion models. In: ICLR (2024)
2024
Closest in time.
Li, D., Kamko, A., Akhgari, E., Sabet, A., Xu, L., Doshi, S.: Playground v2.5: Three insights towards enhancing aesthetic quality in text-to-image generation. In: arXiv (2024)
2024
Closest in time.
Lu, Z., Wang, Z., Huang, D., Wu, C., Liu, X., Ouyang, W., Bai, L.: Fit: Flexible vision transformer for diffusion model. In: arXiv (2024)
2024
Closest in time.
Ma, N., Goldstein, M., Albergo, M.S., Boffi, N.M., Vanden-Eijnden, E., Xie, S.: SiT: Exploring flow and diffusion-based generative models with scalable interpolant transformers. In: arXiv (2024)
2024
Closest in time.
OpenAI: Sora (2024), https://openai.com/sora
2024
Closest in time.
Stability.AI: Stable diffusion 3 (2024), https://stability.ai/news/stable-diffusion-3
2024
Closest in time.
Yin, T., Gharbi, M., Zhang, R., Shechtman, E., Durand, F., Freeman, W.T., Park, T.: One-step diffusion with distribution matching distillation. CVPR (2024)
2024
Closest in time.