Fetching the paper…
Reading the bibliography…
This paper presents PipeFusion, an innovative parallel methodology to tackle the high latency issues associated with generating high-resolution images using diffusion transformers (DiTs) models.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Ronneberger, O., Fischer, P., and Brox, T · 2015
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J., Meng, C., and Ermon, S · 2020
Earlier work this paper cites.
Terapipe: Token-level pipeline parallelism for training large-scale language models
Li, Z., Zhuang, S., Guo, S., Zhuo, D., Zhang, H., Song, D., and Stoica, I · 2021
Earlier work this paper cites.
Classifier-free diffusion guidance
Ho, J. and Salimans, T · 2022
Earlier work this paper cites.
Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps
Lu, C., Zhou, Y., Bao, F., Chen, J., Li, C., and Zhu, J · 2022
Earlier work this paper cites.
On aliased resizing and surprising subtleties in gan evaluation
Parmar, G., Zhang, R., and Zhu, J.-Y · 2022
Cited alongside, same era.
Pixart- α \alpha : Fast training of diffusion transformer for photorealistic text-to-image synthesis
Chen, J., Yu, J., Ge, C., Yao, L., Xie, E., Wu, Y., Wang, Z., Kwok, J., Luo, P., Lu, H., et al · 2023
Cited alongside, same era.
Jacobs, S. A., Tanaka, M., Zhang, C., Zhang, M., Song, L., Rajbhandari, S., and He, Y · 2023
Cited alongside, same era.
Lightseq: Sequence level parallelism for distributed training of long context transformers
Li, D., Shao, R., Xie, A., Xing, E. P., Gonzalez, J. E., Stoica, I., Ma, X., and Zhang, H · 2023
Cited alongside, same era.
Ring attention with blockwise transformers for near-infinite context
Scaling rectified flow transformers for high-resolution image synthesis
Esser, P., Kulal, S., Blattmann, A., Entezari, R., Müller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boesel, F., et al · 2024
Closest in time.
A unified sequence parallelism approach for long context generative ai
Fang, J. and Zhao, S · 2024
Closest in time.
Megascale: Scaling large language model training to more than 10,000 gpus
Jiang, Z., Lin, H., Zhong, Y., Huang, Q., Chen, Y., Zhang, Z., Peng, Y., Li, X., Xie, C., Nong, S., et al · 2024
Closest in time.
Distrifusion: Distributed parallel inference for high-resolution diffusion models
Li, M., Cai, T., Cao, J., Zhang, Q., Cai, H., Bai, J., Jia, Y., Liu, M.-Y., Li, K., and Han, S · 2024
Closest in time.
Movie gen: A cast of media foundation models
MetaAI · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Liu, H., Zaharia, M., and Abbeel, P · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Cited alongside, same era.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2023
Cited alongside, same era.
Frdiff: Feature reuse for exquisite zero-shot acceleration of diffusion models
So, J., Lee, J., and Park, E · 2023
Cited alongside, same era.
Announcing black forest labs
BlackForestLabs · 2024
Cited alongside, same era.
Chen, J., Ge, C., Xie, E., Wu, Y., Yao, L., Ren, X., Wang, Z., Luo, P., Lu, H., and Li, Z
Cited in the paper.
Chen, P., Shen, M., Ye, P., Cao, J., Tu, C., Bouganis, C.-S., Zhao, Y., and Chen, T
Cited in the paper.
Learning-to-cache: Accelerating diffusion transformer via layer caching
Ma, X., Fang, G., Mi, M. B., and Wang, X
Cited in the paper.
Video generation models as world simulators
OpenAI · 2024
Closest in time.
Mooncake: Kimi’s kvcache-centric architecture for llm serving
Qin, R., Li, Z., He, W., Zhang, M., Wu, Y., Zheng, W., and Xu, X · 2024
Closest in time.
Ditfastattn: Attention compression for diffusion transformer models
Yuan, Z., Lu, P., Zhang, H., Ning, X., Zhang, L., Zhao, T., Yan, S., Dai, G., and Wang, Y · 2024
Closest in time.
Real-time video generation with pyramid attention broadcast
Zhao, X., Jin, X., Wang, K., and You, Y · 2024
Closest in time.