Fetching the paper…
Reading the bibliography…
Scaling multi-dimensional transformers to long sequences is indispensable across various domains.
Data parallel algorithms
Hillis, W. D. and Steele Jr, G. L · 1986
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N. M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al · 2018
Earlier work this paper cites.
Axial attention in multidimensional transformers
Ho, J., Kalchbrenner, N., Weissenborn, D., and Salimans, T · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Earlier work this paper cites.
Pipedream: generalized pipeline parallelism for dnn training
Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., Gibbons, P. B., and Zaharia, M · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
Zero: Memory optimization towards training a trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Shazeer, N. M · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
Pytorch distributed: Experiences on accelerating data parallel training
Li, S., Zhao, Y., Varma, R., Salpekar, O., Noordhuis, P., Li, T., Paszke, A., Smith, J., Vaughan, B., Damania, P., et al · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Spatial-temporal transformer networks for traffic flow forecasting
Xu, M., Dai, W., Liu, C., Gao, X., Lin, W., Qi, G.-J., and Xiong, H · 2020
Earlier work this paper cites.
Spatial-temporal transformer for dynamic scene graph generation
Cong, Y., Liao, W., Ackermann, H., Yang, M. Y., and Rosenhahn, B · 2021
Cited alongside, same era.
End-to-end video object detection with spatial-temporal transformers
He, L., Zhou, Q., Li, X., Niu, L., Cheng, G., Li, X., Liu, W., Tong, Y., Ma, L., and Zhang, L · 2021
Cited alongside, same era.
Highly accurate protein structure prediction with alphafold
Jumper, J. M., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Zídek, A., Potapenko, A., Bridgland, A., Meyer, C., Kohl, S. A. A., Ballard, A., Cowie, A., Romera-Paredes, B., Nikolov, S., Jain, R., Adler, J., Back, T., Petersen, S., Reiman, D. A., Clancy, E., Zielinski, M., Steinegger, M., Pacholska, M., Berghammer, T., Bodenstein, S., Silver, D., Vinyals, O., Senior, A. W., Kavukcuoglu, K., Kohli, P., and Hassabis, D · 2021
Cited alongside, same era.
Chimera: Efficiently training large-scale neural networks with bidirectional pipelines
Li, S. and Hoefler, T · 2021
Cited alongside, same era.
Sequence parallelism: Long sequence training from system perspective
Li, S., Xue, F., Li, Y., and You, Y · 2021
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. S. and Xie, S · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., Parikh, D., Gupta, S., and Taigman, Y · 2022
Later among the works it cites.
Tubedetr: Spatio-temporal video grounding with transformers
Yang, A., Miech, A., Sivic, J., Laptev, I., and Schmid, C · 2022
Later among the works it cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebr’on, F., and Sanghai, S. K · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
Rajbhandari, S., Ruwase, O., Rasley, J., Smith, S., and He, Y · 2021
Cited alongside, same era.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Cited alongside, same era.
Learning spatio-temporal transformer for visual tracking
Yan, B., Peng, H., Fu, J., Wang, D., and Lu, H · 2021
Cited alongside, same era.
3d human pose estimation with spatial and temporal transformers
Zheng, C., Zhu, S., Mendieta, M., Yang, T., Chen, C., and Ding, Z · 2021
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and R’e, C · 2022
Cited alongside, same era.
Rstt: Real-time spatial temporal transformer for space-time video super-resolution
Geng, Z., Liang, L., Ding, T., and Zharkov, I · 2022
Cited alongside, same era.
Accelerate: Training and inference at scale made simple, efficient and adaptable., 2022
Gugger, S., Debut, L., Wolf, T., Schmid, P., Mueller, Z., Mangrulkar, S., Sun, M., and Bossan, B · 2022
Cited alongside, same era.
Stable video diffusion: Scaling latent video diffusion models to large datasets
Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., and Lorenz, D · 2023
Later among the works it cites.
Tempee: Temporal-spatial parallel transformer for radar echo extrapolation beyond auto-regression
Chen, S., Shu, T., Zhao, H., Zhong, G., and Chen, X · 2023
Later among the works it cites.
Sttre: A spatio-temporal transformer with relative embeddings for multivariate time series forecasting
Deihim, A., Alonso, E., and Apostolopoulou, D · 2023
Later among the works it cites.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2023
Later among the works it cites.
Jacobs, S. A., Tanaka, M., Zhang, C., Zhang, M., Song, L., Rajbhandari, S., and He, Y · 2023
Later among the works it cites.
Rwkv: Reinventing rnns for the transformer era
Peng, B., Alcaide, E., Anthony, Q. G., Albalak, A., Arcadinho, S., Biderman, S., Cao, H., Cheng, X., Chung, M., Grella, M., Kranthikiran, G., He, X., Hou, H., Kazienko, P., Kocoń, J., Kong, J., Koptyra, B., Lau, H., Mantri, K. S. I., Mom, F., Saito, A., Tang, X., Wang, B., Wind, J. S., Wozniak, S., Zhang, R., Zhang, Z., Zhao, Q., Zhou, P., Zhu, J., and Zhu, R · 2023
Later among the works it cites.
Pytorch fsdp: Experiences on scaling fully sharded data parallel
Zhao, Y., Gu, A., Varma, R., Luo, L., chin Huang, C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., Desmaison, A., Balioglu, C., Nguyen, B., Chauhan, G., Hao, Y., and Li, S · 2023
Later among the works it cites.
Fastfold: Optimizing alphafold training and inference on gpu clusters
Cheng, S., Zhao, X., Lu, G., Fang, J., Zheng, T., Wu, R., Zhang, X., Peng, J., and You, Y · 2024
Closest in time.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2024
Closest in time.
Latte: Latent diffusion transformer for video generation
Ma, X., Wang, Y., Jia, G., Chen, X., Liu, Z., Li, Y.-F., Chen, C., and Qiao, Y · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., baptiste Alayrac, J., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., Antonoglou, I., Anil, R., Borgeaud, S., Dai, A., Millican, K., Dyer, E., Glaese, M., Sottiaux, T., Lee, B., Viola, F., Reynolds, M., Xu, Y., Molloy, J., Chen, J., Isard, M., Barham, P., Hennigan, T., and et al · 2024
Closest in time.
Open-sora: Democratizing efficient video production for all, April 2024
Zangwei Zheng, X. P · 2024
Closest in time.