Fetching the paper…
Reading the bibliography…
In the evolving landscape of neural network models, one prominent challenge stand out: the significant memory overheads associated with training expansive models.
A turing machine time hierarchy
Žák, S · 1983
Earlier work this paper cites.
Disco: Running commodity operating systems on scalable multiprocessors
Bugnion, E., Devine, S., Govil, K., and Rosenblum, M · 1997
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
The flyweight pattern
Harmes, R. and Diaz, D · 2008
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., et al · 2012
Earlier work this paper cites.
Effective multi-gpu communication using multiple cuda streams and threads
Sourouri, M., Gillberg, T., Baden, S. B., and Cai, X · 2014
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in tensorflow
Sergeev, A. and Del Balso, M · 2018
Earlier work this paper cites.
Imagenet training in minutes
You, Y., Zhang, Z., Hsieh, C.-J., Demmel, J., and Keutzer, K · 2018
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Cited alongside, same era.
Pipedream: Generalized pipeline parallelism for dnn training
Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., Gibbons, P. B., and Zaharia, M · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Cited alongside, same era.
torchgpipe: On-the-fly pipeline parallelism for training giant models
Kim, C., Lee, H., Jeong, M., Baek, W., Yoon, B., Kim, I., Lim, S., and Kim, S · 2020
Cited alongside, same era.
Efficient large-scale language model training on gpu clusters using megatron-lm
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., et al · 2021
Later among the works it cites.
{ \{ ZeRO-Offload } \} : Democratizing { \{ Billion-Scale } \} model training
Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y · 2021
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Later among the works it cites.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A. P., Caron, M., Geirhos, R., Alabdulmohsin, I., et al · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Efficiently scaling transformer inference
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pytorch distributed: Experiences on accelerating data parallel training
Li, S., Zhao, Y., Varma, R., Salpekar, O., Noordhuis, P., Li, T., Paszke, A., Smith, J., Vaughan, B., Damania, P., et al · 2020
Cited alongside, same era.
Elastic machine learning algorithms in amazon sagemaker
Liberty, E., Karnin, Z., Xiang, B., Rouesnel, L., Coskun, B., Nallapati, R., Delgado, J., Sadoughi, A., Astashonok, Y., Das, P., et al · 2020
Cited alongside, same era.
Introduction to tensorflow 2.0
Singh, P., Manure, A., Singh, P., and Manure, A · 2020
Cited alongside, same era.
Pipetransformer: Automated elastic pipelining for distributed training of transformers
He, C., Li, S., Soltanolkotabi, M., and Avestimehr, S · 2021
Cited alongside, same era.
https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/usage/inplace.html
In-place Operations
Cited in the paper.
https://www.nvidia.com/en-us/data-center/a100/
NVIDIA A100 Tensor Core GPU
Cited in the paper.
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Pytorch fsdp: experiences on scaling fully sharded data parallel
Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., et al · 2023
Closest in time.