Fetching the paper…
Reading the bibliography…
Many deep learning applications benefit from using large models with billions of parameters.
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 1909
Earlier work this paper cites.
Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position
Fukushima, K · 1980
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Matrix multiplication via arithmetic progressions
Coppersmith, D. and Winograd, S · 1990
Earlier work this paper cites.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Weighted round-robin cell multiplexing in a general-purpose atm switch chip
Katevenis, M., Sidiropoulos, S., and Courcoubetis, C · 1991
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S. and Schmidhuber, J · 1996
Earlier work this paper cites.
Algorithm 799: revolve: an implementation of checkpointing for the reverse or adjoint mode of computational differentiation
Griewank, A. and Walther, A · 2000
Earlier work this paper cites.
Kademlia: A peer-to-peer information system based on the xor metric
Maymounkov, P. and Mazieres, D · 2002
Earlier work this paper cites.
Training large neural networks with constant memory using a new execution algorithm
Pudipeddi, B., Mesmakhosroshahi, M., Xi, J., and Bharadwaj, S · 2002
Earlier work this paper cites.
GLU variants improve transformer
Shazeer, N · 2002
Earlier work this paper cites.
Understanding the efficiency of gpu algorithms for matrix-matrix multiplication
Fatahalian, K., Sugerman, J., and Hanrahan, P · 2004
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Recht, B., Ré, C., Wright, S. J., and Niu, F · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Le, Q. V., Mao, M. Z., Ranzato, M., Senior, A. W., Tucker, P. A., Yang, K., and Ng, A. Y · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
How Mechanics Shaped the Modern World
Allen, D. H · 2013
Earlier work this paper cites.
Deep learning with COTS HPC systems
Coates, A., Huval, B., Wang, T., Wu, D. J., Catanzaro, B., and Ng, A. Y · 2013
Earlier work this paper cites.
Maxout networks
Goodfellow, I. J., Warde-Farley, D., Mirza, M., Courville, A. C., and Bengio, Y · 2013
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
Chilimbi, T., Suzue, Y., Apacible, J., and Kalyanaraman, K · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A · 2014
Earlier work this paper cites.
Matrix multiplication on high-density multi-gpu architectures: Theoretical and experimental investigations
Zhang, P. and Gao, Y · 2015
Earlier work this paper cites.
Ba, L. J., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
8-bit approximations for parallelism in deep learning
Dettmers, T · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
Proteus: Agile ml elasticity through tiered reliability in dynamic resource markets
Harlap, A., Tumanov, A., Chung, A., Ganger, G. R., and Gibbons, P. B · 2017
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? A case study for decentralized parallel stochastic gradient descent
Lian, X., Zhang, C., Zhang, H., Hsieh, C., Zhang, W., and Liu, J · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
A hybrid gpu cluster and volunteer computing platform for scalable deep learning
Kijsipongse, E., Piyatumrong, A., and U-ruekolan, S · 2018
Earlier work this paper cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Lin, Y., Han, S., Mao, H., Wang, Y., and Dally, B · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., Sepassi, R., and Hechtman, B. A · 2018
Cited alongside, same era.
Making asynchronous stochastic gradient descent work for transformers
Aji, A. F. and Heafield, K · 2019
Cited alongside, same era.
Adaptive input representations for neural language modeling
Baevski, A. and Auli, M · 2019
Cited alongside, same era.
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis
Ben-Nun, T. and Hoefler, T · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Openwebtext corpus, 2019
Gokaslan, A. and Cohen, V · 2019
Machine learning on volatile instances
Zhang, X., Wang, J., Joshi, G., and Joe-Wong, C · 2020
Later among the works it cites.
A refined laser method and faster matrix multiplication
Alman, J. and Williams, V. V · 2021
Later among the works it cites.
Distributed deep learning using volunteer computing-like paradigm
Atre, M., Jha, B., and Rao, A · 2021
Later among the works it cites.
Fairscale: A general purpose modular pytorch library for high performance and large scale training
Baines, M., Bhosale, S., Caggiano, V., Goyal, N., Goyal, S., Ott, M., Lefaudeux, B., Liptchinsky, V., Rabbat, M., Sheiffer, S., Sridhar, A., and Xu, M · 2021
Later among the works it cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, 2021
Black, S., Leo, G., Wang, P., Leahy, C., and Biderman, S · 2021
Later among the works it cites.
Coatnet: Marrying convolution and attention for all data sizes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M. X., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., and Chen, Z · 2019
Cited alongside, same era.
Beyond data and model parallelism for deep neural networks
Jia, Z., Zaharia, M., and Aiken, A · 2019
Cited alongside, same era.
Large memory layers with product keys
Lample, G., Sablayrolles, A., Ranzato, M., Denoyer, L., and Jégou, H · 2019
Cited alongside, same era.
Scaling the summit: Deploying the world’s fastest supercomputer
Larrea, V. G. V., Joubert, W., Brim, M. J., Budiardja, R. D., Maxwell, D., Ezell, M., Zimmer, C., Boehm, S., Elwasif, W. R., Oral, S., Fuson, C., Pelfrey, D., Hernandez, O. R., Leverman, D., Hanley, J., Berrill, M. A., and Tharrington, A. N · 2019
Cited alongside, same era.
Speeding up deep learning with transient servers
Li, S., Walls, R. J., Xu, L., and Guo, T · 2019
Cited alongside, same era.
Pipedream: Generalized pipeline parallelism for dnn training
Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., Gibbons, P. B., and Zaharia, M · 2019
Cited alongside, same era.
Dai, Z., Liu, H., Le, Q. V., and Tan, M · 2021
Later among the works it cites.
Diffusion models beat gans on image synthesis
Dhariwal, P. and Nichol, A. Q · 2021
Later among the works it cites.
Distributed deep learning in open collaborations
Diskin, M., Bukhtiyarov, A., Ryabinin, M., Saulnier, L., Lhoest, Q., Sinitsin, A., Popov, D., Pyrkin, D. V., Kashirin, M., Borzunov, A., del Moral, A. V., Mazur, D., Kobelev, I., Jernite, Y., Wolf, T., and Pekhimenko, G · 2021
Later among the works it cites.
Elastic Horovod
ElasticHorovod · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2021
Later among the works it cites.
Deberta: decoding-enhanced bert with disentangled attention
He, P., Liu, X., Gao, J., and Chen, W · 2021
Later among the works it cites.
Microsoft announces new supercomputer, lays out vision for future ai work
Langston, J · 2021
Later among the works it cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2021
Later among the works it cites.
Li, C., Zhang, M., and He, Y · 2021
Later among the works it cites.
Efficient large-scale language model training on gpu clusters
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., et al · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher, 2021
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P.-S., Glaese, A., Welbl, J., Dathathri, S., Huang, S., Uesato, J., Mellor, J., Higgins, I., Creswell, A., McAleese, N., Wu, A., Elsen, E., Jayakumar, S., Buchatskaya, E., Budden, D., Sutherland, E., Simonyan, K., Paganini, M., Sifre, L., Martens, L., Li, X. L., Kuncoro, A., Nematzadeh, A., Gribovskaya, E., Donato, D., Lazaridou, A., Mensch, A., Lespiau, J.-B., Tsimpoukelli, M., Grigorev, N., Fritz, D., Sottiaux, T., Pajarskas, M., Pohlen, T., Gong, Z., Toyama, D., de Masson d’Autume, C., Li, Y., Terzi, T., Mikulik, V., Babuschkin, I., Clark, A., de Las Casas, D., Guy, A., Jones, C., Bradbury, J., Johnson, M., Hechtman, B., Weidinger, L., Gabriel, I., Isaac, W., Lockhart, E., Osindero, S., Rimell, L., Dyer, C., Vinyals, O., Ayoub, K., Stanway, J., Bennett, L., Hassabis, D., Kavukcuoglu, K., and Irving, G · 2021
Later among the works it cites.
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
Rajbhandari, S., Ruwase, O., Rasley, J., Smith, S., and He, Y · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I · 2021
Later among the works it cites.
Zero-offload: Democratizing billion-scale model training, 2021
Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y · 2021
Later among the works it cites.
Moshpit SGD: communication-efficient decentralized training on heterogeneous unreliable devices
Ryabinin, M., Gorbunov, E., Plokhotnyuk, V., and Pekhimenko, G · 2021
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding, 2021
Su, J., Lu, Y., Pan, S., Wen, B., and Liu, Y · 2021
Later among the works it cites.
ERNIE 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation
Sun, Y., Wang, S., Feng, S., Ding, S., Pang, C., Shang, J., Liu, J., Chen, X., Zhao, Y., Lu, Y., Liu, W., Wu, Z., Gong, W., Liang, J., Shang, Z., Sun, P., Liu, W., Ouyang, X., Yu, D., Tian, H., Wu, H., and Wang, H · 2021
Later among the works it cites.
Piper: Multidimensional planner for DNN parallelization
Tarnawski, J., Narayanan, D., and Phanishayee, A · 2021
Later among the works it cites.
PyTorch Elastic
TorchElastic · 2021
Later among the works it cites.
Monthly ip latency data, 2021
Verizon · 2021
Later among the works it cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B. and Komatsuzaki, A · 2021
Later among the works it cites.
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2021
Later among the works it cites.
PaLM: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
Later among the works it cites.
8-bit optimizers via block-wise quantization
Dettmers, T., Lewis, M., Shleifer, S., and Zettlemoyer, L · 2022
Later among the works it cites.
A survey of transformers
Lin, T., Wang, Y., Liu, X., and Qiu, X · 2022
Later among the works it cites.
Bamboo: Making preemptible instances resilient for affordable training of large dnns, 2022
Thorpe, J., Zhao, P., Eyolfson, J., Qiao, Y., Jia, Z., Zhang, M., Netravali, R., and Xu, G. H · 2022
Later among the works it cites.
Fine-tuning language models over slow networks using activation quantization with guarantees
Wang, J., Yuan, B., Rimanic, L., He, Y., Dao, T., Chen, B., Re, C., and Zhang, C · 2022
Later among the works it cites.
Decentralized training of foundation models in heterogeneous environments
Yuan, B., He, Y., Davis, J. Q., Zhang, T., Dao, T., Chen, B., Liang, P., Re, C., and Zhang, C · 2022
Later among the works it cites.
Alpa: Automating inter- and intra-operator parallelism for distributed deep learning, 2022
Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Xing, E. P., Gonzalez, J. E., and Stoica, I · 2022
Later among the works it cites.