Fetching the paper…
Reading the bibliography…
Training LLMs relies on distributed implementations using multiple GPUs to compute gradients in parallel with sharded optimizers.
Distributed training strategies for the structured perceptron
Ryan McDonald, Keith Hall, and Gideon Mann · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex Smola · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc' aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, Quoc Le, and Andrew Ng · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep learning with elastic averaging sgd
Sixin Zhang, Anna Choromanska, and Yann LeCun · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost, 2016
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Earlier work this paper cites.
Federated optimization: Distributed machine learning for on-device intelligence
Jakub Konecný, H. B. McMahan, Daniel Ramage, and Peter Richtárik · 2016
Earlier work this paper cites.
Asynchrony begets momentum, with an application to deep learning
Ioannis Mitliagkas, Ce Zhang, Stefan Hadjis, and Christopher Ré · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernandez · 2016
Earlier work this paper cites.
Communication-Efficient Learning of Deep Networks from Decentralized Data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Asynchronous stochastic gradient descent with delay compensation
Shuxin Zheng, Qi Meng, Taifeng Wang, Wei Chen, Nenghai Yu, Zhi-Ming Ma, and Tie-Yan Liu · 2017
Earlier work this paper cites.
i-revnet: Deep invertible networks
Jörn-Henrik Jacobsen, Arnold W.M. Smeulders, and Edouard Oyallon · 2018
Earlier work this paper cites.
Pipe-sgd: A decentralized pipelined sgd framework for distributed deep net training
Youjie Li, Mingchao Yu, Songze Li, Salman Avestimehr, Nam Sung Kim, and Alexander Schwing · 2018
Earlier work this paper cites.
On the convergence properties of a k-step averaging stochastic gradient descent algorithm for nonconvex optimization
Fan Zhou and Guojing Cong · 2018
Earlier work this paper cites.
Stochastic gradient push for distributed deep learning
Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, and Mike Rabbat · 2019
Earlier work this paper cites.
Efficient and robust parallel dnn training through model parallelism on multi-gpu platform, 2019
Chi-Chung Chen, Chia-Lin Yang, and Hsiang-Yun Cheng · 2019
Earlier work this paper cites.
Openwebtext corpus
Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Earlier work this paper cites.
A study of bfloat16 for deep learning training, 2019
Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi, Dipankar Das, Kunal Banerjee, Sasikanth Avancha, Dharma Teja Vooturi, Nataraj Jammalamadaka, Jianyu Huang, Hector Yuen, Jiyan Yang, Jongsoo Park, Alexander Heinecke, Evangelos Georganas, Sudarshan Srinivasan, Abhisek Kundu, Misha Smelyanskiy, Bharat Kaul, and Pradeep Dubey · 2019
Earlier work this paper cites.
Federated optimization for heterogeneous networks
Tian Li, Anit Kumar Sahu, Manzil Zaheer, Maziar Sanjabi, Ameet Talwalkar, and Virginia Smith · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Pipedream: Generalized pipeline parallelism for dnn training
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R Devanur, Gregory R Ganger, Phillip B Gibbons, and Matei Zaharia · 2019
Earlier work this paper cites.
Pytorch: an imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Faster distributed deep net training: computation and communication decoupled stochastic gradient descent
Shuheng Shen, Linli Xu, Jingchang Liu, Xianfeng Liang, and Yifei Cheng · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Earlier work this paper cites.
Local SGD converges fast and communicates little
Sebastian U. Stich · 2019
Earlier work this paper cites.
Yellowfin and the art of momentum tuning
Jian Zhang and Ioannis Mitliagkas · 2019
Earlier work this paper cites.
A tight convergence analysis for stochastic gradient descent with delayed updates
Yossi Arjevani, Ohad Shamir, and Nathan Srebro · 2020
Cited alongside, same era.
Advances in asynchronous parallel and distributed optimization
By Mahmoud Assran, Arda Aytekin, Hamid Reza Feyzmahdavian, Mikael Johansson, and Michael G. Rabbat · 2020
Cited alongside, same era.
Toward communication efficient adaptive gradient method
Xiangyi Chen, Xiaoyun Li, and P. Li · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Cited alongside, same era.
Mime: Mimicking centralized stochastic algorithms in federated learning
Sai Praneeth Karimireddy, Martin Jaggi, Satyen Kale, Mehryar Mohri, Sashank J. Reddi, Sebastian U. Stich, and Ananda Theertha Suresh · 2020
Zero-offload: Democratizing billion-scale model training, 2021
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He · 2021
Later among the works it cites.
Pipemare: Asynchronous pipeline parallel dnn training
Bowen Yang, Jian Zhang, Jonathan Li, Christopher Re, Christopher Aberger, and Christopher De Sa · 2021
Later among the works it cites.
Fully decoupled neural network learning using delayed gradients
Huiping Zhuang, Yi Wang, Qinglai Liu, and Zhiping Lin · 2021
Later among the works it cites.
Sapipe: Staleness-aware pipeline for data parallel dnn training
Yangrui Chen, Cong Xie, Meng Ma, Juncheng Gu, Yanghua Peng, Haibin Lin, Chuan Wu, and Yibo Zhu · 2022
Later among the works it cites.
Reversible vision transformers
Karttikeya Mangalam, Haoqi Fan, Yanghao Li, Chao-Yuan Wu, Bo Xiong, Christoph Feichtenhofer, and Jitendra Malik · 2022
Later among the works it cites.
Gradskip: Communication-accelerated local gradient methods with better computational complexity, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
torchgpipe: On-the-fly pipeline parallelism for training giant models, 2020
Chiheon Kim, Heungsub Lee, Myungryong Jeong, Woonhyuk Baek, Boogeon Yoon, Ildoo Kim, Sungbin Lim, and Sungwoong Kim · 2020
Cited alongside, same era.
Pytorch distributed: experiences on accelerating data parallel training
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala · 2020
Cited alongside, same era.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U. Stich, Kumar Kshitij Patel, and Martin Jaggi · 2020
Cited alongside, same era.
Zero: memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models, 2020
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Cited alongside, same era.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He · 2020
Cited alongside, same era.
The error-feedback framework: better rates for sgd with delayed gradients and compressed updates
Sebastian U. Stich and Sai Praneeth Karimireddy · 2020
Cited alongside, same era.
Artavazd Maranjyan, Mher Safaryan, and Peter Richtárik · 2022
Later among the works it cites.
Asynchronous SGD beats minibatch SGD under arbitrary delays
Konstantin Mishchenko, Francis Bach, Mathieu Even, and Blake Woodworth · 2022
Later among the works it cites.
Proxskip: Yes! local gradient steps provably lead to communication acceleration! finally!
Konstantin Mishchenko, Grigory Malinovsky, Sebastian Stich, and Peter Richtárik · 2022
Later among the works it cites.
Federated learning with buffered asynchronous aggregation
John Nguyen, Kshitiz Malik, Hongyuan Zhan, Ashkan Yousefpour, Mike Rabbat, Mani Malek, and Dzmitry Huba · 2022
Later among the works it cites.
https://github.com/microsoft/deepspeed/discussions/2461, 2022
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He · 2022
Later among the works it cites.
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al · 2022
Later among the works it cites.
Delay-adaptive step-sizes for asynchronous learning
Xuyang Wu, Sindri Magnusson, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2022
Later among the works it cites.
Mics: Near-linear scaling for training gigantic model on public cloud, 2022
Zhen Zhang, Shuai Zheng, Yida Wang, Justin Chiu, George Karypis, Trishul Chilimbi, Mu Li, and Xin Jin · 2022
Later among the works it cites.
GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch, 9 2023
Alex Andonian, Quentin Anthony, Stella Biderman, Sid Black, Preetham Gali, Leo Gao, Eric Hallahan, Josh Levy-Kramer, Connor Leahy, Lucas Nestler, Kip Parker, Michael Pieler, Jason Phang, Shivanshu Purohit, Hailey Schoelkopf, Dashiell Stander, Tri Songz, Curt Tigges, Benjamin Thérien, Phil Wang, and Samuel Weinbach · 2023
Later among the works it cites.
Diloco: Distributed low-communication training of language models
Arthur Douillard, Qixuan Feng, Andrei A Rusu, Rachita Chhaparia, Yani Donchev, Adhiguna Kuncoro, Marc’Aurelio Ranzato, Arthur Szlam, and Jiajun Shen · 2023
Later among the works it cites.
Tinystories: How small can language models be and still speak coherent english?, 2023
Ronen Eldan and Yuanzhi Li · 2023
Later among the works it cites.
Asynchronous iterations in optimization: New sequence results and sharper algorithmic guarantees
Hamid Reza Feyzmahdavian and Mikael Johansson · 2023
Later among the works it cites.
Colossal-auto: Unified automation of parallelization and activation checkpoint for large-scale models, 2023
Yuliang Liu, Shenggui Li, Jiarui Fang, Yanjun Shao, Boyuan Yao, and Yang You · 2023
Later among the works it cites.
$\textbf{A}^2\textbf{CiD}^2$: Accelerating asynchronous communication in decentralized deep learning
Adel Nabli, Eugene Belilovsky, and Edouard Oyallon · 2023
Later among the works it cites.
DADAO: Decoupled accelerated decentralized asynchronous optimization
Adel Nabli and Edouard Oyallon · 2023
Later among the works it cites.
Computation vs. communication scaling for future transformers on future hardware, 2023
Suchita Pati, Shaizeen Aga, Mahzabeen Islam, Nuwan Jayasena, and Matthew D. Sinclair · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Later among the works it cites.
Zero++: Extremely efficient collective communication for giant model training, 2023
Guanhua Wang, Heyang Qin, Sam Ade Jacobs, Connor Holmes, Samyam Rajbhandari, Olatunji Ruwase, Feng Yan, Lei Yang, and Yuxiong He · 2023
Later among the works it cites.
A survey of large language models, 2023
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen · 2023
Later among the works it cites.
Pytorch fsdp: Experiences on scaling fully sharded data parallel, 2023
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, and Shen Li · 2023
Later among the works it cites.
Sharper convergence guarantees for asynchronous sgd for distributed and federated learning
Anastasia Koloskova, Sebastian U. Stich, and Martin Jaggi · 2024
Closest in time.
CO2: Efficient distributed training with full communication-computation overlap
Weigao Sun, Zhen Qin, Weixuan Sun, Shidi Li, Dong Li, Xuyang Shen, Yu Qiao, and Yiran Zhong · 2024
Closest in time.
Zachary Charles, Gabriel Teston, Lucio Dery, Keith Rush, Nova Fallen, Zachary Garrett, Arthur Szlam, and Arthur Douillard · 2025
Closest in time.
PETRA: Parallel end-to-end training with reversible architectures
Stephane Rivaud, Louis Fournier, Thomas Pumir, Eugene Belilovsky, Michael Eickenberg, and Edouard Oyallon · 2025
Closest in time.