Fetching the paper…
Reading the bibliography…
The fundamental success of large language models hinges upon the efficacious implementation of large-scale distributed training techniques.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Gap aware mitigation of gradient staleness
Saar Barkai, Ido Hakimi, and Assaf Schuster · 1909
Earlier work this paper cites.
Slowmo: Improving communication-efficient distributed sgd with slow momentum
Jianyu Wang, Vinayak Tantia, Nicolas Ballas, and Michael Rabbat · 1910
Earlier work this paper cites.
Understanding top-k sparsification in distributed deep learning
Shaohuai Shi, Xiaowen Chu, Ka Chun Cheung, and Simon See · 1911
Earlier work this paper cites.
rtop-k: A statistical estimation approach to distributed sgd
Leighton Pate Barnes, Huseyin A. Inan, Berivan Isik, and Ayfer Ozgur · 2005
Earlier work this paper cites.
Pytorch distributed: Experiences on accelerating data parallel training
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala · 2006
Earlier work this paper cites.
More effective distributed ml via a stale synchronous parallel parameter server
Qirong Ho, James Cipar, Henggang Cui, Seunghak Lee, Jin Kyu Kim, Phillip B Gibbons, Garth A Gibson, Greg Ganger, and Eric P Xing · 2013
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su · 2014
Earlier work this paper cites.
ShapeNet: An Information-Rich 3D Model Repository
Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu · 2015
Earlier work this paper cites.
3d shapenets: A deep representation for volumetric shapes
Zhirong Wu, Shuran Song, Aditya Khosla, Fisher Yu, Linguang Zhang, Xiaoou Tang, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Staleness-aware async-sgd for distributed deep learning
Wei Zhang, Suyog Gupta, Xiangru Lian, and Ji Liu · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2016
Earlier work this paper cites.
Revisiting distributed synchronous sgd
Jianmin Chen, Xinghao Pan, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie · 2016
Earlier work this paper cites.
Parallel sgd: When does averaging help?
Jian Zhang, Christopher De Sa, Ioannis Mitliagkas, and Christopher Ré · 2016
Cited alongside, same era.
A descent lemma beyond lipschitz gradient continuity: first-order methods revisited and applications
Heinz H Bauschke, Jérôme Bolte, and Marc Teboulle · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U. Stich, Kumar Kshitij Patel, and Martin Jaggi · 2018
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Slow and stale gradients can win the race
Sanghamitra Dutta, Jianyu Wang, and Gauri Joshi · 2021
Later among the works it cites.
1-bit lamb: Communication efficient large-scale large-batch training with lamb’s convergence speed
Conglong Li, Ammar Ahmad Awan, Hanlin Tang, Samyam Rajbhandari, and Yuxiong He · 2021
Later among the works it cites.
1-bit adam: Communication efficient large-scale training with adam’s convergence speed
Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari, Conglong Li, Xiangru Lian, Ji Liu, Ce Zhang, and Yuxiong He · 2021
Later among the works it cites.
Cooperative sgd: A unified framework for the design and analysis of local-update sgd algorithms
Jianyu Wang and Gauri Joshi · 2021
Later among the works it cites.
Sharper convergence guarantees for asynchronous sgd for distributed and federated learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Local sgd converges fast and communicates little
Sebastian U Stich · 2018
Cited alongside, same era.
Stochastic gradient push for distributed deep learning
Mahmoud Assran, Nicolas Loizou, Nicolas Ballas, and Mike Rabbat · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Cited alongside, same era.
Communication-efficient distributed deep learning: A comprehensive survey
Zhenheng Tang, Shaohuai Shi, Xiaowen Chu, Wei Wang, and Bo Li · 2020
Cited alongside, same era.
Overlap local-SGD: An algorithmic approach to hide communication delays in distributed SGD
Jianyu Wang, Hao Liang, and Gauri Joshi · 2020
Cited alongside, same era.
Anastasia Koloskova, Sebastian U. Stich, and Martin Jaggi · 2022
Later among the works it cites.
Near-optimal sparse allreduce for distributed deep learning
Shigang Li and Torsten Hoefler · 2022
Later among the works it cites.
Maximizing communication efficiency for large-scale training via 0/1 adam
Yucheng Lu, Conglong Li, Minjia Zhang, Christopher De Sa, and Yuxiong He · 2022
Later among the works it cites.
Asynchronous sgd beats minibatch sgd under arbitrary delays
Konstantin Mishchenko, Francis Bach, Mathieu Even, and Blake Woodworth · 2022
Later among the works it cites.
Masked autoencoders for point cloud self-supervised learning
Yatian Pang, Wenxiao Wang, Francis EH Tay, Wei Liu, Yonghong Tian, and Li Yuan · 2022
Later among the works it cites.
The devil in linear transformer
Zhen Qin, Xiaodong Han, Weixuan Sun, Dongxu Li, Lingpeng Kong, Nick Barnes, and Yiran Zhong · 2022
Later among the works it cites.
Wenbo Su, Yuanxing Zhang, Yufeng Cai, Kaixu Ren, Pengjie Wang, Huimin Yi, Yue Song, Jing Chen, Hongbo Deng, Jian Xu, Lin Qu, and Bo zheng · 2022
Later among the works it cites.
Alexander Wettig, Tianyu Gao, Zexuan Zhong, and Danqi Chen · 2022
Later among the works it cites.
Sampled transformer for point sets
Shidi Li, Christian Walder, Alexander Soen, Lexing Xie, and Miaomiao Liu · 2023
Later among the works it cites.
Vicinity vision transformer
Weixuan Sun, Zhen Qin, Hui Deng, Jianyuan Wang, Yi Zhang, Kaihao Zhang, Nick Barnes, Stan Birchfield, Lingpeng Kong, and Yiran Zhong · 2023
Later among the works it cites.
Ms-net: A multi-path sparse model for motion prediction in multi-scenes
Xiaqiang Tang, Weigao Sun, Siyuan Hu, Yiyang Sun, and Yafeng Guo · 2023
Later among the works it cites.