Fetching the paper…
Reading the bibliography…
Recent years have witnessed a growing list of systems for distributed data-parallel training.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Optimization of collective reduction operations
Rolf Rabenseifner · 2004
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Feng Niu, Benjamin Recht, Christopher Ré, and Stephen J Wright · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg S Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V Le, Mark Z Mao, Marc’Aurelio Ranzato, Andrew Senior, Paul Tucker, et al · 2012
Earlier work this paper cites.
Mariana: Tencent deep learning platform and its applications
Yongqiang Zou, Xing Jin, Yi Li, Zhimao Guo, Eryu Wang, and Bin Xiao · 2014
Earlier work this paper cites.
Hybrid parallelization strategies for large-scale machine learning in systemml
Matthias Boehm, Shirish Tatikonda, Berthold Reinwald, Prithviraj Sen, Yuanyuan Tian, Douglas R Burdick, and Shivakumar Vaithyanathan · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Malt: distributed data-parallelism for existing ml applications
Hao Li, Asim Kadav, Erik Kruus, and Cristian Ungureanu · 2015
Earlier work this paper cites.
High-performance distributed ml at scale through parameter server consistency models
Wei Dai, Abhimanu Kumar, Jinliang Wei, Qirong Ho, Garth Gibson, and Eric P Xing · 2015
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2016
Earlier work this paper cites.
Systemml: Declarative machine learning on spark
Matthias Boehm, Michael W Dusenberry, Deron Eriksson, Alexandre V Evfimievski, Faraz Makari Manshadi, Niketan Pansare, Berthold Reinwald, Frederick R Reiss, Prithviraj Sen, Arvind C Surve, et al · 2016
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
Tianqi Chen and Carlos Guestrin · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Zipml: Training linear models with end-to-end low precision, and a little bit of deep learning
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Earlier work this paper cites.
Terngrad: ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Earlier work this paper cites.
Poseidon: An efficient communication architecture for distributed deep learning on { \{ GPU } \} clusters
Hao Zhang, Zeyu Zheng, Shizhen Xu, Wei Dai, Qirong Ho, Xiaodan Liang, Zhiting Hu, Jinliang Wei, Pengtao Xie, and Eric P Xing · 2017
Earlier work this paper cites.
Efficient distributed learning with sparsity
Jialei Wang, Mladen Kolar, Nathan Srebro, and Tong Zhang · 2017
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Earlier work this paper cites.
Asynchronous stochastic gradient descent with delay compensation
Shuxin Zheng, Qi Meng, Taifeng Wang, Wei Chen, Nenghai Yu, Zhi-Ming Ma, and Tie-Yan Liu · 2017
Earlier work this paper cites.
Asynchronous distributed variational gaussian process for regression
Hao Peng, Shandian Zhe, Xiao Zhang, and Yuan Qi · 2017
Earlier work this paper cites.
Understanding and optimizing asynchronous low-precision stochastic gradient descent
Christopher De Sa, Matthew Feldman, Christopher Ré, and Kunle Olukotun · 2017
Earlier work this paper cites.
Heterogeneity-aware distributed parameter servers
Jiawei Jiang, Bin Cui, Ce Zhang, and Lele Yu · 2017
Earlier work this paper cites.
A cost-based optimizer for gradient descent optimization
Zoi Kaoudi, Jorge-Arnulfo Quiané-Ruiz, Saravanan Thirumuruganathan, Sanjay Chawla, and Divy Agrawal · 2017
Earlier work this paper cites.
signsgd: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Earlier work this paper cites.
Gradient sparsification for communication-efficient distributed optimization
J Wangni, J Liu, J Wang, and T Zhang · 2018
Cited alongside, same era.
Horovod: fast and easy distributed deep learning in tensorflow
Alexander Sergeev and Mike Del Balso · 2018
Cited alongside, same era.
Angel: a new large-scale machine learning system
Jie Jiang, Lele Yu, Jiawei Jiang, Yuhong Liu, and Bin Cui · 2018
Cited alongside, same era.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Sarit Khirirat, Nikola Konstantinov, and Cédric Renggli · 2018
Cited alongside, same era.
Atomo: communication-efficient learning via atomic sparsification
Hongyi Wang, Scott Sievert, Zachary Charles, Shengchao Liu, Stephen Wright, and Dimitris Papailiopoulos · 2018
Cited alongside, same era.
Beyond data and model parallelism for deep neural networks
Zhihao Jia, Matei Zaharia, and Alex Aiken · 2019
Later among the works it cites.
Supporting very large models using automatic dataflow graph partitioning
Minjie Wang, Chien-chin Huang, and Jinyang Li · 2019
Later among the works it cites.
Declarative recursive computation on an rdbms: or, why you should use a database for distributed machine learning
Dimitrije Jankov, Shangyu Luo, Binhang Yuan, Zhuhua Cai, Jia Zou, Chris Jermaine, and Zekai J Gao · 2019
Later among the works it cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Later among the works it cites.
Pipedream: generalized pipeline parallelism for dnn training
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R Devanur, Gregory R Ganger, Phillip B Gibbons, and Matei Zaharia · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Youjie Li, Mingchao Yu, Songze Li, Salman Avestimehr, Nam Sung Kim, and Alexander Schwing · 2018
Cited alongside, same era.
Asynchronous decentralized parallel stochastic gradient descent
Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu · 2018
Cited alongside, same era.
Communication compression for decentralized training
Hanlin Tang, Shaoduo Gan, Ce Zhang, Tong Zhang, and Ji Liu · 2018
Cited alongside, same era.
D2: Decentralized training over decentralized data
Hanlin Tang, Xiangru Lian, Ming Yan, Ce Zhang, and Ji Liu · 2018
Cited alongside, same era.
Local sgd converges fast and communicates little
Sebastian U Stich · 2018
Cited alongside, same era.
Distributed asynchronous optimization with unbounded delays: How slow can you go?
Zhengyuan Zhou, Panayotis Mertikopoulos, Nicholas Bambos, Peter Glynn, Yinyu Ye, Li-Jia Li, and Li Fei-Fei · 2018
Cited alongside, same era.
Asynchronous stochastic quasi-newton mcmc for non-convex optimization
Umut Simsekli, Cagatay Yildiz, Than Huy Nguyen, Taylan Cemgil, and Gael Richard · 2018
Cited alongside, same era.
In-database distributed machine learning: demonstration using teradata sql engine
Sandeep Singh Sandha, Wellington Cabrera, Mohammed Al-Kateb, Sanjay Nair, and Mani Srivastava · 2019
Later among the works it cites.
Ps2: Parameter server on spark
Zhipeng Zhang, Bin Cui, Yingxia Shao, Lele Yu, Jiawei Jiang, and Xupeng Miao · 2019
Later among the works it cites.
Mllib*: Fast training of glms using spark mllib
Zhipeng Zhang, Jiawei Jiang, Wentao Wu, Ce Zhang, Lele Yu, and Bin Cui · 2019
Later among the works it cites.
Communication-efficient distributed sgd with sketching
Nikita Ivkin, Daniel Rothchild, Enayat Ullah, Ion Stoica, Raman Arora, et al · 2019
Later among the works it cites.
Qsparse-local-sgd: Distributed sgd with quantization, sparsification, and local computations
Debraj Basu, Deepesh Data, Can Karakus, and Suhas Diggavi · 2019
Later among the works it cites.
Deepsqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression
Hanlin Tang, Xiangru Lian, Shuang Qiu, Lei Yuan, Ce Zhang, Tong Zhang, and Ji Liu · 2019
Later among the works it cites.
Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning
Hao Yu, Sen Yang, and Shenghuo Zhu · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Scalable deep learning on distributed infrastructures: Challenges, techniques, and tools
Ruben Mayer and Hans-Arno Jacobsen · 2020
Later among the works it cites.
Memory-efficient pipeline-parallel dnn training
Deepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen, and Matei Zaharia · 2020
Later among the works it cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Later among the works it cites.
Vertica-ml: Distributed machine learning in vertica database
Arash Fard, Anh Le, George Larionov, Waqas Dhillon, and Chuck Bear · 2020
Later among the works it cites.
Db4ml-an in-memory database kernel with machine learning support
Matthias Jasny, Tobias Ziegler, Tim Kraska, Uwe Roehm, and Carsten Binnig · 2020
Later among the works it cites.
Distributed learning systems with first-order methods
Ji Liu, Ce Zhang, et al · 2020
Later among the works it cites.
On biased compression for distributed learning
Aleksandr Beznosikov, Samuel Horváth, Peter Richtárik, and Mher Safaryan · 2020
Later among the works it cites.
Tensor relational algebra for distributed machine learning system design
Binhang Yuan, Dimitrije Jankov, Jia Zou, Yuxin Tang, Daniel Bourgeois, and Chris Jermaine · 2021
Closest in time.
Terapipe: Token-level pipeline parallelism for training large-scale language models
Zhuohan Li, Siyuan Zhuang, Shiyuan Guo, Danyang Zhuo, Hao Zhang, Dawn Song, and Ion Stoica · 2021
Closest in time.
Pipetransformer: Automated elastic pipelining for distributed training of transformers
Chaoyang He, Shen Li, Mahdi Soltanolkotabi, and Salman Avestimehr · 2021
Closest in time.
1-bit adam: Communication efficient large-scale training with adam’s convergence speed
Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari, Conglong Li, Xiangru Lian, Ji Liu, Ce Zhang, and Yuxiong He · 2021
Closest in time.
Distributed learning systems with first-order methods
Ji Liu and Ce Zhang · 2021
Closest in time.