Fetching the paper…
Reading the bibliography…
In this paper, we consider hybrid parallelism -- a paradigm that employs both Data Parallelism (DP) and Model Parallelism (MP) -- to scale distributed training of large recommendation models.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Dark silicon and the end of multicore scaling. In 2011 38th Annual international symposium on computer architecture (ISCA) . IEEE, 365–376
Hadi Esmaeilzadeh, Emily Blem, Renee St Amant, Karthikeyan Sankaralingam, and Doug Burger. 2011 · 2011
Earlier work this paper cites.
Deep sparse rectifier neural networks. In Proceedings of the fourteenth international conference on artificial intelligence and statistics . 315–323
Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011 · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Advances in neural information processing systems . 693–701
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu. 2011 · 2011
Earlier work this paper cites.
Wit3: Web inventory of transcribed and translated talks. In Conference of european association for machine translation . 261–268
Mauro Cettolo, Christian Girardi, and Marcello Federico. 2012 · 2012
Earlier work this paper cites.
Large scale distributed deep networks. In Advances in neural information processing systems . 1223–1231
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Approximate Computation and Implicit Regularization for Very Large-scale Data Analysis. In Proceedings of the 31st ACM Symposium on Principles of Database Systems . 143–154
M. W. Mahoney. 2012 · 2012
Earlier work this paper cites.
Adaptive dropout for training deep neural networks. In Advances in neural information processing systems . 3084–3092
Jimmy Ba and Brendan Frey. 2013 · 2013
Earlier work this paper cites.
Alireza Makhzani and Brendan Frey. 2013 · 2013
Earlier work this paper cites.
On model parallelization and scheduling strategies for distributed machine learning. In Advances in neural inf. processing systems . 2834–2842
Seunghak Lee, Jin Kyu Kim, Xun Zheng, Qirong Ho, Garth A Gibson, and Eric P Xing. 2014 · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server. In 11th { \{ USENIX } \} Symposium on Operating Systems Design and Implementation ( { \{ OSDI } \} 14) . 583–598
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su. 2014 · 2014
Earlier work this paper cites.
Sketching As a Tool for Numerical Linear Algebra
David P. Woodruff. 2014 · 2014
Earlier work this paper cites.
Winner-take-all autoencoders. In Advances in neural information processing systems . 2791–2799
Alireza Makhzani and Brendan J Frey. 2015 · 2015
Earlier work this paper cites.
Revisiting distributed synchronous SGD
Jianmin Chen, Xinghao Pan, Rajat Monga, Samy Bengio, and Rafal Jozefowicz. 2016 · 2016
Earlier work this paper cites.
Deep learning . Vol. 1
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016 · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2016 · 2016
Earlier work this paper cites.
STRADS: A Distributed Framework for Scheduled Model Parallel Machine Learning (EuroSys ’16) . Association for Computing Machinery, New York, NY, USA, Article 5, 16 pages
Jin Kyu Kim, Qirong Ho, Seunghak Lee, Xun Zheng, Wei Dai, Garth A. Gibson, and Eric P. Xing. 2016 · 2016
Earlier work this paper cites.
Parallel Local Graph Clustering
J. Shun, F. Roosta-Khorasani, K. Fountoulakis, and M. W. Mahoney. 2016 · 2016
Earlier work this paper cites.
Sparse Communication for Distributed Gradient Descent. In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing . 440–445
Alham Fikri Aji and Kenneth Heafield. 2017 · 2017
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding. In Advances in Neural Information Processing Systems . 1709–1720
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic. 2017 · 2017
Earlier work this paper cites.
MCM-GPU: Multi-chip-module GPUs for continued performance scalability
Akhil Arunkumar, Evgeny Bolotin, Benjamin Cho, Ugljesa Milic, Eiman Ebrahimi, Oreste Villa, Aamer Jaleel, Carole-Jean Wu, and David Nellans. 2017 · 2017
Earlier work this paper cites.
AdaBatch: Adaptive Batch Sizes for Training Deep Neural Networks
Aditya Devarakonda, Maxim Naumov, and Michael Garland. 2017 · 2017
Earlier work this paper cites.
An optimization approach to locally-biased graph algorithms
K. Fountoulakis, D. F. Gleich, and M. W. Mahoney. 2017 · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks. In NIPS
Elad Hoffer, Itay Hubara, and Daniel Soudry. 2017 · 2017
Cited alongside, same era.
Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4700–4708
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. 2017 · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Cited alongside, same era.
Atomo: Communication-efficient learning via atomic sparsification. In Advances in Neural Information Processing Systems . 9850–9861
Hongyi Wang, Scott Sievert, Shengchao Liu, Zachary Charles, Dimitris Papailiopoulos, and Stephen Wright. 2018 · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization. In Advances in Neural Information Processing Systems . 1299–1309
Jianqiao Wangni, Jialei Wang, Ji Liu, and Tong Zhang. 2018 · 2018
Later among the works it cites.
Hessian-based analysis of large batch training and robustness to adversaries
Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer, and Michael W Mahoney. 2018 · 2018
Later among the works it cites.
Gradiveq: Vector quantization for bandwidth-efficient gradient aggregation in distributed CNN training. In Advances in Neural Information Processing Systems . 5123–5133
Mingchao Yu, Zhifeng Lin, Krishna Narra, Songze Li, Youjie Li, Nam Sung Kim, Alexander Schwing, Murali Annavaram, and Salman Avestimehr. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le. 2017 · 2017
Cited alongside, same era.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning. In Advances in neural information processing systems . 1509–1519
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2017 · 2017
Cited alongside, same era.
Scaling SGD batch size to 32k for ImageNet training
Yang You, Igor Gitman, and Boris Ginsburg. 2017 · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods. In Advances in Neural Information Processing Systems . 5973–5983
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cédric Renggli. 2018 · 2018
Cited alongside, same era.
Efficient and robust parallel dnn training through model parallelism on multi-gpu platform
Chi-Chung Chen, Chia-Lin Yang, and Hsiang-Yun Cheng. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Integrated model, batch, and domain parallelism in training neural networks. In Proceedings of the 30th on Symposium on Parallelism in Algorithms and Architectures . 77–86
Amir Gholami, Ariful Azad, Peter Jin, Kurt Keutzer, and Aydin Buluc. 2018 · 2018
Cited alongside, same era.
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis
Tal Ben-Nun and Torsten Hoefler. 2019 · 2019
Later among the works it cites.
Beidi Chen, Tharun Medini, James Farwell, Sameh Gobriel, Charlie Tai, and Anshumali Shrivastava. 2019 · 2019
Later among the works it cites.
GradZip: Gradient Compression using Alternating Matrix Factorization for Large-scale Deep Learning
Minsik Cho, Vinod Muthusamy, Brad Nemanich, and Ruchir Puri. 2019 · 2019
Later among the works it cites.
XPipe: Efficient Pipeline Model Parallelism for Multi-GPU DNN Training
Lei Guan, Wotao Yin, Dongsheng Li, and Xicheng Lu. 2019 · 2019
Later among the works it cites.
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, and zhifeng Chen. 2019 · 2019
Later among the works it cites.
Communication-efficient distributed sgd with sketching. In Advances in Neural Information Processing Systems . 13144–13154
Nikita Ivkin, Daniel Rothchild, Enayat Ullah, Vladimir Braverman, Ion Stoica, and Raman Arora. 2019 · 2019
Later among the works it cites.
Beyond data and model parallelism for deep neural networks
Zhihao Jia, Matei Zaharia, and Alex Aiken. 2019 · 2019
Later among the works it cites.
Error Feedback Fixes SignSGD and other Gradient Compression Schemes. In International Conference on Machine Learning . 3252–3261
Sai Praneeth Karimireddy, Quentin Rebjock, Sebastian Stich, and Martin Jaggi. 2019 · 2019
Later among the works it cites.
PipeDream: generalized pipeline parallelism for DNN training. In Proceedings of the 27th ACM Symposium on Operating Systems Principles . 1–15
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R Devanur, Gregory R Ganger, Phillip B Gibbons, and Matei Zaharia. 2019 · 2019
Later among the works it cites.
Deep learning recommendation model for personalization and recommendation systems
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G Azzolini, et al · 2019
Later among the works it cites.
fairseq: A Fast, Extensible Toolkit for Sequence Modeling. In Proceedings of NAACL-HLT 2019: Demonstrations
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Sparse gradient compression for distributed SGD. In International Conference on Database Systems for Advanced Applications . Springer, 139–155
Haobo Sun, Yingxia Shao, Jiawei Jiang, Bin Cui, Kai Lei, Yu Xu, and Jiang Wang. 2019 · 2019
Later among the works it cites.
PowerSGD: Practical low-rank gradient compression for distributed optimization. In Advances in Neural Information Processing Systems . 14259–14268
Thijs Vogels, Sai Praneeth Karimireddy, and Martin Jaggi. 2019 · 2019
Later among the works it cites.
PipeMare: Asynchronous Pipeline Parallel DNN Training
Bowen Yang, Jian Zhang, Jonathan Li, Christopher Ré, Christopher R Aberger, and Christopher De Sa. 2019 · 2019
Later among the works it cites.
Stochastic Weight Averaging in Parallel: Large-Batch Training That Generalizes Well. In International Conference on Learning Representations
Vipul Gupta, Santiago Akle Serrano, and Dennis DeCoste. 2020 · 2020
Closest in time.
Don’t Use Large Mini-batches, Use Local SGD. In International Conference on Learning Representations
Tao Lin, Sebastian U. Stich, Kumar Kshitij Patel, and Martin Jaggi. 2020 · 2020
Closest in time.
Jay H Park, Gyeongchan Yun, Chang M Yi, Nguyen T Nguyen, Seungmin Lee, Jaesik Choi, Sam H Noh, and Young-ri Choi. 2020 · 2020
Closest in time.
Distributed Equivalent Substitution Training for Large-Scale Recommender Systems. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval . 911–920
Haidong Rong, Yangzihao Wang, Feihu Zhou, Junjie Zhai, Haiyang Wu, Rui Lan, Fan Li, Han Zhang, Yuekui Yang, Zhenyu Guo, et al · 2020
Closest in time.
Compressed Communication for Distributed Deep Learning: Survey and Quantitative Evaluation
Hang Xu, Chen-Yu Ho, Ahmed M Abdelmoniem, Aritra Dutta, El Houcine Bergou, Konstantinos Karatsenidis, Marco Canini, and Panos Kalnis. 2020 · 2020
Closest in time.