Fetching the paper…
Reading the bibliography…
Batch-splitting (data-parallelism) is the dominant distributed Deep Neural Network (DNN) training strategy, due to its universal applicability and its amenability to Single-Program-Multiple-Data (SPMD) programming.
“Communication efficient matrix multiplication on hypercubes”
Jarle Berntsen · 1989
Earlier work this paper cites.
“More Iteration Space Tiling”
M. Wolfe · 1989
Earlier work this paper cites.
“Communication Complexity of PRAMs”
Alok Aggarwal, Ashok. Chandra and Marc Snir · 1990
Earlier work this paper cites.
“Communication lower bounds for distributed-memory matrix multiplication”
Dror Irony, Sivan Toledo and Alexander Tiskin · 2004
Earlier work this paper cites.
“Bandwidth optimal all-reduce algorithms for clusters of workstations”
Pitch Patarasuk and Xin Yuan · 2009
Earlier work this paper cites.
“Y.: Optimal bucket algorithms for large mpi collectives on torus interconnects”
Nikhil Jain and Yogish Sabharwal · 2010
Earlier work this paper cites.
“Minimizing communication in numerical linear algebra”
Grey Ballard, James Demmel, Olga Holtz and Oded Schwartz · 2011
Earlier work this paper cites.
“Communication-Optimal Parallel 2.5D Matrix Multiplication and LU Factorization Algorithms”
Edgar Solomonik and James Demmel · 2011
Earlier work this paper cites.
“Large Scale Distributed Deep Networks”
Jeffrey Dean et al · 2012
Earlier work this paper cites.
“Imagenet classification with deep convolutional neural networks”
Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton · 2012
Cited alongside, same era.
“Communication optimal parallel multiplication of sparse random matrices”
Grey Ballard et al · 2013
Cited alongside, same era.
“Communication Lower Bounds and Optimal Algorithms for Programs That Reference Arrays - Part 1”, 2013
Michael Christ et al · 2013
Cited alongside, same era.
“Communication-optimal parallel recursive rectangular matrix multiplication”
James Demmel et al · 2013
Cited alongside, same era.
“A massively parallel tensor contraction framework for coupled-cluster computations”
Edgar Solomonik et al · 2014
Cited alongside, same era.
“Scalable Task-based Algorithm for Multiplication of Block-rank-sparse Matrices”
Justus. Calvin, Cannada. Lewis and Edward. Valeev · 2015
“Integrated Model and Data Parallelism in Training Neural Networks”
Amir Gholami, Ariful Azad, Kurt Keutzer and Aydin Buluc · 2017
Later among the works it cites.
“Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer”
Noam Shazeer et al · 2017
Later among the works it cites.
Ashish Vaswani et al · 2017
Later among the works it cites.
“Communication-Optimal Convolutional Neural Nets”
James Demmel and Grace Dinh · 2018
Closest in time.
“Beyond Data and Model Parallelism for Deep Neural Networks”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems” Software available from tensorflow.org, 2015
Martin et al · 2015
Cited alongside, same era.
“Exploring the Limits of Language Modeling”
Rafal Jozefowicz et al · 2016
Cited alongside, same era.
“Communication-Avoiding Parallel Sparse-Dense Matrix-Matrix Multiplication”
Penporn Koanantakool et al · 2016
Cited alongside, same era.
Zhihao Jia, Matei Zaharia and Alex Aiken · 2018
Closest in time.
“Exploring Hidden Dimensions in Parallelizing Convolutional Neural Networks”
Zhihao Jia, Sina Lin, Charles Qi and Alex Aiken · 2018
Closest in time.
“Spatially Parallel Convolutions”, 2018
Peter Jin, Boris Ginsburg and Kurt Keutzer · 2018
Closest in time.
“Communication-Avoiding Optimization Methods for Distributed Massive-Scale Sparse Inverse Covariance Estimation”
Penporn Koanantakool et al · 2018
Closest in time.