Fetching the paper…
Reading the bibliography…
Distributed model training suffers from communication bottlenecks due to frequent model updates transmitted across compute nodes.
Understanding top-k sparsification in distributed deep learning
S. Shi, X. Chu, K. C. Cheung, and S. See · 1911
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky et al · 2009
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. aurelio Ranzato, A. Senior, P. Tucker, K. Yang, Q. V. Le, and A. Y. Ng · 2012
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Firecaffe: near-linear acceleration of deep neural network training on compute clusters
F. N. Iandola, M. W. Moskewicz, K. Ashraf, and K. Keutzer · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Earlier work this paper cites.
Sparse communication for distributed gradient descent
A. F. Aji and K. Heafield · 2017
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Earlier work this paper cites.
Adabatch: Adaptive batch sizes for training deep neural networks
A. Devarakonda, M. Naumov, and M. Garland · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
E. Hoffer, I. Hubara, and D. Soudry · 2017
Earlier work this paper cites.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
M. Honnibal and I. Montani · 2017
Earlier work this paper cites.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Earlier work this paper cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally · 2017
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
S. L. Smith, P.-J. Kindermans, C. Ying, and Q. V. Le · 2017
Cited alongside, same era.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Cited alongside, same era.
Gradient diversity empowers distributed learning
D. Yin, A. Pananjady, M. Lam, D. Papailiopoulos, K. Ramchandran, and P. Bartlett · 2017
Cited alongside, same era.
Distributed learning with sublinear communication
J. Acharya, C. De Sa, D. J. Foster, and K. Sridharan · 2019
Later among the works it cites.
Critical learning periods in deep networks
A. Achille, M. Rovere, and S. Soatto · 2019
Later among the works it cites.
On the relation between the sharpest directions of DNN loss and the SGD step length
S. Jastrzębski, Z. Kenton, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Later among the works it cites.
A distributed synchronous sgd algorithm with global top-k sparsification for low bandwidth networks
S. Shi, Q. Wang, K. Zhao, Z. Tang, Y. Wang, X. Huang, and X. Chu · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. You, I. Gitman, and B. Ginsburg · 2017
Cited alongside, same era.
signsgd: Compressed optimisation for non-convex problems
J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar · 2018
Cited alongside, same era.
Adacomp: Adaptive residual gradient compression for data-parallel distributed training
C.-Y. Chen, J. Choi, D. Brand, A. Agrawal, W. Zhang, and K. Gopalakrishnan · 2018
Cited alongside, same era.
On the computational inefficiency of large batch sizes for stochastic gradient descent
N. Golmant, N. Vemuri, Z. Yao, V. Feinberg, A. Gholami, K. Rothauge, M. W. Mahoney, and J. Gonzalez · 2018
Cited alongside, same era.
Gradient descent happens in a tiny subspace
G. Gur-Ari, D. A. Roberts, and E. Dyer · 2018
Cited alongside, same era.
Squeeze-and-excitation networks
J. Hu, L. Shen, and G. Sun · 2018
Cited alongside, same era.
Horovod: fast and easy distributed deep learning in tensorflow
A. Sergeev and M. Del Balso · 2018
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training
C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl · 2018
Cited alongside, same era.
Local sgd converges fast and communicates little
S. U. Stich · 2019
Later among the works it cites.
S. U. Stich and S. P. Karimireddy · 2019
Later among the works it cites.
Powersgd: Practical low-rank gradient compression for distributed optimization
T. Vogels, S. P. Karimireddy, and M. Jaggi · 2019
Later among the works it cites.
Slow and stale gradients can win the race
S. Dutta, J. Wang, and G. Joshi · 2020
Closest in time.
The early phase of neural network training
J. Frankle, D. J. Schwab, and A. S. Morcos · 2020
Closest in time.
Accelerating distributed deep learning by adaptive gradient quantization
J. Guo, W. Liu, W. Wang, J. Han, R. Li, Y. Lu, and S. Hu · 2020
Closest in time.
The break-even point on optimization trajectories of deep neural networks
S. Jastrzebski, M. Szymczak, S. Fort, D. Arpit, J. Tabor, K. Cho*, and K. Geras* · 2020
Closest in time.
Don’t use large mini-batches, use local sgd
T. Lin, S. U. Stich, K. K. Patel, and M. Jaggi · 2020
Closest in time.
Plink: Discovering and exploiting locality for accelerated distributed training on the public cloud
L. Luo, P. West, J. Nelson, A. Krishnamurthy, and L. Ceze · 2020
Closest in time.
Mlperf training benchmark
P. Mattson, C. Cheng, G. Diamos, C. Coleman, P. Micikevicius, D. Patterson, H. Tang, G.-Y. Wei, P. Bailis, V. Bittorf, D. Brooks, D. Chen, D. Dutta, U. Gupta, K. Hazelwood, A. Hock, X. Huang, D. Kang, D. Kanter, N. Kumar, J. Liao, D. Narayanan, T. Oguntebi, G. Pekhimenko, L. Pentecost, V. Janapa Reddi, T. Robie, T. St John, C.-J. Wu, L. Xu, C. Young, and M. Zaharia · 2020
Closest in time.
Overlap local-sgd: An algorithmic approach to hide communication delays in distributed sgd
J. Wang, H. Liang, and G. Joshi · 2020
Closest in time.