Fetching the paper…
Reading the bibliography…
To mitigate communication overheads in distributed model training, several studies propose the use of compressed stochastic gradients, usually achieved by sparsification or quantization.
Some mathematical notes on three-mode factor analysis
Tucker, L. R · 1966
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Optimization of collective communication operations in mpich
Thakur, R., Rabenseifner, R., and Gropp, W · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Senior, A., Tucker, P., Yang, K., Le, Q. V., et al · 2012
Earlier work this paper cites.
Lin, M., Chen, Q., and Yan, S · 2013
Earlier work this paper cites.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
Sainath, T. N., Kingsbury, B., Sindhwani, V., Arisoy, E., and Ramabhadran, B · 2013
Earlier work this paper cites.
Restructuring of deep neural network acoustic models with singular value decomposition
Xue, J., Li, J., and Gong, Y · 2013
Earlier work this paper cites.
Speeding up convolutional neural networks with low rank expansions
Jaderberg, M., Vedaldi, A., and Zisserman, A · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Li, M., Andersen, D. G., Park, J. W., Smola, A. J., Ahmed, A., Josifovski, V., Long, J., Shekita, E. J., and Su, B.-Y · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Mean-normalized stochastic gradient for large-scale deep learning
Wiesler, S., Richard, A., Schluter, R., and Ney, H · 2014
Earlier work this paper cites.
Taming the wild: A unified analysis of hogwild-style algorithms
De Sa, C. M., Zhang, C., Olukotun, K., and Ré, C · 2015
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
Training cnns with low-rank filters for efficient image classification
Ioannou, Y., Robertson, D., Shotton, J., Cipolla, R., and Criminisi, A · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Perturbed iterate analysis for asynchronous stochastic optimization
Mania, H., Pan, X., Papailiopoulos, D., Recht, B., Ramchandran, K., and Jordan, M. I · 2015
Earlier work this paper cites.
Scalable distributed DNN training using commodity gpu cloud computing
Strom, N · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Earlier work this paper cites.
Multi30k: Multilingual english-german image descriptions
Elliott, D., Frank, S., Sima’an, K., and Specia, L · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Network trimming: A data-driven neuron pruning approach towards efficient deep architectures
Hu, H., Peng, R., Tai, Y.-W., and Tang, C.-K · 2016
Earlier work this paper cites.
Binarized neural networks
Hubara, I., Courbariaux, M., Soudry, D., El-Yaniv, R., and Bengio, Y · 2016
Earlier work this paper cites.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size
Iandola, F. N., Han, S., Moskewicz, M. W., Ashraf, K., Dally, W. J., and Keutzer, K · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Randomized distributed mean estimation: Accuracy vs communication
Konečnỳ, J. and Richtárik, P · 2016
Earlier work this paper cites.
Federated learning: Strategies for improving communication efficiency
Konečnỳ, J., McMahan, H. B., Yu, F. X., Richtárik, P., Suresh, A. T., and Bacon, D · 2016
Earlier work this paper cites.
ASAGA: asynchronous parallel SAGA
Leblond, R., Pedregosa, F., and Lacoste-Julien, S · 2016
Earlier work this paper cites.
Pruning filters for efficient convnets
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Cited alongside, same era.
Using the output embedding to improve language models
Press, O. and Wolf, L · 2016
Cited alongside, same era.
Xnor-net: Imagenet classification using binary convolutional neural networks
Rastegari, M., Ordonez, V., Redmon, J., and Farhadi, A · 2016
Cited alongside, same era.
Distributed mean estimation with limited communication
Suresh, A. T., Yu, F. X., Kumar, S., and McMahan, H. B · 2016
Critical learning periods in deep networks
Achille, A., Rovere, M., and Soatto, S · 2018
Later among the works it cites.
High-accuracy low-precision training
De Sa, C., Leszczynski, M., Zhang, J., Marzoev, A., Aberger, C. R., Olukotun, K., and Ré, C · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Later among the works it cites.
Synchronous multi-GPU deep learning with low-precision communication: An experimental study
Grubic, D., Tam, L., Alistarh, D., and Zhang, C · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning structured sparsity in deep neural networks
Wen, W., Wu, C., Wang, Y., Chen, Y., and Li, H · 2016
Cited alongside, same era.
Quantized convolutional neural networks for mobile devices
Wu, J., Leng, C., Wang, Y., Hu, Q., and Cheng, J · 2016
Cited alongside, same era.
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
DoReFa-Net: training low bitwidth convolutional neural networks with low bitwidth gradients
Zhou, S., Wu, Y., Ni, Z., Zhou, X., Wen, H., and Zou, Y · 2016
Cited alongside, same era.
Zhu, C., Han, S., Mao, H., and Dally, W. J · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
Aji, A. F. and Heafield, K · 2017
Cited alongside, same era.
Qsgd: Communication-efficient SGD via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Cited alongside, same era.
Rethinking the value of network pruning
Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T · 2018
Later among the works it cites.
SparCML: high-performance sparse communication for machine learning
Renggli, C., Alistarh, D., and Hoefler, T · 2018
Later among the works it cites.
Sparsified sgd with memory
Stich, S. U., Cordonnier, J.-B., and Jaggi, M · 2018
Later among the works it cites.
Variance-based gradient compression for efficient distributed deep learning
Tsuzuku, Y., Imachi, H., and Akiba, T · 2018
Later among the works it cites.
Atomo: Communication-efficient learning via atomic sparsification
Wang, H., Sievert, S., Liu, S., Charles, Z., Papailiopoulos, D., and Wright, S · 2018
Later among the works it cites.
Error compensated quantized sgd and its applications to large-scale distributed optimization
Wu, J., Huang, W., Huang, J., and Zhang, T · 2018
Later among the works it cites.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
Zhang, X., Zhou, X., Lin, M., and Sun, J · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Later among the works it cites.
Error feedback fixes signsgd and other gradient compression schemes
Karimireddy, S. P., Rebjock, Q., Stich, S., and Jaggi, M · 2019
Later among the works it cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q. V · 2019
Later among the works it cites.
Doublesqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression
Tang, H., Yu, C., Lian, X., Zhang, T., and Liu, J · 2019
Later among the works it cites.
Powersgd: Practical low-rank gradient compression for distributed optimization
Vogels, T., Karimireddy, S. P., and Jaggi, M · 2019
Later among the works it cites.
Drawing early-bird tickets: Towards more efficient training of deep networks
You, H., Li, C., Xu, P., Fu, Y., Wang, Y., Chen, X., Baraniuk, R. G., Wang, Z., and Lin, Y · 2019
Later among the works it cites.
Accordion: Adaptive gradient communication via critical learning regime identification
Agarwal, S., Wang, H., Lee, K., Venkataraman, S., and Papailiopoulos, D · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
Low-rank compression of neural nets: Learning the rank of each layer
Idelbayev, Y. and Carreira-Perpinán, M. A · 2020
Later among the works it cites.
The break-even point on optimization trajectories of deep neural networks
Jastrzebski, S., Szymczak, M., Fort, S., Arpit, D., Tabor, J., Cho, K., and Geras, K · 2020
Later among the works it cites.
A unified architecture for accelerating distributed { \{ DNN } \} training in heterogeneous gpu/cpu clusters
Jiang, Y., Zhu, Y., Lan, C., Yi, B., Cui, Y., and Guo, C · 2020
Later among the works it cites.
The two regimes of deep network training
Leclerc, G. and Madry, A · 2020
Later among the works it cites.
Fixing the train-test resolution discrepancy: Fixefficientnet
Touvron, H., Vedaldi, A., Douze, M., and Jégou, H · 2020
Later among the works it cites.
Principal component networks: Parameter reduction early in training
Waleffe, R. and Rekatsinas, T · 2020
Later among the works it cites.
Go wide, then narrow: Efficient training of deep thin networks
Zhou, D., Ye, M., Chen, C., Meng, T., Tan, M., Song, X., Le, Q., Liu, Q., and Schuurmans, D · 2020
Later among the works it cites.