Fetching the paper…
Reading the bibliography…
Decentralized training of deep learning models enables on-device learning over networks, as well as efficient scaling to large compute clusters.
Problems in decentralized decision making and computation
Tsitsiklis, J. N · 1984
Earlier work this paper cites.
Gossip-based computation of aggregate information
Kempe, D., Dobra, A., and Gehrke, J · 2003
Earlier work this paper cites.
Fast linear iterations for distributed averaging
Xiao, L. and Boyd, S · 2004
Earlier work this paper cites.
Randomized gossip algorithms
Boyd, S., Ghosh, A., Prabhakar, B., and Shah, D · 2006
Earlier work this paper cites.
The tradeoffs of large scale learning
Bottou, L. and Bousquet, O · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Nedić, A. and Ozdaglar, A · 2009
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Dekel, O., Gilad-Bachrach, R., Shamir, O., and Xiao, L · 2012
Earlier work this paper cites.
Dual averaging for distributed optimization: Convergence analysis and network scaling
Duchi, J. C., Agarwal, A., and Wainwright, M. J · 2012
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Multi30k: Multilingual english-german image descriptions
Elliott, D., Frank, S., Sima’an, K., and Specia, L · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Efficient distributed online prediction and stochastic optimization with approximate distributed averaging
Tsianos, K. I. and Rabbat, M. G · 2016
Earlier work this paper cites.
On the convergence of decentralized gradient descent
Yuan, K., Ling, Q., and Yin, W · 2016
Earlier work this paper cites.
A downsampled variant of ImageNet as an alternative to the CIFAR datasets
Chrabaszcz, P., Loshchilov, I., and Hutter, F · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Prox-PDA: The proximal primal-dual algorithm for fast distributed nonconvex optimization and learning over networks
Hong, M., Hajinezhad, D., and Zhao, M.-M · 2017
Earlier work this paper cites.
Collaborative deep learning in fixed topology networks
Jiang, Z., Balu, A., Hegde, C., and Sarkar, S · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Lian, X., Zhang, C., Zhang, H., Hsieh, C.-J., Zhang, W., and Liu, J · 2017
Earlier work this paper cites.
Implicit regularization in deep learning
Neyshabur, B · 2017
Cited alongside, same era.
Optimal algorithms for smooth and strongly convex distributed optimization in networks
Scaman, K., Bach, F., Bubeck, S., Lee, Y. T., and Massoulié, L · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Personalized and private peer-to-peer machine learning
Bellet, A., Guerraoui, R., Taziki, M., and Tommasi, M · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
On the relation between the sharpest directions of DNN loss and the SGD step length
Jastrzebski, S., Kenton, Z., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A · 2019
Later among the works it cites.
Advances and open problems in federated learning
Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K., Charles, Z., Cormode, G., Cummings, R., D’Oliveira, R. G. L., Rouayheb, S. E., Evans, D., Gardner, J., Garrett, Z., Gascón, A., Ghazi, B., Gibbons, P. B., Gruteser, M., Harchaoui, Z., He, C., He, L., Huo, Z., Hutchinson, B., Hsu, J., Jaggi, M., Javidi, T., Joshi, G., Khodak, M., Konečný, J., Korolova, A., Koushanfar, F., Koyejo, S., Lepoint, T., Liu, Y., Mittal, P., Mohri, M., Nock, R., Özgür, A., Pagh, R., Raykova, M., Qi, H., Ramage, D., Raskar, R., Song, D., Song, W., Stich, S. U., Sun, Z., Suresh, A. T., Tramèr, F., Vepakomma, P., Wang, J., Xiong, L., Xu, Z., Yang, Q., Yu, F. X., Yu, H., and Zhao, S · 2019
Later among the works it cites.
Decentralized stochastic optimization and gossip algorithms with compressed communication
Koloskova, A., Stich, S. U., and Jaggi, M · 2019
Later among the works it cites.
Hop: Heterogeneity-aware decentralized training
Luo, Q., Lin, J., Zhuo, Y., and Qian, X · 2019
Later among the works it cites.
A simple and fast distributed accelerated gradient method
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G · 2018
Cited alongside, same era.
On consensus-optimality trade-offs in collaborative deep learning
Jiang, Z., Balu, A., Hegde, C., and Sarkar, S · 2018
Cited alongside, same era.
Asynchronous decentralized parallel stochastic gradient descent
Lian, X., Zhang, W., Zhang, C., and Liu, J · 2018
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2018
Cited alongside, same era.
How does batch normalization help optimization?
Santurkar, S., Tsipras, D., Ilyas, A., and Madry, A · 2018
Cited alongside, same era.
Optimal algorithms for non-smooth distributed optimization in networks
Scaman, K., Bach, F., Bubeck, S., Massoulié, L., and Lee, Y. T · 2018
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training
Shallue, C. J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2018
Cited alongside, same era.
Sharma, C., Narayanan, V., and Balamurugan, P · 2019
Later among the works it cites.
Unified optimal analysis of the (stochastic) gradient method
Stich, S. U · 2019
Later among the works it cites.
Distributed non-convex first-order optimization and information processing: Lower complexity bounds and rate optimal algorithms
Sun, H. and Hong, M · 2019
Later among the works it cites.
Matcha: Speeding up decentralized sgd via matching decomposition sampling
Wang, J., Sahu, A. K., Yang, Z., Joshi, G., and Kar, S · 2019
Later among the works it cites.
The early phase of neural network training
Frankle, J., Schwab, D. J., and Morcos, A. S · 2020
Later among the works it cites.
Stochastic weight averaging in parallel: Large-batch training that generalizes well
Gupta, V., Serrano, S. A., and DeCoste, D · 2020
Later among the works it cites.
The non-IID data quagmire of decentralized machine learning
Hsieh, K., Phanishayee, A., Mutlu, O., and Gibbons, P · 2020
Later among the works it cites.
The break-even point on optimization trajectories of deep neural networks
Jastrzebski, S., Szymczak, M., Fort, S., Arpit, D., Tabor, J., Cho, K., and Geras, K · 2020
Later among the works it cites.
AdaScale SGD: A user-friendly algorithm for distributed training
Johnson, T., Agrawal, P., Gu, H., and Guestrin, C · 2020
Later among the works it cites.
Decentralized SGD with asynchronous, local and quantized updates
Nadiradze, G., Sabour, A., Alistarh, D., Sharma, A., Markov, I., and Aksenov, V · 2020
Later among the works it cites.
Decentralized gradient methods: does topology matter?
Neglia, G., Xu, C., Towsley, D., and Calbi, G · 2020
Later among the works it cites.
The error-feedback framework: Better rates for SGD with delayed gradients and compressed updates
Stich, S. U. and Karimireddy, S. P · 2020
Later among the works it cites.
Powergossip: Practical low-rank communication compression in decentralized deep learning
Vogels, T., Karimireddy, S. P., and Jaggi, M · 2020
Later among the works it cites.
Exploring the error-runtime trade-off in decentralized optimization
Wang, J., Sahu, A. K., Joshi, G., and Kar, S · 2020
Later among the works it cites.
Critical parameters for scalable distributed learning with large batches and asynchronous updates
Stich, S. U., Mohtashami, A., and Jaggi, M · 2021
Closest in time.