Fetching the paper…
Reading the bibliography…
SOTA decentralized SGD algorithms can overcome the bandwidth bottleneck at the parameter server by using communication collectives like Ring All-Reduce for synchronization.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Nedic, A. and Ozdaglar, A · 2009
Earlier work this paper cites.
Dual averaging for distributed optimization: Convergence analysis and network scaling
Duchi, J. C., Agarwal, A., and Wainwright, M. J · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Dekel, O., Gilad-Bachrach, R., Shamir, O., and Xiao, L · 2012
Earlier work this paper cites.
The tail at scale
Dean, J. and Barroso, L. A · 2013
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Exploiting bounded staleness to speed up big data analytics
Cui, H., Cipar, J., Ho, Q., Kim, J. K., Lee, S., Kumar, A., Wei, J., Dai, W., Ganger, G. R., Gibbons, P. B., et al · 2014
Earlier work this paper cites.
Communication efficient distributed machine learning with the parameter server
Li, M., Andersen, D. G., Smola, A. J., and Yu, K · 2014
Earlier work this paper cites.
{ \{ TensorFlow } \} : a system for { \{ Large-Scale } \} machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Earlier work this paper cites.
Gossip training for deep learning
Blot, M., Picard, D., Cord, M., and Thome, N · 2016
Earlier work this paper cites.
Short-dot: Computing large linear transforms distributedly using coded short dot products
Dutta, S., Cadambe, V., and Grover, P · 2016
Cited alongside, same era.
Omnivore: An optimizer for multi-device deep learning on cpus and gpus
Hadjis, S., Zhang, C., Mitliagkas, I., Iter, D., and Ré, C · 2016
Cited alongside, same era.
How to scale distributed deep learning?
Jin, P. H., Yuan, Q., Iandola, F., and Keutzer, K · 2016
Cited alongside, same era.
Asynchrony begets momentum, with an application to deep learning
Mitliagkas, I., Zhang, C., Hadjis, S., and Ré, C · 2016
Cited alongside, same era.
Excess-risk of distributed stochastic learners
Towfic, Z. J., Chen, J., and Sayed, A. H · 2016
Cited alongside, same era.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Later among the works it cites.
Slow and stale gradients can win the race: Error-runtime trade-offs in distributed sgd
Dutta, S., Joshi, G., Ghosh, S., Dube, P., and Nagpurkar, P · 2018
Later among the works it cites.
Convergence rates for distributed stochastic optimization over random networks
Jakovetic, D., Bajovic, D., Sahu, A. K., and Kar, S · 2018
Later among the works it cites.
Asynchronous decentralized parallel stochastic gradient descent
Lian, X., Zhang, W., Zhang, C., and Liu, J · 2018
Later among the works it cites.
Revisiting small batch training for deep neural networks
Masters, D. and Luschi, C · 2018
Later among the works it cites.
Optimal algorithms for non-smooth distributed optimization in networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the convergence of decentralized gradient descent
Yuan, K., Ling, Q., and Yin, W · 2016
Cited alongside, same era.
Hannah, R. and Yin, W · 2017
Cited alongside, same era.
Speeding up distributed machine learning using codes
Lee, K., Lam, M., Pedarsani, R., Papailiopoulos, D., and Ramchandran, K · 2017
Cited alongside, same era.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Lian, X., Zhang, C., Zhang, H., Hsieh, C.-J., Zhang, W., and Liu, J · 2017
Cited alongside, same era.
Scaman, K., Bach, F., Bubeck, S., Massoulié, L., and Lee, Y. T · 2018
Later among the works it cites.
On nonconvex decentralized gradient descent
Zeng, J. and Yin, W · 2018
Later among the works it cites.
Matcha: Speeding up decentralized sgd via matching decomposition sampling
Wang, J., Sahu, A. K., Yang, Z., Joshi, G., and Kar, S · 2019
Later among the works it cites.
Heterogeneous cpu+ gpu stochastic gradient descent algorithms
Ma, Y. and Rusu, F · 2020
Later among the works it cites.