Fetching the paper…
Reading the bibliography…
We consider decentralized stochastic optimization with the objective function (e.g.
Television by pulse code modulation
Goodall, W. M · 1951
Earlier work this paper cites.
A Stochastic Approximation Method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Picture coding using pseudo-random noise
Roberts, L · 1962
Earlier work this paper cites.
Problems in decentralized decision making and computation
Tsitsiklis, J. N · 1984
Earlier work this paper cites.
Rcv1: A new benchmark collection for text categorization research
Lewis, D. D., Yang, Y., Rose, T. G., and Li, F · 2004
Earlier work this paper cites.
Consensus problems in networks of agents with switching topology and time-delays
Olfati-Saber, R. and Murray, R. M · 2004
Earlier work this paper cites.
Fast linear iterations for distributed averaging
Xiao, L. and Boyd, S · 2004
Earlier work this paper cites.
A scheme for robust distributed sensor fusion based on average consensus
Xiao, L., Boyd, S., and Lall, S · 2005
Earlier work this paper cites.
Randomized gossip algorithms
Boyd, S., Ghosh, A., Prabhakar, B., and Shah, D · 2006
Earlier work this paper cites.
Average consensus on networks with transmission noise or quantization
Carli, R., Fagnani, F., Frasca, P., Taylor, T., and Zampieri, S · 2007
Earlier work this paper cites.
Distributed average consensus with dithered quantization
Aysal, T. C., Coates, M. J., and Rabbat, M. G · 2008
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Nedić, A. and Ozdaglar, A · 2008
Earlier work this paper cites.
Distributed subgradient methods and quantization effects
Nedić, A., Olshevsky, A., Ozdaglar, A., and Tsitsiklis, J. N · 2008
Earlier work this paper cites.
Pascal large scale learning challenge
Sonnenburg, S., Franc, V., Yom-Tov, E., and Sebag, M · 2008
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L · 2010
Earlier work this paper cites.
Distributed estimation of gauss - markov random fields with one-bit quantized data
Fang, J. and Li, H · 2010
Earlier work this paper cites.
A randomized incremental subgradient method for distributed optimization in networked systems
Johansson, B., Rabi, M., and Johansson, M · 2010
Earlier work this paper cites.
Distributed consensus with limited communication data rate
Li, T., Fu, M., Xie, L., and Zhang, J · 2010
Cited alongside, same era.
Dual averaging for distributed optimization: Convergence analysis and network scaling
Duchi, J. C., Agarwal, A., and Wainwright, M. J · 2011
Cited alongside, same era.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E · 2011
Cited alongside, same era.
Optimal distributed online prediction using mini-batches
Dekel, O., Gilad-Bachrach, R., Shamir, O., and Xiao, L · 2012
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K · 2012
Cited alongside, same era.
Distributed average consensus with quantization refinement
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Lian, X., Zhang, C., Zhang, H., Hsieh, C.-J., Zhang, W., and Liu, J · 2017
Later among the works it cites.
Optimal algorithms for smooth and strongly convex distributed optimization in networks
Scaman, K., Bach, F., Bubeck, S., Lee, Y. T., and Massoulié, L · 2017
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Later among the works it cites.
ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning
Zhang, H., Li, J., Kara, K., Alistarh, D., Liu, J., and Zhang, C · 2017
Later among the works it cites.
The convergence of sparsified gradient methods
Alistarh, D., Hoefler, T., Johansson, M., Konstantinov, N., Khirirat, S., and Renggli, C · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thanou, D., Kokiopoulou, E., Pu, Y., and Frossard, P · 2012
Cited alongside, same era.
Distributed alternating direction method of multipliers
Wei, E. and Ozdaglar, A · 2012
Cited alongside, same era.
Distributed dual averaging method for multi-agent optimization with quantized communication
Yuan, D., Xu, S., Zhao, H., and Rong, L · 2012
Cited alongside, same era.
Asynchronous distributed optimization using a randomized alternating direction method of multipliers
Iutzeler, F., Bianchi, P., Ciblat, P., and Hachem, W · 2013
Cited alongside, same era.
Reversible markov chains and random walks on graphs, 2002
Aldous, D. and Fill, J. A · 2014
Cited alongside, same era.
Fast distributed gradient methods
Jakovetić, D., Xavier, J., and Moura, J. M. F · 2014
Cited alongside, same era.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Cited alongside, same era.
Stochastic Gradient Push for Distributed Deep Learning
Assran, M., Loizou, N., Ballas, N., and Rabbat, M · 2018
Later among the works it cites.
Cola: Decentralized linear learning
He, L., Bian, A., and Jaggi, M · 2018
Later among the works it cites.
Randomized Distributed Mean Estimation: Accuracy vs. Communication
Konecny, J. and Richtárik, P · 2018
Later among the works it cites.
Communication-efficient algorithms for decentralized and stochastic optimization
Lan, G., Lee, S., and Zhou, Y · 2018
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Lin, Y., Han, S., Mao, H., Wang, Y., and Dally, B · 2018
Later among the works it cites.
Optimal algorithms for non-smooth distributed optimization in networks
Scaman, K., Bach, F., Bubeck, S., Massoulié, L., and Lee, Y. T · 2018
Later among the works it cites.
Local SGD Converges Fast and Communicates Little
Stich, S. U · 2018
Later among the works it cites.
Sparsified SGD with memory
Stich, S. U., Cordonnier, J.-B., and Jaggi, M · 2018
Later among the works it cites.
A Dual Approach for Optimal Algorithms in Distributed Optimization over Networks
Uribe, C. A., Lee, S., and Gasnikov, A · 2018
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Wangni, J., Wang, J., Liu, J., and Zhang, T · 2018
Later among the works it cites.
Gossip-based computation of aggregate information
Kempe, D., Dobra, A., and Gehrke, J · 2040
Closest in time.