Fetching the paper…
Reading the bibliography…
We develop and analyze MARINA: a new communication efficient method for non-convex distributed learning over heterogeneous datasets.
Decentralized deep learning with arbitrary communication compression
Koloskova, A., Lin, T., Stich, S. U., and Jaggi, M · 1907
Earlier work this paper cites.
A topological property of real analytic subsets
Łojasiewicz, S · 1963
Earlier work this paper cites.
Gradient methods for the minimisation of functionals
Polyak, B. T · 1963
Earlier work this paper cites.
Some np-complete problems in quadratic and nonlinear programming
Murty, K. and Kabadi, S · 1987
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
From convex to nonconvex: a loss function analysis for binary classification
Zhao, L., Mammadov, M., and Yearwood, J · 2010
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chang, C.-C. and Lin, C.-J · 2011
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Earlier work this paper cites.
Dropping convexity for faster semi-definite optimization
Bhojanapalli, S., Kyrillidis, A., and Sanghavi, S · 2016
Earlier work this paper cites.
Deep learning , volume 1
Goodfellow, I., Bengio, Y., Courville, A., and Bengio, Y · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Federated learning: Strategies for improving communication efficiency
Konečný, J., McMahan, H. B., Yu, F. X., Richtárik, P., Suresh, A. T., and Bacon, D · 2016
Earlier work this paper cites.
Federated learning: strategies for improving communication efficiency
Konečný, J., McMahan, H. B., Yu, F., Richtárik, P., Suresh, A. T., and Bacon, D · 2016
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Lian, X., Zhang, C., Zhang, H., Hsieh, C.-J., Zhang, W., and Liu, J · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
McMahan, H. B., Moore, E., Ramage, D., Hampson, S., and Agüera y Arcas, B · 2017
Cited alongside, same era.
Distributed mean estimation with limited communication
Suresh, A. T., Yu, F. X., Kumar, S., and McMahan, H. B · 2017
Cited alongside, same era.
Near-optimal non-convex optimization via stochastic path integrated differential estimator
Fang, C., Li, C., Lin, Z., and Zhang, T · 2018
Cited alongside, same era.
Implicit regularization in nonconvex statistical estimation: Gradient descent converges linearly for phase retrieval and matrix completion
Ma, C., Wang, K., Chi, Y., and Chen, Y · 2018
Cited alongside, same era.
Optimization for deep learning: theory and algorithms
Sun, R · 2019
Later among the works it cites.
On biased compression for distributed learning
Beznosikov, A., Horváth, S., Richtárik, P., and Safaryan, M · 2020
Later among the works it cites.
Recent theoretical advances in non-convex optimization
Danilova, M., Dvurechensky, P., Gasnikov, A., Gorbunov, E., Guminov, S., Kamzolov, D., and Shibaev, I · 2020
Later among the works it cites.
Improved convergence rates for non-convex federated learning with compression
Das, R., Hashemi, A., Sanghavi, S., and Dhillon, I. S · 2020
Later among the works it cites.
Linearly converging error compensated sgd
Gorbunov, E., Kovalev, D., Makarenko, D., and Richtárik, P · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arjevani, Y., Carmon, Y., Duchi, J. C., Foster, D. J., Srebro, N., and Woodworth, B · 2019
Cited alongside, same era.
Qsparse-local-SGD: Distributed SGD with quantization, sparsification and local computations
Basu, D., Data, D., Karakus, C., and Diggavi, S · 2019
Cited alongside, same era.
Lower bounds for finding stationary points i
Carmon, Y., Duchi, J. C., Hinder, O., and Sidford, A · 2019
Cited alongside, same era.
Natural compression for distributed deep learning
Horváth, S., Ho, C.-Y., Ľudovít Horváth, Sahu, A. N., Canini, M., and Richtárik, P · 2019
Cited alongside, same era.
Stochastic distributed learning with gradient quantization and variance reduction
Horváth, S., Kovalev, D., Mishchenko, K., Stich, S., and Richtárik, P · 2019
Cited alongside, same era.
Advances and open problems in federated learning
Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K., Charles, Z., Cormode, G., Cummings, R., et al · 2019
Cited alongside, same era.
Error feedback fixes signSGD and other gradient compression schemes
Karimireddy, S. P., Rebjock, Q., Stich, S., and Jaggi, M · 2019
Cited alongside, same era.
Later among the works it cites.
Federated learning with compression: Unified analysis and sharp guarantees
Haddadpour, F., Kamani, M. M., Mokhtari, A., and Mahdavi, M · 2020
Later among the works it cites.
Scaffold: Stochastic controlled averaging for federated learning
Karimireddy, S. P., Kale, S., Mohri, M., Reddi, S., Stich, S., and Suresh, A. T · 2020
Later among the works it cites.
Distributed stochastic non-convex optimization: Momentum-based variance reduction
Khanduri, P., Sharma, P., Kafle, S., Bulusu, S., Rajawat, K., and Varshney, P. K · 2020
Later among the works it cites.
A unified analysis of stochastic gradient methods for nonconvex federated optimization
Li, Z. and Richtárik, P · 2020
Later among the works it cites.
Page: A simple and optimal probabilistic gradient estimator for nonconvex optimization
Li, Z., Bao, H., Zhang, X., and Richtárik, P · 2020
Later among the works it cites.
Error compensated distributed sgd can be accelerated
Qian, X., Richtárik, P., and Zhang, T · 2020
Later among the works it cites.
Safaryan, M., Shulgin, E., and Richtárik, P · 2020
Later among the works it cites.
The error-feedback framework: Better rates for sgd with delayed gradients and compressed updates
Stich, S. U. and Karimireddy, S. P · 2020
Later among the works it cites.
Improving the sample and communication complexity for decentralized non-convex optimization: Joint gradient estimation and tracking
Sun, H., Lu, S., and Hong, M · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2020
Later among the works it cites.