Fetching the paper…
Reading the bibliography…
We introduce a framework - Artemis - to tackle the problem of learning in a distributed or federated setting with communication constraints and device partial participation.
Distributed Learning with Compressed Gradient Differences
Mishchenko, K., Gorbunov, E., Takáč, M., and Richtárik, P · 1901
Earlier work this paper cites.
Stochastic Distributed Learning with Gradient Quantization and Variance Reduction
Horváth, S., Kovalev, D., Mishchenko, K., Stich, S., and Richtárik, P · 1904
Earlier work this paper cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei, S. and Montanari, A · 1908
Earlier work this paper cites.
Advances and Open Problems in Federated Learning
Kairouz, P., McMahan, H. B., Avent, B., Bellet, A., Bennis, M., Bhagoji, A. N., Bonawitz, K., Charles, Z., Cormode, G., Cummings, R., D’Oliveira, R. G. L., Rouayheb, S. E., Evans, D., Gardner, J., Garrett, Z., Gascón, A., Ghazi, B., Gibbons, P. B., Gruteser, M., Harchaoui, Z., He, C., He, L., Huo, Z., Hutchinson, B., Hsu, J., Jaggi, M., Javidi, T., Joshi, G., Khodak, M., Konečný, J., Korolova, A., Koushanfar, F., Koyejo, S., Lepoint, T., Liu, Y., Mittal, P., Mohri, M., Nock, R., Özgür, A., Pagh, R., Raykova, M., Qi, H., Ramage, D., Raskar, R., Song, D., Song, W., Stich, S. U., Sun, Z., Suresh, A. T., Tramèr, F., Vepakomma, P., Wang, J., Xiong, L., Xu, Z., Yang, Q., Yu, F. X., Yu, H., and Zhao, S · 1912
Earlier work this paper cites.
KDD-Cup 2004: results and analysis
Caruana, R., Joachims, T., and Backstrom, L · 1931
Earlier work this paper cites.
A Double Residual Compression Algorithm for Efficient Distributed Learning
Liu, X., Li, Y., Tang, J., and Yan, M · 1938
Earlier work this paper cites.
A Stochastic Approximation Method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Universal codeword sets and representations of the integers, September 1975
Elias, P · 1975
Earlier work this paper cites.
Zhu, D. L. and Marcotte, P · 1996
Earlier work this paper cites.
On-line learning and stochastic approximations
Bottou, L · 1999
Earlier work this paper cites.
Communication Efficient Sparsification for Large Scale Machine Learning
Khirirat, S., Magnússon, S., Aytekin, A., and Johansson, M · 2003
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Nesterov, Y · 2004
Earlier work this paper cites.
A Better Alternative to Error Feedback for Communication-Efficient Distributed Learning
Horváth, S. and Richtárik, P · 2006
Earlier work this paper cites.
Green Algorithms: Quantifying the carbon emissions of computation
Lannelongue, L., Grealey, J., and Inouye, M · 2007
Earlier work this paper cites.
Visualizing Data using t-SNE
Maaten, L. v. d. and Hinton, G · 2008
Earlier work this paper cites.
Markov Chains and Stochastic Stability
Meyn, S. and Tweedie, R · 2009
Earlier work this paper cites.
Optimal transport : old and new
Villani, C · 2009
Cited alongside, same era.
Large Scale Distributed Deep Networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., Le, Q., and Ng, A · 2012
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, A., Shamir, O., and Sridharan, K · 2012
Cited alongside, same era.
Scaling distributed machine learning with the parameter server
Li, M., Andersen, D. G., Park, J. W., Smola, A. J., Ahmed, A., Josifovski, V., Long, J., Shekita, E. J., and Su, B.-Y · 2014
Cited alongside, same era.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Cited alongside, same era.
Scalable distributed DNN training using commodity GPU cloud computing
Sparsified SGD with Memory
Stich, S. U., Cordonnier, J.-B., and Jaggi, M · 2018
Later among the works it cites.
Error Compensated Quantized SGD and its Applications to Large-scale Distributed Optimization
Wu, J., Huang, W., Huang, J., and Zhang, T · 2018
Later among the works it cites.
DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
Zhou, S., Wu, Y., Ni, Z., Zhou, X., Wen, H., and Zou, Y · 2018
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2019
Later among the works it cites.
SGD: General Analysis and Improved Rates
Gower, R. M., Loizou, N., Qian, X., Sailanbayev, A., Shulgin, E., and Richtárik, P · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Strom, N · 2015
Cited alongside, same era.
Federated Optimization: Distributed Machine Learning for On-Device Intelligence
Konečný, J., McMahan, H. B., Ramage, D., and Richtárik, P · 2016
Cited alongside, same era.
Sparse Communication for Distributed Gradient Descent
Aji, A. F. and Heafield, K · 2017
Cited alongside, same era.
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Cited alongside, same era.
Communication-Efficient Learning of Deep Networks from Decentralized Data
McMahan, B., Moore, E., Ramage, D., Hampson, S., and Arcas, B. A. y · 2017
Cited alongside, same era.
TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Cited alongside, same era.
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., and Chen, z · 2019
Later among the works it cites.
Error Feedback Fixes SignSGD and other Gradient Compression Schemes
Karimireddy, S. P., Rebjock, Q., Stich, S., and Jaggi, M · 2019
Later among the works it cites.
Robust and Communication-Efficient Federated Learning From Non-i.i.d. Data
Sattler, F., Wiedemann, S., Müller, K.-R., and Samek, W · 2019
Later among the works it cites.
Local SGD Converges Fast and Communicates Little
Stich, S. U · 2019
Later among the works it cites.
DoubleSqueeze: Parallel Stochastic Gradient Descent with Double-pass Error-Compensated Compression
Tang, H., Yu, C., Lian, X., Zhang, T., and Liu, J · 2019
Later among the works it cites.
Double Quantization for Communication-Efficient Distributed Optimization
Yu, Y., Wu, J., and Huang, L · 2019
Later among the works it cites.
Communication-Efficient Distributed Blockwise Momentum SGD with Error-Feedback
Zheng, S., Huang, Z., and Kwok, J · 2019
Later among the works it cites.
Acceleration for Compressed Gradient Descent in Distributed and Federated Optimization
Li, Z., Kovalev, D., Qian, X., and Richtarik, P · 2020
Closest in time.
RATQ: A Universal Fixed-Length Quantizer for Stochastic Optimization
Mayekar, P. and Tyagi, H · 2020
Closest in time.
FedPAQ: A Communication-Efficient Federated Learning Method with Periodic Averaging and Quantization
Reisizadeh, A., Mokhtari, A., Hassani, H., Jadbabaie, A., and Pedarsani, R · 2020
Closest in time.
Preserved central model for faster bidirectional compression in distributed settings
Philippenko, C. and Dieuleveut, A · 2021
Closest in time.