Fetching the paper…
Reading the bibliography…
Distributed Mean Estimation (DME), in which $n$ clients communicate vectors to a parameter server that estimates their average, is a fundamental building block in communication-efficient federated learning.
A Note on a Method for Generating Points Uniformly on N-Dimensional Spheres
Muller, M. E · 1959
Earlier work this paper cites.
Quantizing for Minimum Distortion
Max, J · 1960
Earlier work this paper cites.
Picture coding using pseudo-random noise
Roberts, L · 1962
Earlier work this paper cites.
Asymptotically optimal block quantization
Gersho, A · 1979
Earlier work this paper cites.
An algorithm for vector quantizer design
Linde, Y., Buzo, A., and Gray, R · 1980
Earlier work this paper cites.
Least Squares Quantization in PCM
Lloyd, S · 1982
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Quantization
Gray, R. and Neuhoff, D · 1998
Earlier work this paper cites.
Gradient-Based Learning Applied to Document Recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Finding frequent items in data streams
Charikar, M., Chen, K., and Farach-Colton, M · 2002
Earlier work this paper cites.
The Fast Johnson–Lindenstrauss Transform and Approximate Nearest Neighbors
Ailon, N. and Chazelle, B · 2009
Earlier work this paper cites.
Learning Multiple Layers of Features From Tiny Images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Mnist handwritten digit database
LeCun, Y., Cortes, C., and Burges, C · 2010
Earlier work this paper cites.
Uncertainty Principles and Vector Quantization
Lyubarskii, Y. and Vershynin, R · 2010
Earlier work this paper cites.
Nonlinear modeling, estimation and predictive control in APMonitor
Hedengren, J. D., Shishavan, R. A., Powell, K. M., and Edgar, T. F · 2014
Earlier work this paper cites.
1-Bit Stochastic Gradient Descent and Its Application to Data-Parallel Distributed Training of Speech DNNs
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Earlier work this paper cites.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Practical and Optimal LSH for Angular Distance
Andoni, A., Indyk, P., Laarhoven, T., Razenshteyn, I., and Schmidt, L · 2015
Earlier work this paper cites.
A tight gaussian bound for weighted sums of rademacher random variables
Bentkus, V. K. and Dzindzalieta, D · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Orthogonal Random Features
Yu, F. X. X., Suresh, A. T., Choromanski, K. M., Holtmann-Rice, D. N., and Kumar, S · 2016
Earlier work this paper cites.
Sparse Communication for Distributed Gradient Descent
Aji, A. F. and Heafield, K · 2017
Earlier work this paper cites.
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
Alistarh, D., Grubic, D., Li, J., Tomioka, R., and Vojnovic, M · 2017
Earlier work this paper cites.
Federated Learning: Strategies for Improving Communication Efficiency, 2017
Konečný, J., McMahan, H. B., Yu, F. X., Richtárik, P., Suresh, A. T., and Bacon, D · 2017
Earlier work this paper cites.
Communication-Efficient Learning of Deep Networks from Decentralized Data
McMahan, H. B., Moore, E., Ramage, D., Hampson, S., and y Arcas, B. A · 2017
Earlier work this paper cites.
Distributed Mean Estimation With Limited Communication
Suresh, A. T., Felix, X. Y., Kumar, S., and McMahan, H. B · 2017
Cited alongside, same era.
TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Cited alongside, same era.
The Convergence of Sparsified Gradient Methods
Alistarh, D.-A., Hoefler, T., Johansson, M., Konstantinov, N. H., Khirirat, S., and Renggli, C · 2018
Cited alongside, same era.
Gekko optimization suite
Beal, L., Hill, D., Martin, R., and Hedengren, J · 2018
Cited alongside, same era.
signSGD: Compressed Optimisation for Non-Convex Problems
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A · 2018
Cited alongside, same era.
Expanding the Reach of Federated Learning by Reducing Client Resource Requirements
On large-cohort training for federated learning
Charles, Z., Garrett, Z., Huo, Z., Shmulyian, S., and Smith, V · 2021
Later among the works it cites.
New Bounds For Distributed Mean Estimation and Variance Reduction
Davies, P., Gurunanthan, V., Moshrefi, N., Ashkboos, S., and Alistarh, D · 2021
Later among the works it cites.
Efficient Sparse Collective Communication and its Application to Accelerate Distributed Deep Learning
Fei, J., Ho, C.-Y., Sahu, A. N., Canini, M., and Sapio, A · 2021
Later among the works it cites.
vqsgd: Vector quantized stochastic gradient descent
Gandikota, V., Kane, D., Maity, R. K., and Mazumdar, A · 2021
Later among the works it cites.
MARINA: Faster Non-Convex Distributed Learning with Compression
Gorbunov, E., Burlachenko, K. P., Li, Z., and Richtarik, P · 2021
Later among the works it cites.
A Better Alternative to Error Feedback for Communication-Efficient Distributed Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Caldas, S., Konečný, J., McMahan, H. B., and Talwalkar, A · 2018
Cited alongside, same era.
Randomized Distributed Mean Estimation: Accuracy vs. Communication
Konečnỳ, J. and Richtárik, P · 2018
Cited alongside, same era.
Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
Lin, Y., Han, S., Mao, H., Wang, Y., and Dally, B · 2018
Cited alongside, same era.
Sparsified SGD with Memory
Stich, S. U., Cordonnier, J.-B., and Jaggi, M · 2018
Cited alongside, same era.
Gradient sparsification for communication-efficient distributed optimization
Wangni, J., Wang, J., Liu, J., and Zhang, T · 2018
Cited alongside, same era.
Qsparse-local-sgd: Distributed sgd with quantization, sparsification and local computations
Basu, D., Data, D., Karakus, C., and Diggavi, S · 2019
Cited alongside, same era.
Communication-Efficient Distributed SGD With Sketching
Ivkin, N., Rothchild, D., Ullah, E., Braverman, V., Stoica, I., and Arora, R · 2019
Cited alongside, same era.
Horváth, S. and Richtarik, P · 2021
Later among the works it cites.
ATP: In-network Aggregation for Multi-tenant Learning
Lao, C., Le, Y., Mahajan, K., Chen, Y., Wu, W., Akella, A., and Swift, M · 2021
Later among the works it cites.
Adaptive Federated Optimization
Reddi, S. J., Charles, Z., Zaheer, M., Garrett, Z., Rush, K., Konečný, J., Kumar, S., and McMahan, H. B · 2021
Later among the works it cites.
EF21: A New, Simpler, Theoretically Better, and Practically Faster Error Feedback
Richtárik, P., Sokolov, I., and Fatkhullin, I · 2021
Later among the works it cites.
Scaling Distributed Machine Learning with In-Network Aggregation
Sapio, A., Canini, M., Ho, C.-Y., Nelson, J., Kalnis, P., Kim, C., Krishnamurthy, A., Moshref, M., Ports, D., and Richtarik, P · 2021
Later among the works it cites.
SOAR: Minimizing Network Utilization with Bounded In-network Computing
Segal, R., Avin, C., and Scalosub, G · 2021
Later among the works it cites.
DRIVE: One-bit Distributed Mean Estimation
Vargaftik, S., Ben Basat, R., Portnoy, A., Mendelson, G., Ben-Itzhak, Y., and Mitzenmacher, M · 2021
Later among the works it cites.
A Field Guide to Federated Optimization
Wang, J., Charles, Z., Xu, Z., Joshi, G., McMahan, H. B., Al-Shedivat, M., Andrew, G., Avestimehr, S., Daly, K., Data, D., et al · 2021
Later among the works it cites.
Murana: A generic framework for stochastic variance-reduced optimization
Condat, L. and Richtárik, P · 2022
Closest in time.
Natural compression for distributed deep learning
Horvóth, S., Ho, C.-Y., Horvath, L., Sahu, A. N., Canini, M., and Richtárik, P · 2022
Closest in time.
IntSGD: Adaptive floatless compression of stochastic gradients
Mishchenko, K., Wang, B., Kovalev, D., and Richtárik, P · 2022
Closest in time.
Optimizing the communication-accuracy trade-off in federated learning with rate-distortion theory
Mitchell, N., Ballé, J., Charles, Z., and Konečnỳ, J · 2022
Closest in time.
3pc: Three point compressors for communication-efficient distributed training and a better theory for lazy aggregation
Richtárik, P., Sokolov, I., Gasanov, E., Fatkhullin, I., Li, Z., and Gorbunov, E · 2022
Closest in time.
Correlated quantization for distributed mean estimation and optimization
Suresh, A. T., Sun, Z., Ro, J. H., and Yu, F · 2022
Closest in time.
Permutation compressors for provably faster distributed nonconvex optimization
Szlendak, R., Tyurin, A., and Richtárik, P · 2022
Closest in time.
SNARF: A learning-enhanced range filter
Vaidya, K., Kraska, T., Chatterjee, S., Knorr, E. R., Mitzenmacher, M., and Idreos, S · 2022
Closest in time.
EDEN: Communication-Efficient and Robust Distributed Mean Estimation for Federated Learning
Vargaftik, S., Ben Basat, R., Portnoy, A., Mendelson, G., Ben-Itzhak, Y., and Mitzenmacher, M · 2022
Closest in time.
Understanding clipping for federated learning: Convergence and client-level differential privacy
Zhang, X., Chen, X., Hong, M., Wu, S., and Yi, J · 2022
Closest in time.
DoCoFL: Downlink Compression for Cross-Device Federated Learning
Dorfman, R., Vargaftik, S., Ben-Itzhak, Y., and Levy, K. Y · 2023
Closest in time.
Stochastic distributed learning with gradient quantization and double-variance reduction
Horváth, S., Kovalev, D., Mishchenko, K., Richtárik, P., and Stich, S · 2023
Closest in time.
Thc: Accelerating distributed deep learning using tensor homomorphic compression
Li, M., Basat, R. B., Vargaftik, S., Lao, C., Xu, K., Tang, X., Mitzenmacher, M., and Yu, M · 2023
Closest in time.