Fetching the paper…
Reading the bibliography…
Quantization is a fundamental optimization for many machine-learning use cases, including compressing gradients, model weights and activations, and datasets.
A. Aggarwal, M. Klawe, S. Moran, P. Shor, and R. Wilber, “Geometric applications of a matrix searching algorithm,” in Proceedings of the second annual symposium on Computational geometry , 1986, pp. 285–292
1986
Earlier work this paper cites.
Z. Galil and K. Park, “A Linear-time Algorithm for Concave One-dimensional Dynamic Programming,” Information Processing Letters , vol. 33, no. 6, pp. 309–311, 1990
1990
Earlier work this paper cites.
A. T. Suresh, X. Y. Felix, S. Kumar, and H. B. McMahan, “Distributed Mean Estimation With Limited Communication,” in International Conference on Machine Learning . PMLR, 2017, pp. 3329–3337
2017
Earlier work this paper cites.
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic, “QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding,” Advances in Neural Information Processing Systems , vol. 30, pp. 1709–1720, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. Zhang, J. Li, K. Kara, D. Alistarh, J. Liu, and C. Zhang, “ZipML: Training Linear Models with End-to-End Low Precision, and a Little Bit of Deep Learning,” in International Conference on Machine Learning . PMLR, 2017, pp. 4035–4043
2017
Earlier work this paper cites.
J. Konečnỳ and P. Richtárik, “Randomized Distributed Mean Estimation: Accuracy vs. Communication,” Frontiers in Applied Mathematics and Statistics , vol. 4, p. 62, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh et al. , “Mixed Precision Training,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
M. Vladimirova, J. Arbel, and P. Mesejo, “Bayesian Neural Networks Become Heavier-Tailed With Depth,” in NeurIPS 2018-Thirty-second Conference on Neural Information Processing Systems , 2018, pp. 1–7
2018
Earlier work this paper cites.
S. U. Stich, J.-B. Cordonnier, and M. Jaggi, “Sparsified sgd with memory,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
R. Banner, Y. Nahshan, and D. Soudry, “Post Training 4-Bit Quantization of Convolutional Networks for Rapid-Deployment,” in NeurIPS , 2019
2019
Earlier work this paper cites.
T. Vogels, S. P. Karimireddy, and M. Jaggi, “PowerSGD: Practical Low-Rank Gradient Compression for Distributed Optimization,” in Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Cited alongside, same era.
S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer, “Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 05, 2020, pp. 8815–8821
2020
Cited alongside, same era.
M. Safaryan, E. Shulgin, and P. Richtárik, “Uncertainty Principle for Communication Compression in Distributed and Federated Learning and the Search for an Optimal Compressor,” Information and Inference: A Journal of the IMA , 2020
2020
Cited alongside, same era.
X. Ye, P. Dai, J. Luo, X. Guo, Y. Qi, J. Yang, and Y. Chen, “Accelerating CNN Training by Pruning Activation Gradients,” in European Conference on Computer Vision . Springer, 2020, pp. 322–338
2020
V. Monga, Y. Li, and Y. C. Eldar, “Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing,” IEEE Signal Processing Magazine , vol. 38, no. 2, pp. 18–44, 2021
2021
Later among the works it cites.
S. Vargaftik, R. B. Basat, A. Portnoy, G. Mendelson, Y. B. Itzhak, and M. Mitzenmacher, “EDEN: Communication-Efficient and Robust Distributed Mean Estimation for Federated Learning,” in International Conference on Machine Learning . PMLR, 2022, pp. 21 984–22 014
2022
Later among the works it cites.
S. Horvóth, C.-Y. Ho, L. Horvath, A. N. Sahu, M. Canini, and P. Richtárik, “Natural Compression for Distributed Deep Learning,” in Mathematical and Scientific Machine Learning . PMLR, 2022, pp. 129–141
2022
Later among the works it cites.
R. Dorfman, S. Vargaftik, Y. Ben-Itzhak, and K. Y. Levy, “DoCoFL: downlink compression for cross-device federated learning,” in International Conference on Machine Learning . PMLR, 2023, pp. 8356–8388
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
F. Fu, Y. Hu, Y. He, J. Jiang, Y. Shao, C. Zhang, and B. Cui, “Don’t waste your bits! squeeze activations and gradients for deep neural networks via tinyscript,” in International Conference on Machine Learning . PMLR, 2020, pp. 3304–3314
2020
Cited alongside, same era.
F. Faghri, I. Tabrizian, I. Markov, D. Alistarh, D. M. Roy, and A. Ramezani-Kebrya, “Adaptive Gradient Quantization for Data-parallel SGD,” Advances in neural information processing systems , vol. 33, pp. 3174–3185, 2020
2020
Cited alongside, same era.
M. Nagel, R. A. Amjad, M. Van Baalen, C. Louizos, and T. Blankevoort, “Up or Down? Adaptive Rounding for Post-training Quantization,” in International Conference on Machine Learning . PMLR, 2020, pp. 7197–7206
2020
Cited alongside, same era.
Ben Basat, Ran and Mitzenmacher, Michael and Vargaftik, Shay, “How to Send a Real Number Using a Single Bit (And Some Shared Randomness),” in 48th International Colloquium on Automata, Languages, and Programming (ICALP 2021) , vol. 198, 2021, p. 25
2021
Cited alongside, same era.
S. Vargaftik, R. Ben-Basat, A. Portnoy, G. Mendelson, Y. Ben-Itzhak, and M. Mitzenmacher, “Drive: One-bit Distributed Mean Estimation,” Advances in Neural Information Processing Systems , vol. 34, pp. 362–377, 2021
2021
Cited alongside, same era.
B. Chmiel, L. Ben-Uri, M. Shkolnik, E. Hoffer, R. Banner, and D. Soudry, “Neural Gradients are Near-Lognormal: Improved Quantized and Sparse Training,” in International Conference on Learning Representations . OpenReview.net, 2021. [Online]. Available: https://openreview.net/forum?id=EoFNy62JGd
2021
Cited alongside, same era.
A. Ramezani-Kebrya, F. Faghri, I. Markov, V. Aksenov, D. Alistarh, and D. M. Roy, “NUQSGD: Provably Communication-efficient Data-parallel SGD via Nonuniform Quantization,” The Journal of Machine Learning Research , vol. 22, no. 1, pp. 5074–5116, 2021
2021
Cited alongside, same era.
J. Fei, C.-Y. Ho, A. N. Sahu, M. Canini, and A. Sapio, “Efficient Sparse Collective Communication and its Application to Accelerate Distributed Deep Learning,” in Proceedings of the 2021 ACM SIGCOMM 2021 Conference , 2021, pp. 676–691
2021
Cited alongside, same era.
D. Zhou, K. Wang, J. Gu, X. Peng, D. Lian, Y. Zhang, Y. You, and J. Feng, “Dataset Quantization,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 17 205–17 216
2023
Later among the works it cites.
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh, “OPTQ: Accurate quantization for generative pre-trained transformers,” in The Eleventh International Conference on Learning Representations , 2023
2023
Later among the works it cites.
Y. Jeon, C. Lee, K. Park, and H.-y. Kim, “A Frustratingly Easy Post-Training Quantization Scheme for LLMs,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 14 446–14 461
2023
Later among the works it cites.
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Ré, I. Stoica, and C. Zhang, “Flexgen: High-Throughput Generative Inference of Large Language Models With a Single GPU,” in International Conference on Machine Learning . PMLR, 2023, pp. 31 094–31 116
2023
Later among the works it cites.
W. Han, S. Vargaftik, M. Mitzenmacher, B. Karp, and R. Ben Basat, “Beyond Throughput and Compression Ratios: Towards High End-to-end Utility of Gradient Compression,” in HotNets , 2024
2024
Closest in time.
R. B. Basat, S. Vargaftik, A. Portnoy, G. Einziger, Y. Ben-Itzhak, and M. Mitzenmacher, “Accelerating Federated Learning with Quick Distributed Mean Estimation,” in International Conference on Machine Learning , 2024
2024
Closest in time.
X. Chen, S. Vargaftik, and R. Ben-Basat, “When ML Training Cuts Through Congestion: Just-in-Time Gradient Compression via Packet Trimming,” in Hotnets , 2024
2024
Closest in time.
M. Li, R. B. Basat, S. Vargaftik, C. Lao, K. Xu, X. Tang, M. Mitzenmacher, and M. Yu, “THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression,” in USENIX Symposium on Networked Systems Design and Implementation , 2024
2024
Closest in time.