Fetching the paper…
Reading the bibliography…
Communication compression is a crucial technique for modern distributed learning systems to alleviate their communication bottlenecks over slower networks.
Deepsqueeze: Parallel stochastic gradient descent with double-pass error-compensated compression
Hanlin Tang, Xiangru Lian, Shuang Qiu, Lei Yuan, Ce Zhang, Tong Zhang, and Ji Liu · 1907
Earlier work this paper cites.
On the convergence of sgd with biased gradients, 2020
Ahmad Ajalloeian and Sebastian U. Stich · 2008
Earlier work this paper cites.
Base-delta-immediate compression: Practical data compression for on-chip caches
Gennady Pekhimenko, Vivek Seshadri, Onur Mutlu, Michael A Kozuch, Phillip B Gibbons, and Todd C Mowry · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2016
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J Dally · 2016
Earlier work this paper cites.
Zipml: Training linear models with end-to-end low precision, and a little bit of deep learning
Hantian Zhang, Jerry Li, Kaan Kara, Dan Alistarh, Ji Liu, and Ce Zhang · 2017
Earlier work this paper cites.
Terngrad: ternary gradients to reduce communication in distributed deep learning
Wei Wen, Cong Xu, Feng Yan, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2017
Earlier work this paper cites.
Efficient distributed learning with sparsity
Jialei Wang, Mladen Kolar, Nathan Srebro, and Tong Zhang · 2017
Earlier work this paper cites.
signsgd: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Earlier work this paper cites.
Pipe-sgd: a decentralized pipelined sgd framework for distributed deep net training
Youjie Li, Mingchao Yu, Songze Li, Salman Avestimehr, Nam Sung Kim, and Alexander Schwing · 2018
Earlier work this paper cites.
Asynchronous decentralized parallel stochastic gradient descent
Xiangru Lian, Wei Zhang, Ce Zhang, and Ji Liu · 2018
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in tensorflow
Alexander Sergeev and Mike Del Balso · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Earlier work this paper cites.
Gist: Efficient data encoding for deep neural network training
Animesh Jain, Amar Phanishayee, Jason Mars, Lingjia Tang, and Gennady Pekhimenko · 2018
Earlier work this paper cites.
Gradient sparsification for communication-efficient distributed optimization
J Wangni, J Liu, J Wang, and T Zhang · 2018
Earlier work this paper cites.
The convergence of sparsified gradient methods
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Sarit Khirirat, Nikola Konstantinov, and Cédric Renggli · 2018
Earlier work this paper cites.
Atomo: communication-efficient learning via atomic sparsification
Hongyi Wang, Scott Sievert, Zachary Charles, Shengchao Liu, Stephen Wright, and Dimitris Papailiopoulos · 2018
Earlier work this paper cites.
Sketchml: Accelerating distributed machine learning with data sketches
Jiawei Jiang, Fangcheng Fu, Tong Yang, and Bin Cui · 2018
Earlier work this paper cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H Nguyen, Madeleine Gibescu, and Antonio Liotta · 2018
Earlier work this paper cites.
Compressing dma engine: Leveraging activation sparsity for training deep neural networks
Minsoo Rhu, Mike O’Connor, Niladrish Chatterjee, Jeff Pool, Youngeun Kwon, and Stephen W Keckler · 2018
Earlier work this paper cites.
Dnn feature map compression using learned representation over gf (2)
Denis Gudovskiy, Alec Hodgkinson, and Luca Rigazio · 2018
Cited alongside, same era.
Decentralized stochastic optimization and gossip algorithms with compressed communication
Anastasia Koloskova, Sebastian Stich, and Martin Jaggi · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
Backprop with approximate activations for memory-efficient network training
Ayan Chakrabarti and Benjamin Moseley · 2019
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Cited alongside, same era.
Jpeg-act: accelerating deep learning via transform-based lossy compression
R David Evans, Lufei Liu, and Tor M Aamodt · 2020
Later among the works it cites.
Neural network weight compression with nnw-bdi
Andrei Bersatti, Nima Shoghi Ghalehshahi, and Hyesoon Kim · 2020
Later among the works it cites.
Distributed deep learning in open collaborations
Michael Diskin, Alexey Bukhtiyarov, Max Ryabinin, Lucile Saulnier, Anton Sinitsin, Dmitry Popov, Dmitry V Pyrkin, Maxim Kashirin, Alexander Borzunov, Albert Villanova del Moral, et al · 2021
Later among the works it cites.
BAGUA: scaling up distributed learning with system relaxations
Gan, Shaoduo and Lian, Xiangru and Wang, Rui and Chang, Jianbin and Liu, Chengjun and Shi, Hongmei and Zhang, Shengzhuo and Li, Xianghong and Sun, Tengxu and Jiang, Jiawei and others · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri S. Chatterji, Annie S. Chen, Kathleen Creel, Jared Quincy Davis, Dorottya Demszky, Chris Donahue, Moussa Doumbouya, Esin Durmus, Stefano Ermon, John Etchemendy, Kawin Ethayarajh, Li Fei-Fei, Chelsea Finn, Trevor Gale, Lauren Gillespie, Karan Goel, Noah D. Goodman, Shelby Grossman, Neel Guha, Tatsunori Hashimoto, Peter Henderson, John Hewitt, Daniel E. Ho, Jenny Hong, Kyle Hsu, Jing Huang, Thomas Icard, Saahil Jain, Dan Jurafsky, Pratyusha Kalluri, Siddharth Karamcheti, Geoff Keeling, Fereshte Khani, Omar Khattab, Pang Wei Koh, Mark S. Krass, Ranjay Krishna, Rohith Kuditipudi, and et al · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pipedream: generalized pipeline parallelism for dnn training
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R Devanur, Gregory R Ganger, Phillip B Gibbons, and Matei Zaharia · 2019
Cited alongside, same era.
Communication-efficient distributed sgd with sketching
Nikita Ivkin, Daniel Rothchild, Enayat Ullah, Ion Stoica, Raman Arora, et al · 2019
Cited alongside, same era.
Sparse networks from scratch: Faster training without losing performance
Tim Dettmers and Luke Zettlemoyer · 2019
Cited alongside, same era.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Hesham Mostafa and Xin Wang · 2019
Cited alongside, same era.
Dynamic sparse training: Find efficient sparse network from scratch with trainable masked layers
LIU Junjie, XU Zhe, SHI Runbin, Ray CC Cheung, and Hayden KH So · 2019
Cited alongside, same era.
Shapeshifter: Enabling fine-grain data width adaptation in deep learning
Alberto Delmás Lascorz, Sayeh Sharify, Isak Edo, Dylan Malone Stuart, Omar Mohamed Awad, Patrick Judd, Mostafa Mahmoud, Milos Nikolic, Kevin Siu, Zissis Poulos, et al · 2019
Cited alongside, same era.
Accelerating convolutional neural networks via activation map compression
Georgios Georgiadis · 2019
Cited alongside, same era.
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Fairscale: A general purpose modular pytorch library for high performance and large scale training, 2021
Mandeep Baines, Shruti Bhosale, Vittorio Caggiano, Naman Goyal, Siddharth Goyal, Myle Ott, Benjamin Lefaudeux, Vitaliy Liptchinsky, Mike Rabbat, Sam Sheiffer, et al · 2021
Later among the works it cites.
Ac-gc: Lossy activation compression with guaranteed convergence
R David Evans and Tor Aamodt · 2021
Later among the works it cites.
Distributed learning systems with first-order methods
Ji Liu and Ce Zhang · 2021
Later among the works it cites.
1-bit adam: Communication efficient large-scale training with adam’s convergence speed
Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari, Conglong Li, Xiangru Lian, Ji Liu, Ce Zhang, and Yuxiong He · 2021
Later among the works it cites.
Pipemare: Asynchronous pipeline parallel dnn training
Bowen Yang, Jian Zhang, Jonathan Li, Christopher Ré, Christopher Aberger, and Christopher De Sa · 2021
Later among the works it cites.
Memory-efficient pipeline-parallel dnn training
Deepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen, and Matei Zaharia · 2021
Later among the works it cites.
Stochastic sign descent methods: New algorithms and better theory
Mher Safaryan and Peter Richtárik · 2021
Later among the works it cites.
Marina: Faster non-convex distributed learning with compression
Eduard Gorbunov, Konstantin P Burlachenko, Zhize Li, and Peter Richtárik · 2021
Later among the works it cites.
Smoothness matrices beat smoothness constants: better communication compression techniques for distributed optimization
Mher Safaryan, Filip Hanzely, and Peter Richtárik · 2021
Later among the works it cites.
Error compensated distributed sgd can be accelerated
Xun Qian, Peter Richtárik, and Tong Zhang · 2021
Later among the works it cites.
An efficient statistical-based gradient compression technique for distributed training systems
Ahmed M Abdelmoniem, Ahmed Elzanaty, Mohamed-Slim Alouini, and Marco Canini · 2021
Later among the works it cites.
Pufferfish: Communication-efficient models at no extra cost
Hongyi Wang, Saurabh Agarwal, and Dimitris Papailiopoulos · 2021
Later among the works it cites.
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste · 2021
Later among the works it cites.
A novel memory-efficient deep learning training framework via error-bounded lossy compression
Sian Jin, Guanpeng Li, Shuaiwen Leon Song, and Dingwen Tao · 2021
Later among the works it cites.
Fpraker: A processing element for accelerating neural network training
Omar Mohamed Awad, Mostafa Mahmoud, Isak Edo, Ali Hadi Zadeh, Ciaran Bannon, Anand Jayarajan, Gennady Pekhimenko, and Andreas Moshovos · 2021
Later among the works it cites.
Training transformers together
Alexander Borzunov, Max Ryabinin, Tim Dettmers, Quentin Lhoest, Lucile Saulnier, Michael Diskin, and Yacine Jernite · 2022
Closest in time.
Swarm parallelism: Training large models can be surprisingly communication-efficient
Max Ryabinin, Tim Dettmers, Michael Diskin, and Alexander Borzunov · 2023
Closest in time.