Fetching the paper…
Reading the bibliography…
To alleviate hardware scarcity in training large deep neural networks (DNNs), particularly large language models (LLMs), we present FusionLLM, a decentralized training system designed and implemented for training DNNs using geo-distributed GPUs across different computing clusters or individual devices.
Probabilistic rounding in neural network learning with limited precision
M. Höhfeld and S. E. Fahlman · 1992
Earlier work this paper cites.
Local Area Network Traffic Locality: Characterization and Application
N. Gulati, C. L. Williamson, and R. B. Bunt · 1993
Earlier work this paper cites.
A fast and high quality multilevel scheme for partitioning irregular graphs
G. Karypis and V. Kumar · 1998
Earlier work this paper cites.
Locality-aware request distribution in cluster-based network servers
V. S. Pai, M. Aron, G. Banga, M. Svendsen, P. Druschel, W. Zwaenepoel, and E. Nahum · 1998
Earlier work this paper cites.
Seti@home-massively distributed computing for seti
E. Korpela, D. Werthimer, D. Anderson, J. Cobb, and M. Leboisky · 2001
Earlier work this paper cites.
Optimization of collective communication operations in mpich
R. Thakur, R. Rabenseifner, and W. Gropp · 2005
Earlier work this paper cites.
Coolstreaming/donet: a data-driven overlay network for peer-to-peer live media streaming
X. Zhang, J. Liu, B. Li, and Y.-S. Yum · 2005
Earlier work this paper cites.
Fast unfolding of communities in large networks
V. D. Blondel, J.-L. Guillaume, R. Lambiotte, and E. Lefebvre · 2008
Earlier work this paper cites.
Performance analysis of cloud computing services for many-tasks scientific computing
A. Iosup, S. Ostermann, M. N. Yigitbasi, R. Prodan, T. Fahringer, and D. Epema · 2011
Earlier work this paper cites.
Efficient BackProp
Y. A. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 2012
Earlier work this paper cites.
Dropout training as adaptive regularization
S. Wager, S. Wang, and P. S. Liang · 2013
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning
S. Chetlur, C. Woolley, P. Vandermersch, J. Cohen, J. Tran, B. Catanzaro, and E. Shelhamer · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Intel Math Kernel Library
E. Wang, Q. Zhang, B. Shen, G. Zhang, X. Lu, Q. Wu, and Y. Wang · 2014
Earlier work this paper cites.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang · 2015
Earlier work this paper cites.
Tiny imagenet visual recognition challenge
Y. Le and X. Yang · 2015
Earlier work this paper cites.
Parallel graph partitioning for complex networks
H. Meyerhenke, P. Sanders, and C. Schulz · 2015
Earlier work this paper cites.
Adding gradient noise improves learning for very deep networks
A. Neelakantan, L. Vilnis, Q. V. Le, I. Sutskever, L. Kaiser, K. Kurach, and J. Martens · 2015
Earlier work this paper cites.
Incentive-compatible experimental design
P. Toulis, D. C. Parkes, E. Pfeffer, and J. Zou · 2015
Earlier work this paper cites.
Heterogeneous environment aware streaming graph partitioning
N. Xu, B. Cui, L. Chen, Z. Huang, and Y. Shao · 2015
Earlier work this paper cites.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, et al · 2016
Earlier work this paper cites.
Qsgd: Randomized quantization for communication-optimal stochastic gradient descent
D. Alistarh, J. Li, R. Tomioka, and M. Vojnovic · 2016
Earlier work this paper cites.
Training deep nets with sublinear memory cost, 2016
T. Chen, B. Xu, C. Zhang, and C. Guestrin · 2016
Earlier work this paper cites.
Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Data Protection Regulation) (Text with EEA relevance), 2016
European Commission · 2016
Earlier work this paper cites.
Memory-efficient backpropagation through time, 2016
A. Gruslys, R. Munos, I. Danihelka, M. Lanctot, and A. Graves · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
S. Zhou, Y. Wu, Z. Ni, X. Zhou, H. Wen, and Y. Zou · 2016
Earlier work this paper cites.
Efficient use of limited-memory accelerators for linear learning on heterogeneous systems
C. Dünner, T. P. Parnell, and M. Jaggi · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Earlier work this paper cites.
Paleo: A performance model for deep neural networks
H. Qi, E. R. Sparks, and A. Talwalkar · 2017
Earlier work this paper cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
A. C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht · 2017
Cited alongside, same era.
Graph edge partitioning via neighborhood heuristic
C. Zhang, F. Wei, Q. Liu, Z. G. Tang, and Z. Li · 2017
Cited alongside, same era.
Jalad: Joint accuracy-and latency-aware deep structure decoupling for edge-cloud execution
H. Li, C. Hu, J. Jiang, Z. Wang, Y. Wen, and W. Zhu · 2018
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Y. Lin, S. Han, H. Mao, Y. Wang, and B. Dally · 2018
Cited alongside, same era.
Ac-gc: Lossy activation compression with guaranteed convergence
R. D. Evans and T. Aamodt · 2021
Later among the works it cites.
Efficient sparse collective communication and its application to accelerate distributed deep learning
J. Fei, C.-Y. Ho, A. N. Sahu, M. Canini, and A. Sapio · 2021
Later among the works it cites.
Lora: Low-rank adaptation of large language models
E. J. Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, et al · 2021
Later among the works it cites.
Beta shapley: a unified and noise-reduced data valuation framework for machine learning
Y. Kwon and J. Zou · 2021
Later among the works it cites.
Efficient large-scale language model training on gpu clusters using megatron-lm
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. McCandlish, J. Kaplan, D. Amodei, and O. D. Team · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Cited alongside, same era.
Training deep neural networks with 8-bit floating point numbers
N. Wang, J. Choi, D. Brand, C.-Y. Chen, and K. Gopalakrishnan · 2018
Cited alongside, same era.
Three mechanisms of weight decay regularization
G. Zhang, C. Wang, B. Xu, and R. Grosse · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2019
Cited alongside, same era.
Mosaic: Heterogeneity-, communication-, and constraint-aware model slicing and execution for accurate and efficient inference
M. Han, J. Hyun, S. Park, J. Park, and W. Baek · 2019
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, et al · 2019
Cited alongside, same era.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Later among the works it cites.
A quantitative survey of communication optimizations in distributed deep learning
S. Shi, Z. Tang, X. Chu, C. Liu, W. Wang, and B. Li · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2021
Later among the works it cites.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2022
Later among the works it cites.
An empirical analysis of compute-optimal large language model training
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, O. Vinyals, J. Rae, and L. Sifre · 2022
Later among the works it cites.
Comai: Enabling lightweight, collaborative intelligence by retrofitting vision dnns
K. Jayarajah, D. Wanniarachchige, T. Abdelzaher, and A. Misra · 2022
Later among the works it cites.
On distributed adaptive optimization with gradient compression
X. Li, B. Karimi, and P. Li · 2022
Later among the works it cites.
Walle: An End-to-End, General-Purpose, and Large-Scale production system for Device-Cloud collaborative machine learning
C. Lv, C. Niu, R. Gu, X. Jiang, Z. Wang, B. Liu, Z. Wu, Q. Yao, C. Huang, P. Huang, T. Huang, H. Shu, J. Song, B. Zou, P. Lan, G. Xu, F. Wu, S. Tang, F. Wu, and G. Chen · 2022
Later among the works it cites.
Introducing chatgpt
OpenAI · 2022
Later among the works it cites.
Gossipfl: a decentralized federated learning framework with sparsified and adaptive communication
Z. Tang, S. Shi, B. Li, and X. Chu · 2022
Later among the works it cites.
Fine-tuning language models over slow networks using activation quantization with guarantees
J. WANG, B. Yuan, L. Rimanic, Y. He, T. Dao, B. Chen, C. Ré, and C. Zhang · 2022
Later among the works it cites.
Edgepipe: Tailoring pipeline parallelism with deep neural networks for volatile wireless edge devices
J. Yoon, Y. Byeon, J. Kim, and H. Lee · 2022
Later among the works it cites.
Decentralized training of foundation models in heterogeneous environments
B. Yuan, Y. He, J. Davis, T. Zhang, T. Dao, B. Chen, P. S. Liang, C. Re, and C. Zhang · 2022
Later among the works it cites.
Summary of chatgpt/gpt-4 research and perspective towards the future of large language models, 2023
Y. Liu, T. Han, S. Ma, J. Zhang, Y. Yang, J. Tian, H. He, A. Li, M. He, Z. Liu, Z. Wu, D. Zhu, X. Li, N. Qiang, D. Shen, T. Liu, and B. Ge · 2023
Later among the works it cites.
Swarm parallelism: Training large models can be surprisingly communication-efficient
M. Ryabinin, T. Dettmers, M. Diskin, and A. Borzunov · 2023
Later among the works it cites.
Fusionai: Decentralized training and deploying llms with massive consumer-level gpus
Z. Tang, Y. Wang, X. He, L. Zhang, X. Pan, Q. Wang, R. Zeng, K. Zhao, S. Shi, B. He, et al · 2023
Later among the works it cites.
Chatgpt is fun, but not an author
H. H. Thorp · 2023
Later among the works it cites.
CocktailSGD: Fine-tuning foundation models over 500Mbps networks
J. Wang, Y. Lu, B. Yuan, B. Chen, P. Liang, C. De Sa, C. Re, and C. Zhang · 2023
Later among the works it cites.
Incentive-aware decentralized data collaboration
Y. Wang, Y. Wu, X. Chen, G. Feng, and B. C. Ooi · 2023
Later among the works it cites.
Hi-speed dnn training with espresso: Unleashing the full potential of gradient compression with near-optimal usage strategies
Z. Wang, H. Lin, Y. Zhu, and T. S. E. Ng · 2023
Later among the works it cites.
Evaluation and optimization of gradient compression for distributed deep learning
L. Zhang, L. Zhang, S. Shi, X. Chu, and B. Li · 2023
Later among the works it cites.
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt, 2023
C. Zhou, Q. Li, C. Li, J. Yu, Y. Liu, G. Wang, K. Zhang, C. Ji, Q. Yan, L. He, H. Peng, J. Li, J. Wu, Z. Liu, P. Xie, C. Xiong, J. Pei, P. S. Yu, and L. Sun · 2023
Later among the works it cites.
Nnfacet: Splitting neural network for concurrent smart sensors
J. Chen, D. Van Le, R. Tan, and D. Ho · 2024
Closest in time.
Fusefl: One-shot federated learning through the lens of causality with progressive model fusion
Z. Tang, Y. Zhang, P. Dong, Y. Cheung, A. C. Zhou, B. Han, and X. Chu · 2024
Closest in time.
Fedimpro: Measuring and improving client update in federated learning
Z. Tang, Y. Zhang, S. Shi, X. Tian, T. Liu, B. Han, and X. Chu · 2024
Closest in time.