Fetching the paper…
Reading the bibliography…
The rapid growth of memory and computation requirements of large language models (LLMs) has outpaced the development of hardware, hindering people who lack large-scale high-end GPUs from training or deploying LLMs.
M. Höhfeld and S. E. Fahlman, “Probabilistic rounding in neural network learning with limited precision,” Neurocomputing
1992
Earlier work this paper cites.
S. Ratnasamy, P. Francis, M. Handley, R. Karp, and S. Shenker, “A scalable content-addressable network,” in SIGCOMM
2001
Earlier work this paper cites.
X. Zhang, J. Liu, B. Li, and Y.-S. Yum, “Coolstreaming/donet: a data-driven overlay network for peer-to-peer live media streaming,” in Proceedings IEEE 24th Annual Joint Conference of the IEEE Computer and Communications Societies
2005
Earlier work this paper cites.
R. Thakur, R. Rabenseifner, and W. Gropp, “Optimization of collective communication operations in mpich,” IJHPCA
2005
Earlier work this paper cites.
A. Iosup, S. Ostermann, M. N. Yigitbasi, R. Prodan, T. Fahringer, and D. Epema, “Performance analysis of cloud computing services for many-tasks scientific computing,” IEEE TPDS
2011
Earlier work this paper cites.
F. Niu, B. Recht, C. Re, and S. J. Wright, “Hogwild!: A lock-free approach to parallelizing stochastic gradient descent,” in NeurIPS
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” NeurIPS
2012
Earlier work this paper cites.
X. Chu, X. Chen, A. L. Jia, J. A. Pouwelse, and D. H. Epema, “Dissecting darknets: measurement and performance analysis,” ACM TOIT
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
European Commission, “Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Data Protection Regulation) (Text with EEA relevance),” 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Gruslys, R. Munos, I. Danihelka, M. Lanctot, and A. Graves, “Memory-efficient backpropagation through time,” NeurIPS
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
T. Araki, J. Furukawa, Y. Lindell, A. Nof, and K. Ohara, “High-throughput semi-honest secure three-party computation with an honest majority,” in ACM CCS
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
C. Dünner, T. P. Parnell, and M. Jaggi, “Efficient use of limited-memory accelerators for linear learning on heterogeneous systems,” in NeurIPS
2017
Earlier work this paper cites.
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li, “Terngrad: Ternary gradients to reduce communication in distributed deep learning,” in NeurIPS
2017
Earlier work this paper cites.
B. Hitaj, G. Ateniese, and F. Perez-Cruz, “Deep models under the gan: information leakage from collaborative deep learning,” in ACM SIGSAC
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
H. Qi, E. R. Sparks, and A. Talwalkar, “Paleo: A performance model for deep neural networks,” in ICLR
2017
Earlier work this paper cites.
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differentiation in pytorch,” in NIPS-W
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
H. Li, C. Hu, J. Jiang, Z. Wang, Y. Wen, and W. Zhu, “Jalad: Joint accuracy-and latency-aware deep structure decoupling for edge-cloud execution,” in ICPADS
2018
Earlier work this paper cites.
H. Yu, S. X. Yang, and S. Zhu, “Parallel restarted sgd with faster convergence and less communication: Demystifying why model averaging works for deep learning,” in AAAI
2018
Earlier work this paper cites.
N. Wang, J. Choi, D. Brand, C.-Y. Chen, and K. Gopalakrishnan, “Training deep neural networks with 8-bit floating point numbers,” in Advances in Neural Information Processing Systems
2018
Earlier work this paper cites.
D. Justus, J. Brennan, S. Bonner, and A. S. McGough, “Predicting the computational cost of deep learning models,” in IEEE Big Data
2018
Earlier work this paper cites.
Z. Zhou, X. Chen, E. Li, L. Zeng, K. Luo, and J. Zhang, “Edge intelligence: Paving the last mile of artificial intelligence with edge computing,” Proceedings of the IEEE
2019
Earlier work this paper cites.
S. Shi, Q. Wang, K. Zhao, Z. Tang, Y. Wang, X. Huang, and X. Chu, “A distributed synchronous sgd algorithm with global top-k sparsification for low bandwidth networks,” in ICDCS
2019
Cited alongside, same era.
S. Shi, K. Zhao, Q. Wang, Z. Tang, and X. Chu, “A convergence analysis of distributed sgd with communication-efficient gradient sparsification,” in IJCAI
2019
Cited alongside, same era.
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu, et al
2019
Cited alongside, same era.
D. Narayanan, A. Harlap, A. Phanishayee, V. Seshadri, N. R. Devanur, G. R. Ganger, P. B. Gibbons, and M. Zaharia, “Pipedream: Generalized pipeline parallelism for dnn training,” in SOSP
2019
Cited alongside, same era.
M. Diskin, A. Bukhtiyarov, M. Ryabinin, L. Saulnier, A. Sinitsin, D. Popov, D. V. Pyrkin, M. Kashirin, A. Borzunov, A. Villanova del Moral, et al
2021
Later among the works it cites.
M. Diskin, A. Bukhtiyarov, M. Ryabinin, L. Saulnier, A. Sinitsin, D. Popov, D. V. Pyrkin, M. Kashirin, A. Borzunov, A. Villanova del Moral, et al
2021
Later among the works it cites.
S. Shi, X. Zhou, S. Song, X. Wang, Z. Zhu, X. Huang, X. Jiang, F. Zhou, Z. Guo, L. Xie, et al
2021
Later among the works it cites.
A. Spiridonoff, A. Olshevsky, and I. Paschalidis, “Communication-efficient SGD: From local SGD to one-shot averaging,” in NeurIPS
2021
Later among the works it cites.
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
M. Han, J. Hyun, S. Park, J. Park, and W. Baek, “Mosaic: Heterogeneity-, communication-, and constraint-aware model slicing and execution for accurate and efficient inference,” in PACT
2019
Cited alongside, same era.
F. Haddadpour, M. M. Kamani, M. Mahdavi, and V. Cadambe, “Local sgd with periodic averaging: Tighter analysis and adaptive synchronization,” in NeurIPS
2019
Cited alongside, same era.
L. Zhu, Z. Liu, and S. Han, “Deep leakage from gradients,” NeurIPS
2019
Cited alongside, same era.
M. Nasr, R. Shokri, and A. Houmansadr, “Comprehensive privacy analysis of deep learning: Passive and active white-box inference attacks against centralized and federated learning,” in IEEE S&P
2019
Cited alongside, same era.
Z. Tang, Y. Wang, Q. Wang, and X. Chu, “The impact of gpu dvfs on the energy and performance of deep learning: An empirical study,” in Proceedings of the Tenth ACM International Conference on Future Energy Systems
2019
Cited alongside, same era.
Y. Ma, D. Yu, T. Wu, and H. Wang, “Paddlepaddle: An open-source deep learning platform from industrial practice,” Frontiers of Data and Domputing
2019
Cited alongside, same era.
S. Shi, X. Chu, and B. Li, “Mg-wfbp: Efficient data communication for distributed synchronous sgd algorithms,” in INFOCOM
2019
Cited alongside, same era.
Later among the works it cites.
2021
Later among the works it cites.
OpenAI, “Introducing chatgpt.” https://openai.com/blog/chatgpt , 2022
2022
Later among the works it cites.
B. Ghorbani, O. Firat, M. Freitag, A. Bapna, M. Krikun, X. Garcia, C. Chelba, and C. Cherry, “Scaling laws for neural machine translation,” in ICLR
2022
Later among the works it cites.
I. Alabdulmohsin, B. Neyshabur, and X. Zhai, “Revisiting neural scaling laws in language and vision,” in NeurIPS
2022
Later among the works it cites.
2022
Later among the works it cites.
N. Corporation, “2022 nvidia corporation annual review.” https://s201.q4cdn.com/141608511/files/doc_financials/2022/ar/2022-Annual-Review.pdf , 2022
2022
Later among the works it cites.
C. Lv, C. Niu, R. Gu, X. Jiang, Z. Wang, B. Liu, Z. Wu, Q. Yao, C. Huang, P. Huang, T. Huang, H. Shu, J. Song, B. Zou, P. Lan, G. Xu, F. Wu, S. Tang, F. Wu, and G. Chen, “Walle: An End-to-End, General-Purpose, and Large-Scale production system for Device-Cloud collaborative machine learning,” in OSDI
2022
Later among the works it cites.
K. Jayarajah, D. Wanniarachchige, T. Abdelzaher, and A. Misra, “Comai: Enabling lightweight, collaborative intelligence by retrofitting vision dnns,” in IEEE INFOCOM 2022 - IEEE Conference on Computer Communications
2022
Later among the works it cites.
J. Yoon, Y. Byeon, J. Kim, and H. Lee, “Edgepipe: Tailoring pipeline parallelism with deep neural networks for volatile wireless edge devices,” IEEE Internet of Things Journal
2022
Later among the works it cites.
Z. Tang, S. Shi, B. Li, and X. Chu, “Gossipfl: a decentralized federated learning framework with sparsified and adaptive communication,” IEEE TPDS
2022
Later among the works it cites.
Z. Tang, Y. Zhang, S. Shi, X. He, B. Han, and X. Chu, “Virtual homogeneity learning: Defending against data heterogeneity in federated learning,” in ICML
2022
Later among the works it cites.
J. WANG, B. Yuan, L. Rimanic, Y. He, T. Dao, B. Chen, C. Ré, and C. Zhang, “Fine-tuning language models over slow networks using activation quantization with guarantees,” in NeurIPS
2022
Later among the works it cites.
B. Yuan, Y. He, J. Q. Davis, T. Zhang, T. Dao, B. Chen, P. Liang, C. Re, and C. Zhang, “Decentralized training of foundation models in heterogeneous environments,” in NeurIPS
2022
Later among the works it cites.
X. Li, B. Karimi, and P. Li, “On distributed adaptive optimization with gradient compression,” in ICLR
2022
Later among the works it cites.
K. Mishchenko, F. Bach, M. Even, and B. E. Woodworth, “Asynchronous sgd beats minibatch sgd under arbitrary delays,” in NeurIPS
2022
Later among the works it cites.
F. Wang, W. Zhang, S. Lai, M. Hao, and Z. Wang, “Dynamic gpu energy optimization for machine learning training workloads,” IEEE TPDS
2022
Later among the works it cites.
S. M. Nabavinejad, S. Reda, and M. Ebrahimi, “Coordinated batching and dvfs for dnn inference on gpu accelerators,” IEEE TPDS
2022
Later among the works it cites.
Y. Wang, Q. Wang, and X. Chu, “Energy-efficient online scheduling of transformer inference services on gpu servers,” IEEE Transactions on GreenCom
2022
Later among the works it cites.
H. H. Thorp, “Chatgpt is fun, but not an author,” Science
2023
Closest in time.
2023
Closest in time.
J. Wang, Y. Lu, B. Yuan, B. Chen, P. Liang, C. De Sa, C. Re, and C. Zhang, “CocktailSGD: Fine-tuning foundation models over 500Mbps networks,” in ICML
2023
Closest in time.
L. Zhang, L. Zhang, S. Shi, X. Chu, and B. Li, “Evaluation and optimization of gradient compression for distributed deep learning,” in IEEE ICDCS
2023
Closest in time.
J. You, J.-W. Chung, and M. Chowdhury, “Zeus: Understanding and optimizing GPU energy consumption of DNN training,” in NSDI
2023
Closest in time.