Fetching the paper…
Reading the bibliography…
Federated learning systems have been identified as an efficient approach to scaling distributed model training with a large amount of participants or data owners while guaranteeing data privacy.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification
Hsu, T.-M. H.; Qi, H.; and Brown, M. 2019 · 1909
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B.; and Brockett, C. 2005 · 2005
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Dagan, I.; Glickman, O.; and Magnini, B. 2006 · 2006
Earlier work this paper cites.
First Quora Dataset Release: Question Pairs
Iyer, S.; Dandekar, N.; and Csernai, K. 2017 · 2017
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Howard, J.; and Ruder, S. 2018 · 2018
Earlier work this paper cites.
SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization
Gliwa, B.; Mochol, I.; Biesek, M.; and Wawer, A. 2019 · 2019
Earlier work this paper cites.
Towards a unified view of parameter-efficient transfer learning
He, J.; Zhou, C.; Ma, X.; Berg-Kirkpatrick, T.; and Neubig, G. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B.; Al-Rfou, R.; and Constant, N. 2021 · 2021
Earlier work this paper cites.
Federated Learning: Opportunities and Challenges
Mammen, P. M. 2021 · 2021
Earlier work this paper cites.
Large scale private learning via low-rank reparametrization
Yu, D.; Zhang, H.; Chen, W.; Yin, J.; and Liu, T.-Y. 2021 · 2021
Earlier work this paper cites.
SaPus: Self-adaptive parameter update strategy for DNN training on Multi-GPU clusters
Zhang, Z.; and Wang, C. 2021 · 2021
Earlier work this paper cites.
Training Compute-Optimal Large Language Models
Hoffmann, J.; Borgeaud, S.; Mensch, A.; Buchatskaya, E.; Cai, T.; Rutherford, E.; de Las Casas, D.; Hendricks, L. A.; Welbl, J.; Clark, A.; Hennigan, T.; Noland, E.; Millican, K.; van den Driessche, G.; Damoc, B.; Guy, A.; Osindero, S.; Simonyan, K.; Elsen, E.; Rae, J. W.; Vinyals, O.; and Sifre, L. 2022 · 2022
Earlier work this paper cites.
UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning
Mao, Y.; Mathias, L.; Hou, R.; Almahairi, A.; Ma, H.; Han, J.; Yih, W.-t.; and Khabsa, M. 2022 · 2022
Cited alongside, same era.
FedBERT: When federated learning meets pre-training
Tian, Y.; Wan, Y.; Lyu, L.; Yao, D.; Jin, H.; and Sun, L. 2022 · 2022
Cited alongside, same era.
GLM-130B: An Open Bilingual Pre-trained Model
Zeng, A.; Liu, X.; Du, Z.; Wang, Z.; Lai, H.; Ding, M.; Yang, Z.; Xu, Y.; Zheng, W.; Xia, X.; et al. 2022 · 2022
Cited alongside, same era.
Momentum-driven adaptive synchronization model for distributed DNN training on HPC clusters
Zhang, Z.; Ji, Z.; and Wang, C. 2022 · 2022
Cited alongside, same era.
MIPD: An adaptive gradient sparsification framework for distributed DNNs training
Zhang, Z.; and Wang, C. 2022 · 2022
Cited alongside, same era.
Introducing ChatGPT
OpenAI. ???? · 2023
Later among the works it cites.
GPT-4 technical report
OpenAI. 2023 · 2023
Later among the works it cites.
S-lora: Serving thousands of concurrent lora adapters
Sheng, Y.; Cao, S.; Li, D.; Hooper, C.; Lee, N.; Yang, S.; Chou, C.; Zhu, B.; Zheng, L.; Keutzer, K.; et al. 2023 · 2023
Later among the works it cites.
Redpajama-data-v2: An open dataset with 30 trillion tokens for training large language models
Together.AI. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Anil, R.; Dai, A. M.; Firat, O.; Johnson, M.; Lepikhin, D.; Passos, A.; Shakeri, S.; Taropa, E.; Bailey, P.; Chen, Z.; et al. 2023 · 2023
Cited alongside, same era.
One-for-All: Generalized LoRA for Parameter-Efficient Fine-tuning
Chavan, A.; Liu, Z.; Gupta, D.; Xing, E.; and Shen, Z. 2023 · 2023
Cited alongside, same era.
Heterogeneous lora for federated fine-tuning of on-device foundation models
Cho, Y. J.; Liu, L.; Xu, Z.; Fahrezi, A.; Barnes, M.; and Joshi, G. 2023 · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2023 · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, D.; Xia, F.; Sajjadi, M. S.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; et al. 2023 · 2023
Cited alongside, same era.
C-Coll: Introducing error-bounded lossy compression into MPI collectives
Huang, J.; Di, S.; Yu, X.; Zhai, Y.; Liu, J.; Raffenetti, K.; Zhou, H.; Zhao, K.; Chen, Z.; Cappello, F.; et al. 2023 · 2023
Cited alongside, same era.
Yi, L.; Yu, H.; Wang, G.; and Liu, X. 2023 · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. 2023 · 2023
Later among the works it cites.
A survey on error-bounded lossy compression for scientific datasets
Di, S.; Liu, J.; Zhao, K.; Liang, X.; Underwood, R.; Zhang, Z.; Shah, M.; Huang, Y.; Huang, J.; Yu, X.; et al. 2024 · 2024
Later among the works it cites.
Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey
Han, Z.; Gao, C.; Liu, J.; Zhang, J.; and Zhang, S. Q. 2024 · 2024
Later among the works it cites.
An optimized error-controlled mpi collective framework integrated with lossy compression
Huang, J.; Di, S.; Yu, X.; Zhai, Y.; Zhang, Z.; Liu, J.; Lu, X.; Raffenetti, K.; Zhou, H.; Zhao, K.; et al. 2024 · 2024
Later among the works it cites.
Consent in Crisis: The Rapid Decline of the AI Data Commons
Longpre, S.; Mahari, R.; Lee, A.; Lund, C.; Oderinwale, H.; Brannon, W.; Saxena, N.; Obeng-Marnu, N.; South, T.; Hunter, C.; et al. 2024 · 2024
Later among the works it cites.
Federated Learning With Non-IID Data: A Survey
Lu, Z.; Pan, H.; Dai, Y.; Si, X.; and Zhang, Y. 2024 · 2024
Later among the works it cites.
FedFa: A Fully Asynchronous Training Paradigm for Federated Learning
Xu, H.; Zhang, Z.; Di, S.; Liu, B.; Khalid, A.; and Cao, J. 2024 · 2024
Later among the works it cites.
ZCCL: Significantly Improving Collective Communication With Error-Bounded Lossy Compression
Huang, J.; Di, S.; Yu, X.; Zhai, Y.; Zhang, Z.; Liu, J.; Lu, X.; Raffenetti, K.; Zhou, H.; Zhao, K.; et al. 2025 · 2025
Closest in time.
CLLoRA: An Approach to Measure the Effects of the Context Length for LLM Fine-Tuning
Zhang, P.; Zhang, Z.; Di, S.; Xin, Y.; and Liu, B. 2025 · 2025
Closest in time.