Fetching the paper…
Reading the bibliography…
Large language model fine-tuning has been identified as an efficient approach to applying the pre-trained Large language models to other domains.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification
Hsu, T.-M. H.; Qi, H.; and Brown, M. 2019 · 1909
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2023 · 1910
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020 · 2001
Earlier work this paper cites.
Parameter-Efficient Transfer Learning with Diff Pruning
Guo, D.; Rush, A. M.; and Kim, Y. 2021 · 2012
Earlier work this paper cites.
Seq2SQL: Generating Structured Queries from Natural Language Using Reinforcement Learning
Zhong, V.; Xiong, C.; and Socher, R. 2017 · 2017
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. 2018 · 2018
Earlier work this paper cites.
SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization
Gliwa, B.; Mochol, I.; Biesek, M.; and Wawer, A. 2019 · 2019
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; Laroussilhe, Q. D.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
SCAFFOLD: Stochastic Controlled Averaging for Federated Learning
Karimireddy, S. P.; Kale, S.; Mohri, M.; Reddi, S.; Stich, S.; and Suresh, A. T. 2020 · 2020
Earlier work this paper cites.
Federated Learning Based on Dynamic Regularization
Acar, D. A. E.; Zhao, Y.; Navarro, R. M.; Mattina, M.; Whatmough, P. N.; and Saligrama, V. 2021 · 2021
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li, X. L.; and Liang, P. 2021 · 2021
Earlier work this paper cites.
A Survey on Federated Learning
Zhang, C.; Xie, Y.; Bai, H.; Yu, B.; Li, W.; and Gao, Y. 2021 · 2021
Earlier work this paper cites.
SaPus: Self-adaptive parameter update strategy for DNN training on Multi-GPU clusters
Zhang, Z.; and Wang, C. 2021 · 2021
Earlier work this paper cites.
Federated Learning on Non-IID Data: A Survey
Zhu, H.; Xu, J.; Liu, S.; and Jin, Y. 2021 · 2021
Earlier work this paper cites.
Towards a Unified View of Parameter-Efficient Transfer Learning
He, J.; Zhou, C.; Ma, X.; Berg-Kirkpatrick, T.; and Neubig, G. 2022 · 2022
Cited alongside, same era.
Federated Learning on Non-IID Data Silos: An Experimental Study
Li, Q.; Diao, Y.; Chen, Q.; and He, B. 2022 · 2022
Cited alongside, same era.
UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning
Mao, Y.; Mathias, L.; Hou, R.; Almahairi, A.; Ma, H.; Han, J.; Yih, W.-t.; and Khabsa, M. 2022 · 2022
Cited alongside, same era.
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
Zaken, E. B.; Ravfogel, S.; and Goldberg, Y. 2022 · 2022
Cited alongside, same era.
GPT-4 technical report
OpenAI. 2023 · 2023
Later among the works it cites.
Redpajama-data-v2: An open dataset with 30 trillion tokens for training large language models
Together.AI. 2023 · 2023
Later among the works it cites.
Valipour, M.; Rezagholizadeh, M.; Kobyzev, I.; and Ghodsi, A. 2023 · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. 2023 · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
Abdin, M.; Jacobs, S. A.; Awan, A. A.; Aneja, J.; Awadallah, A.; Awadalla, H.; Bach, N.; Bahree, A.; Bakhtiari, A.; Behl, H.; et al. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; Mihaylov, T.; Ott, M.; Shleifer, S.; Shuster, K.; Simig, D.; Koura, P. S.; Sridhar, A.; Wang, T.; and Zettlemoyer, L. 2022 · 2022
Cited alongside, same era.
Momentum-driven adaptive synchronization model for distributed DNN training on HPC clusters
Zhang, Z.; Ji, Z.; and Wang, C. 2022 · 2022
Cited alongside, same era.
MIPD: An adaptive gradient sparsification framework for distributed DNNs training
Zhang, Z.; and Wang, C. 2022 · 2022
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Heterogeneous lora for federated fine-tuning of on-device foundation models
Cho, Y. J.; Liu, L.; Xu, Z.; Fahrezi, A.; Barnes, M.; and Joshi, G. 2023 · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2023 · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, D.; Xia, F.; Sajjadi, M. S.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; et al. 2023 · 2023
Cited alongside, same era.
C-Coll: Introducing error-bounded lossy compression into MPI collectives
Huang, J.; Di, S.; Yu, X.; Zhai, Y.; Liu, J.; Raffenetti, K.; Zhou, H.; Zhao, K.; Chen, Z.; Cappello, F.; et al. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Not All Federated Learning Algorithms Are Created Equal: A Performance Evaluation Study
Baumgart, G. A.; Shin, J.; Payani, A.; Lee, M.; and Kompella, R. R. 2024 · 2024
Later among the works it cites.
A survey on error-bounded lossy compression for scientific datasets
Di, S.; Liu, J.; Zhao, K.; Liang, X.; Underwood, R.; Zhang, Z.; Shah, M.; Huang, Y.; Huang, J.; Yu, X.; et al. 2024 · 2024
Later among the works it cites.
Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey
Han, Z.; Gao, C.; Liu, J.; Zhang, J.; and Zhang, S. Q. 2024 · 2024
Later among the works it cites.
An optimized error-controlled mpi collective framework integrated with lossy compression
Huang, J.; Di, S.; Yu, X.; Zhai, Y.; Zhang, Z.; Liu, J.; Lu, X.; Raffenetti, K.; Zhou, H.; Zhao, K.; et al. 2024 · 2024
Later among the works it cites.
Liu, W.; Zeng, W.; He, K.; Jiang, Y.; and He, J. 2024 · 2024
Later among the works it cites.
Consent in Crisis: The Rapid Decline of the AI Data Commons
Longpre, S.; Mahari, R.; Lee, A.; Lund, C.; Oderinwale, H.; Brannon, W.; Saxena, N.; Obeng-Marnu, N.; South, T.; Hunter, C.; et al. 2024 · 2024
Later among the works it cites.
Federated Learning With Non-IID Data: A Survey
Lu, Z.; Pan, H.; Dai, Y.; Si, X.; and Zhang, Y. 2024 · 2024
Later among the works it cites.
Introducing meta llama 3: The most capable openly available llm to date
Meta, A. 2024 · 2024
Later among the works it cites.
Rethinking Data Selection for Supervised Fine-Tuning
Shen, M. 2024 · 2024
Later among the works it cites.
FedFa: A Fully Asynchronous Training Paradigm for Federated Learning
Xu, H.; Zhang, Z.; Di, S.; Liu, B.; Khalid, A.; and Cao, J. 2024 · 2024
Later among the works it cites.