Fetching the paper…
Reading the bibliography…
Despite demonstrating superior performance across a variety of linguistic tasks, pre-trained large language models (LMs) often require fine-tuning on specific datasets to effectively address different downstream tasks.
Measuring the effects of non-identical data distribution for federated visual classification
Hsu, T.-M. H.; Qi, H.; and Brown, M. 2019 · 1909
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
LightPAFF: A two-stage distillation framework for pre-training and fine-tuning
Song, K.; Sun, H.; Tan, X.; Qin, T.; Lu, J.; Liu, H.; and Liu, T.-Y. 2020 · 2004
Earlier work this paper cites.
Fedml: A research library and benchmark for federated machine learning
He, C.; Li, S.; So, J.; Zeng, X.; Zhang, M.; Wang, H.; Wang, X.; Vepakomma, P.; Singh, A.; Qiu, H.; et al. 2020 · 2007
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Bird, S.; Klein, E.; and Loper, E. 2009 · 2009
Earlier work this paper cites.
Contextual dueling bandits
Dudík, M.; Hofmann, K.; Schapire, R. E.; Slivkins, A.; and Zoghi, M. 2015 · 2015
Earlier work this paper cites.
Sequence-Level Knowledge Distillation
Kim, Y.; and Rush, A. M. 2016 · 2016
Earlier work this paper cites.
A Diversity-Promoting Objective Function for Neural Conversation Models
Li, J.; Galley, M.; Brockett, C.; Gao, J.; and Dolan, W. B. 2016 · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I.; and Hutter, F. 2016 · 2016
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, B.; Moore, E.; Ramage, D.; Hampson, S.; and y Arcas, B. A. 2017 · 2017
Earlier work this paper cites.
Kudo, T.; and Richardson, J. 2018 · 2018
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Earlier work this paper cites.
Patient Knowledge Distillation for BERT Model Compression
Sun, S.; Cheng, Y.; Gan, Z.; and Liu, J. 2019 · 2019
Earlier work this paper cites.
TinyBERT: Distilling BERT for Natural Language Understanding
Jiao, X.; Yin, Y.; Shang, L.; Jiang, X.; Chen, X.; Li, L.; Wang, F.; and Liu, Q. 2020 · 2020
Earlier work this paper cites.
MixKD: Towards Efficient Distillation of Large-scale Language Models
Liang, K. J.; Hao, W.; Shen, D.; Zhou, Y.; Chen, W.; Chen, C.; and Carin, L. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S.; Rasley, J.; Ruwase, O.; and He, Y. 2020 · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J.; Rajbhandari, S.; Ruwase, O.; and He, Y. 2020 · 2020
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wang, W.; Wei, F.; Dong, L.; Bao, H.; Yang, N.; and Zhou, M. 2020 · 2020
Earlier work this paper cites.
Transformers: State-of-the-Art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Scao, T. L.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. M. 2020 · 2020
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B.; Al-Rfou, R.; and Constant, N. 2021 · 2021
Cited alongside, same era.
Datasets: A Community Library for Natural Language Processing
Lhoest, Q.; Villanova del Moral, A.; Jernite, Y.; Thakur, A.; von Platen, P.; Patil, S.; Chaumond, J.; Drame, M.; Plu, J.; Tunstall, L.; Davison, J.; Šaško, M.; Chhablani, G.; Malik, B.; Brandeis, S.; Le Scao, T.; Sanh, V.; Xu, C.; Patry, N.; McMillan-Major, A.; Schmid, P.; Gugger, S.; Delangue, C.; Matussière, T.; Debut, L.; Bekman, S.; Cistac, P.; Goehringer, T.; Mustar, V.; Lagunas, F.; Rush, A.; and Wolf, T. 2021 · 2021
Cited alongside, same era.
Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM
Conover, M.; Hayes, M.; Mathur, A.; Xie, J.; Wan, J.; Shah, S.; Ghodsi, A.; Wendell, P.; Zaharia, M.; and Xin, R. 2023 · 2023
Later among the works it cites.
MiniLLM: Knowledge distillation of large language models
Gu, Y.; Dong, L.; Wei, F.; and Huang, M. 2023 · 2023
Later among the works it cites.
The false promise of imitating proprietary llms
Gudibande, A.; Wallace, E.; Snell, C.; Geng, X.; Liu, H.; Abbeel, P.; Levine, S.; and Song, D. 2023 · 2023
Later among the works it cites.
DS-1000: A natural and reliable benchmark for data science code generation
Lai, Y.; Li, C.; Wang, Y.; Zhang, T.; Zhong, R.; Zettlemoyer, L.; Yih, W.-t.; Fried, D.; Wang, S.; and Yu, T. 2023 · 2023
Later among the works it cites.
Contrastive Decoding: Open-ended Text Generation as Optimization
Li, X. L.; Holtzman, A.; Fried, D.; Liang, P.; Eisner, J.; Hashimoto, T. B.; Zettlemoyer, L.; and Lewis, M. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li, X. L.; and Liang, P. 2021 · 2021
Cited alongside, same era.
DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts
Liu, A.; Sap, M.; Lu, X.; Swayamdipta, S.; Bhagavatula, C.; Smith, N. A.; and Choi, Y. 2021 · 2021
Cited alongside, same era.
CTR-BERT: Cost-effective knowledge distillation for billion-parameter teacher models
Muhamed, A.; Keivanloo, I.; Perera, S.; Mracek, J.; Xu, Y.; Cui, Q.; Rajagopalan, S.; Zeng, B.; and Chilimbi, T. 2021 · 2021
Cited alongside, same era.
{ \{ Zero-offload } \} : Democratizing { \{ billion-scale } \} model training
Ren, J.; Rajbhandari, S.; Aminabadi, R. Y.; Ruwase, O.; Yang, S.; Zhang, M.; Li, D.; and He, Y. 2021 · 2021
Cited alongside, same era.
MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers
Wang, W.; Bao, H.; Huang, S.; Dong, L.; and Wei, F. 2021 · 2021
Cited alongside, same era.
Accelerate: Training and inference at scale made simple, efficient and adaptable
Gugger, S.; Debut, L.; Wolf, T.; Schmid, P.; Mueller, Z.; Mangrulkar, S.; Sun, M.; and Bossan, B. 2022 · 2022
Cited alongside, same era.
RL with KL penalties is better viewed as Bayesian inference
Korbak, T.; Perez, E.; and Buckley, C. L. 2022 · 2022
Cited alongside, same era.
Fedscale: Benchmarking model and system performance of federated learning at scale
Lai, F.; Dai, Y.; Singapuram, S.; Liu, J.; Zhu, X.; Madhyastha, H.; and Chowdhury, M. 2022 · 2022
Cited alongside, same era.
GPT understands, too
Liu, X.; Zheng, Y.; Du, Z.; Ding, M.; Qian, Y.; Yang, Z.; and Tang, J. 2023 · 2023
Later among the works it cites.
An emulator for fine-tuning large language models using small language models
Mitchell, E.; Rafailov, R.; Sharma, A.; Finn, C.; and Manning, C. D. 2023 · 2023
Later among the works it cites.
CombLM: Adapting Black-Box Language Models through Small Fine-Tuned Models
Ormazabal, A.; Artetxe, M.; and Agirre, E. 2023 · 2023
Later among the works it cites.
Peng, B.; Li, C.; He, P.; Galley, M.; and Gao, J. 2023 · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Roziere, B.; Gehring, J.; Gloeckle, F.; Sootla, S.; Gat, I.; Tan, X. E.; Adi, Y.; Liu, J.; Remez, T.; Rapin, J.; et al. 2023 · 2023
Later among the works it cites.
Large Language Models are Not Yet Human-Level Evaluators for Abstractive Summarization
Shen, C.; Cheng, L.; Nguyen, X.-P.; You, Y.; and Bing, L. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
Pmc-llama: Further finetuning llama on medical papers
Wu, C.; Zhang, X.; Zhang, Y.; Wang, Y.; and Xie, W. 2023 · 2023
Later among the works it cites.
Fedprompt: Communication-efficient and privacy-preserving prompt tuning in federated learning
Zhao, H.; Du, W.; Li, F.; Li, P.; and Liu, G. 2023a · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Dettmers, T.; Pagnoni, A.; Holtzman, A.; and Zettlemoyer, L. 2024 · 2024
Closest in time.
How much RAM does the iPhone have?
iLex. 2024 · 2024
Closest in time.
Prodigy: An Expeditiously Adaptive Parameter-Free Learner
Mishchenko, K.; and Defazio, A. 2024 · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C. D.; Ermon, S.; and Finn, C. 2024 · 2024
Closest in time.
FwdLLM: Efficient FedLLM using Forward Gradient
Xu, M.; Cai, D.; Wu, Y.; Li, X.; and Wang, S. 2024 · 2024
Closest in time.
Navigating Text-To-Image Customization: From LyCORIS Fine-Tuning to Model Evaluation
YEH, S.-Y.; Hsieh, Y.-G.; Gao, Z.; Yang, B. B. W.; Oh, G.; and Gong, Y. 2024 · 2024
Closest in time.
Towards building the federatedGPT: Federated instruction tuning
Zhang, J.; Vahidian, S.; Kuo, M.; Li, C.; Zhang, R.; Yu, T.; Wang, G.; and Chen, Y. 2024a · 2024
Closest in time.
Fine-tuning language models from human preferences
Ziegler, D. M.; Stiennon, N.; Wu, J.; Brown, T. B.; Radford, A.; Amodei, D.; Christiano, P.; and Irving, G. 2019 · 2024
Closest in time.