Fetching the paper…
Reading the bibliography…
Parameter-Efficient Fine-Tuning (PEFT) methods have gained significant popularity for adapting pre-trained Large Language Models (LLMs) to downstream tasks, primarily due to their potential to significantly reduce memory and computational overheads.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Socialiqa: Commonsense reasoning about social interactions
Sap, M.; Rashkin, H.; Chen, D.; LeBras, R.; and Choi, Y. 2019 · 1904
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 1905
Earlier work this paper cites.
Vapnik-Chervonenkis Theory
Devroye, L.; Györfi, L.; Lugosi, G.; Devroye, L.; Györfi, L.; and Lugosi, G. 1996 · 1996
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T.; Xu, B.; Zhang, C.; and Guestrin, C. 2016 · 2016
Earlier work this paper cites.
MAWPS: A Math Word Problem Repository
Koncel-Kedziorski, R.; Roy, S.; Amini, A.; Kushman, N.; and Hajishirzi, H. 2016 · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2016 · 2016
Earlier work this paper cites.
Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems
Ling, W.; Yogatama, D.; Dyer, C.; and Blunsom, P. 2017 · 2017
Earlier work this paper cites.
Hierarchical representations for efficient architecture search
Liu, H.; Simonyan, K.; Vinyals, O.; Fernando, C.; and Kavukcuoglu, K. 2017 · 2017
Earlier work this paper cites.
Neural Architecture Search with Reinforcement Learning
Zoph, B.; and Le, Q. V. 2017 · 2017
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Howard, J.; and Ruder, S. 2018 · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T.; Clark, P.; Khot, T.; and Sabharwal, A. 2018 · 2018
Earlier work this paper cites.
Efficient Neural Architecture Search via Parameter Sharing
Pham, H.; Guan, M. Y.; Zoph, B.; Le, Q. V.; and Dean, J. 2018 · 2018
Earlier work this paper cites.
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Clark, C.; Lee, K.; Chang, M.-W.; Kwiatkowski, T.; Collins, M.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
PIQA: Reasoning about Physical Commonsense in Natural Language
Bisk, Y.; Zellers, R.; Bras, R. L.; Gao, J.; and Choi, Y. 2020 · 2020
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z.; Chen, M.; Goodman, S.; Gimpel, K.; Sharma, P.; and Soricut, R. 2020 · 2020
Earlier work this paper cites.
Accelerating training of transformer-based language models with progressive layer dropping
Zhang, M.; and He, Y. 2020 · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Hilton, J.; Nakano, R.; Hesse, C.; and Schulman, J. 2021 · 2021
Cited alongside, same era.
Layer-wise model pruning based on mutual information
Fan, C.; Li, J.; Ao, X.; Wu, F.; Meng, Y.; and Sun, X. 2021 · 2021
Cited alongside, same era.
Learn-to-Share: A Hardware-friendly Transfer Learning Framework Exploiting Computation and Parameter Sharing
Fu, C.; Huang, H.; Chen, X.; Tian, Y.; and Zhao, J. 2021 · 2021
Cited alongside, same era.
Towards a unified view of parameter-efficient transfer learning
He, J.; Zhou, C.; Ma, X.; Berg-Kirkpatrick, T.; and Neubig, G. 2021 · 2021
Cited alongside, same era.
Compacter: Efficient low-rank hypercomplex adapter layers
Henderson, J.; Ruder, S.; et al. 2021 · 2021
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Later among the works it cites.
Parameter-Efficient Fine-Tuning Design Spaces
Chen, J.; Zhang, A.; Shi, X.; Li, M.; Smola, A.; and Yang, D. 2023 · 2023
Later among the works it cites.
LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models
Hu, Z.; Lan, Y.; Wang, L.; Xu, W.; Lim, E.-P.; Lee, R. K.-W.; Bing, L.; and Poria, S. 2023 · 2023
Later among the works it cites.
Less is More: Selective Layer Finetuning with SubTuning
Kaplun, G.; Gurevich, A.; Swisa, T.; David, M.; Shalev-Shwartz, S.; and Malach, E. 2023 · 2023
Later among the works it cites.
Surgical fine-tuning improves adaptation to distribution shifts
Lee, Y.; Chen, A. S.; Tajwar, F.; Kumar, A.; Yao, H.; Liang, P.; and Finn, C. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B.; Al-Rfou, R.; and Constant, N. 2021 · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L.; and Liang, P. 2021 · 2021
Cited alongside, same era.
UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning
Mao, Y.; Mathias, L.; Hou, R.; Almahairi, A.; Ma, H.; Han, J.; tau Yih, W.; and Khabsa, M. 2021 · 2021
Cited alongside, same era.
ZeRO-Offload: Democratizing Billion-Scale Model Training
Ren, J.; Rajbhandari, S.; Aminabadi, R. Y.; Ruwase, O.; Yang, S.; Zhang, M.; Li, D.; and He, Y. 2021 · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K.; Bras, R. L.; Bhagavatula, C.; and Choi, Y. 2021 · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B.; and Komatsuzaki, A. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Peng, B.; Li, C.; He, P.; Galley, M.; and Gao, J. 2023 · 2023
Later among the works it cites.
On the effect of dropping layers of pre-trained transformer models
Sajjad, H.; Dalvi, F.; Durrani, N.; and Nakov, P. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
Llama 3 Model Card
AI@Meta. 2024 · 2024
Closest in time.
Layer skip: Enabling early exit inference and self-speculative decoding
Elhoushi, M.; Shrivastava, A.; Liskovich, D.; Hosmer, B.; Wasti, B.; Lai, L.; Mahmoud, A.; Acun, B.; Agarwal, S.; Roman, A.; et al. 2024 · 2024
Closest in time.
Parameter-efficient fine-tuning for large models: A comprehensive survey
Han, Z.; Gao, C.; Liu, J.; Zhang, J.; and Zhang, S. Q. 2024 · 2024
Closest in time.
Conditional adapters: Parameter-efficient transfer learning with fast inference
Lei, T.; Bai, J.; Brahma, S.; Ainslie, J.; Lee, K.; Zhou, Y.; Du, N.; Zhao, V.; Wu, Y.; Li, B.; et al. 2024 · 2024
Closest in time.
DoRA: Weight-Decomposed Low-Rank Adaptation
Liu, S.-Y.; Wang, C.-Y.; Yin, H.; Molchanov, P.; Wang, Y.-C. F.; Cheng, K.-T.; and Chen, M.-H. 2024 · 2024
Closest in time.
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
Pan, R.; Liu, X.; Diao, S.; Pi, R.; Zhang, J.; Han, C.; and Zhang, T. 2024 · 2024
Closest in time.
AutoLoRA: Automatically Tuning Matrix Ranks in Low-Rank Adaptation Based on Meta Learning
Zhang, R.; Qiang, R.; Somayajula, S. A.; and Xie, P. 2024 · 2024
Closest in time.
LIFT: Efficient Layer-wise Fine-tuning for Large Model Models
Zhu, L.; Hu, L.; Lin, J.; and Han, S. 2024 · 2024
Closest in time.
Toolqa: A dataset for llm question answering with external tools
Zhuang, Y.; Yu, Y.; Wang, K.; Sun, H.; and Zhang, C. 2024 · 2024
Closest in time.
Are NLP Models really able to Solve Simple Math Word Problems?
Patel, A.; Bhattamishra, S.; and Goyal, N. 2021 · 2094
Closest in time.