Fetching the paper…
Reading the bibliography…
With the increasing size of pre-trained language models (PLMs), fine-tuning all the parameters in the model is not efficient, especially when there are a large number of downstream tasks, which incur significant training and storage costs.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1907
Earlier work this paper cites.
DeCAF: A Deep Convolutional Activation Feature for Generic Visual Recognition
Donahue, J.; Jia, Y.; Vinyals, O.; Hoffman, J.; Zhang, N.; Tzeng, E.; and Darrell, T. 2014 · 2014
Earlier work this paper cites.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Goyal, P.; Dollár, P.; Girshick, R. B.; Noordhuis, P.; Wesolowski, L.; Kyrola, A.; Tulloch, A.; Jia, Y.; and He, K. 2017 · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Rebuffi, S.; Bilen, H.; and Vedaldi, A. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Deep Contextualized Word Representations
Peters, M. E.; Neumann, M.; Iyyer, M.; Gardner, M.; Clark, C.; Lee, K.; and Zettlemoyer, L. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; Sutskever, I.; et al. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; de Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019a · 2019
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; de Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019b · 2019
Earlier work this paper cites.
Importance Estimation for Neural Network Pruning
Molchanov, P.; Mallya, A.; Tyree, S.; Frosio, I.; and Kautz, J. 2019 · 2019
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Köpf, A.; Yang, E. Z.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J.; and Chintala, S. 2019 · 2019
Earlier work this paper cites.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. R. 2019 · 2019
Earlier work this paper cites.
PLATON: Pruning Large Transformer Models with Upper Confidence Bound of Weight Importance
Zhang, Q.; Zuo, S.; Liang, C.; Bukharin, A.; He, P.; Chen, W.; and Zhao, T. 2022 · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Cited alongside, same era.
Cross-Attention is All You Need: Adapting Pretrained Transformers for Machine Translation
Gheini, M.; Ren, X.; and May, J. 2021 · 2021
Cited alongside, same era.
Parameter-Efficient Transfer Learning with Diff Pruning
Guo, D.; Rush, A. M.; and Kim, Y. 2021a · 2021
Cited alongside, same era.
Parameter-Efficient Transfer Learning with Diff Pruning
Guo, D.; Rush, A. M.; and Kim, Y. 2021b · 2021
Cited alongside, same era.
The Power of Scale for Parameter-Efficient Prompt Tuning
Lester, B.; Al-Rfou, R.; and Constant, N. 2021a · 2021
Cited alongside, same era.
The Power of Scale for Parameter-Efficient Prompt Tuning
Lester, B.; Al-Rfou, R.; and Constant, N. 2021b · 2021
SparseAdapter: An Easy Approach for Improving the Parameter-Efficiency of Adapters
He, S.; Ding, L.; Dong, D.; Zhang, J.; and Tao, D. 2022b · 2022
Later among the works it cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022 · 2022
Later among the works it cites.
UniPELT: A Unified Framework for Parameter-Efficient Language Model Tuning
Mao, Y.; Mathias, L.; Hou, R.; Almahairi, A.; Ma, H.; Han, J.; Yih, S.; and Khabsa, M. 2022 · 2022
Later among the works it cites.
LST: Ladder Side-Tuning for Parameter and Memory Efficient Transfer Learning
Sung, Y.; Cho, J.; and Bansal, M. 2022 · 2022
Later among the works it cites.
Efficient Fine-Tuning of BERT Models on the Edge
Vucetic, D.; Tayaranian, M.; Ziaeefard, M.; Clark, J. J.; Meyer, B. H.; and Gross, W. J. 2022 · 2022
Later among the works it cites.
AdaMix: Mixture-of-Adapter for Parameter-efficient Tuning of Large Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li, X. L.; and Liang, P. 2021 · 2021
Cited alongside, same era.
Compacter: Efficient Low-Rank Hypercomplex Adapter Layers
Mahabadi, R. K.; Henderson, J.; and Ruder, S. 2021 · 2021
Cited alongside, same era.
AdapterFusion: Non-Destructive Task Composition for Transfer Learning
Pfeiffer, J.; Kamath, A.; Rücklé, A.; Cho, K.; and Gurevych, I. 2021 · 2021
Cited alongside, same era.
Learning Transferable Visual Models From Natural Language Supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021 · 2021
Cited alongside, same era.
Training Neural Networks with Fixed Sparse Masks
Sung, Y.; Nair, V.; and Raffel, C. 2021a · 2021
Cited alongside, same era.
Training Neural Networks with Fixed Sparse Masks
Sung, Y.; Nair, V.; and Raffel, C. 2021b · 2021
Cited alongside, same era.
Wang, Y.; Mukherjee, S.; Liu, X.; Gao, J.; Awadallah, A. H.; and Gao, J. 2022 · 2022
Later among the works it cites.
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
Zaken, E. B.; Goldberg, Y.; and Ravfogel, S. 2022 · 2022
Later among the works it cites.
Parameter-Efficient Fine-Tuning Design Spaces
Chen, J.; Zhang, A.; Shi, X.; Li, M.; Smola, A.; and Yang, D. 2023 · 2023
Closest in time.
QLoRA: Efficient Finetuning of Quantized LLMs
Dettmers, T.; Pagnoni, A.; Holtzman, A.; and Zettlemoyer, L. 2023 · 2023
Closest in time.
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
He, P.; Gao, J.; and Chen, W. 2023 · 2023
Closest in time.
Neural Architecture Search for Parameter-Efficient Fine-tuning of Large Pre-trained Language Models
Lawton, N.; Kumar, A.; Thattai, G.; Galstyan, A.; and Steeg, G. V. 2023 · 2023
Closest in time.
Continual Diffusion: Continual Customization of Text-to-Image Diffusion with C-LoRA
Smith, J. S.; Hsu, Y.; Zhang, L.; Hua, T.; Kira, Z.; Shen, Y.; and Jin, H. 2023 · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lample, G. 2023 · 2023
Closest in time.
One Network, Many Masks: Towards More Parameter-Efficient Transfer Learning
Zeng, G.; Zhang, P.; and Lu, W. 2023 · 2023
Closest in time.
Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
Zhang, Q.; Chen, M.; Bukharin, A.; He, P.; Cheng, Y.; Chen, W.; and Zhao, T. 2023 · 2023
Closest in time.