Fetching the paper…
Reading the bibliography…
With the continuous growth in the number of parameters of transformer-based pretrained language models (PLMs), particularly the emergence of large language models (LLMs) with billions of parameters, many natural language processing (NLP) tasks have demonstrated remarkable success.
D. J. MacKay, “A practical bayesian framework for backpropagation networks,” Neural Comput. , vol. 4, no. 3, pp. 448–472, 1992
1992
Earlier work this paper cites.
S.-i. Amari, “Backpropagation and stochastic gradient descent method,” Neurocomputing , vol. 5, no. 4-5, pp. 185–196, 1993
1993
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proc. Annu. Meeting Assoc. Comput. Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
Q. Le, T. Sarlós, and A. Smola, “Fastfood-computing hilbert space expansions in loglinear time,” in Proc. Int. Conf. Mach. Learn. PMLR, 2013, pp. 244–252
2013
Earlier work this paper cites.
J. Solomon, F. De Goes, G. Peyré, M. Cuturi, A. Butscher, A. Nguyen, T. Du, and L. Guibas, “Convolutional wasserstein distances: Efficient optimal transportation on geometric domains,” ACM Trans. Graph. , vol. 34, no. 4, pp. 1–11, 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2016, pp. 770–778
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Proc. Adv. Neural Inf. Process. Syst. , vol. 30, 2017
2017
Earlier work this paper cites.
S.-A. Rebuffi, H. Bilen, and A. Vedaldi, “Learning multiple visual domains with residual adapters,” Proc. Adv. Neural Inf. Process. Syst. , vol. 30, 2017
2017
Earlier work this paper cites.
D. Ha, A. M. Dai, and Q. V. Le, “Hypernetworks,” in Proc. Int. Conf. Learn. Representations , 2017
2017
Earlier work this paper cites.
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning filters for efficient convnets,” in Proc. Int. Conf. Learn. Representations , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. C. Mocanu, E. Mocanu, P. Stone, P. H. Nguyen, M. Gibescu, and A. Liotta, “Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science,” Nature communications , vol. 9, no. 1, p. 2383, 2018
2018
Earlier work this paper cites.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “GLUE: A multi-task benchmark and analysis platform for natural language understanding,” in Proc. of 2018 EMNLP Workshop BlackboxNLP , 2018, pp. 353–355
2018
Earlier work this paper cites.
S.-A. Rebuffi, H. Bilen, and A. Vedaldi, “Efficient parametrization of multi-domain deep neural networks,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2018, pp. 8119–8127
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. Conf. North Amer. Chapter Assoc. Comput. Linguistics: Hum. Lang. Technol. , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for nlp,” in Proc. Int. Conf. Mach. Learn. PMLR, 2019, pp. 2790–2799
2019
Earlier work this paper cites.
R. Dabre, A. Fujita, and C. Chu, “Exploiting multilingualism through multistage fine-tuning for low-resource neural machine translation,” in Proc. Conf. Empir. Methods Natural Lang. Process., Int. Joint Conf. Natural Lang. Process. , 2019, pp. 1410–1416
2019
Earlier work this paper cites.
K. Clark, U. Khandelwal, O. Levy, and C. D. Manning, “What does BERT look at? an analysis of BERT’s attention,” in Proc. of 2019 ACL Workshop BlackboxNLP , 2019, pp. 276–286
2019
Earlier work this paper cites.
O. Kovaleva, A. Romanov, A. Rogers, and A. Rumshisky, “Revealing the dark secrets of BERT,” in Proc. Conf. Empir. Methods Natural Lang. Process., Int. Joint Conf. Natural Lang. Process. , 2019, pp. 4365–4374
2019
Earlier work this paper cites.
T. Elsken, J. H. Metzen, and F. Hutter, “Neural architecture search: A survey,” J. Mach. Learn. Res. , vol. 20, no. 1, pp. 1997–2017, 2019
2019
Earlier work this paper cites.
J. Frankle and M. Carbin, “The lottery ticket hypothesis: Finding sparse, trainable neural networks,” in Proc. Int. Conf. Learn. Representations , 2019
2019
Earlier work this paper cites.
P. Molchanov, A. Mallya, S. Tyree, I. Frosio, and J. Kautz, “Importance estimation for neural network pruning,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. , 2019, pp. 11 264–11 272
2019
Earlier work this paper cites.
N. Lee, T. Ajanthan, and P. Torr, “SNIP: Single-shot network pruning based on connection sensitivity,” in Proc. Int. Conf. Learn. Representations , 2019
2019
Earlier work this paper cites.
A. Bapna and O. Firat, “Simple, scalable adaptation for neural machine translation,” in Proc. Conf. Empir. Methods Natural Lang. Process., Int. Joint Conf. Natural Lang. Process. , 2019, pp. 1538–1548
2019
Earlier work this paper cites.
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “Roberta: A robustly optimized bert pretraining approach,” in Proc. Int. Conf. Learn. Representations , 2020
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , vol. 21, no. 1, pp. 5485–5551, 2020
2020
Earlier work this paper cites.
Z. Lin, A. Madotto, and P. Fung, “Exploring versatile generative language model via parameter-efficient transfer learning,” in Proc. Findings Conf. Empir. Methods Natural Lang. Process. , 2020, pp. 441–459
2020
Earlier work this paper cites.
M. Zhao, T. Lin, F. Mi, M. Jaggi, and H. Schütze, “Masking as an efficient alternative to finetuning for pretrained language models,” in Proc. Conf. Empir. Methods Natural Lang. Process. , 2020, pp. 2226–2241
2020
Earlier work this paper cites.
Y. Xie, W. Yang, L. Tan, K. Xiong, N. J. Yuan, B. Huai, M. Li, and J. Lin, “Distant supervision for multi-stage fine-tuning in retrieval-based question answering,” in Proceedings of The Web Conference , 2020, pp. 2934–2940
2020
Earlier work this paper cites.
J. Pfeiffer, I. Vulić, I. Gurevych, and S. Ruder, “MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer,” in Proc. Conf. Empir. Methods Natural Lang. Process. , 2020, pp. 7654–7673
2020
Earlier work this paper cites.
S. P. Singh and M. Jaggi, “Model fusion via optimal transport,” Proc. Adv. Neural Inf. Process. Syst. , vol. 33, pp. 22 045–22 055, 2020
2020
Earlier work this paper cites.
R. Aharoni and Y. Goldberg, “Unsupervised domain clusters in pretrained language models,” in Proc. Annu. Meeting Assoc. Comput. Linguistics , 2020, pp. 7747–7763
2020
Earlier work this paper cites.
V. Sanh, T. Wolf, and A. Rush, “Movement pruning: Adaptive sparsity by fine-tuning,” Proc. Adv. Neural Inf. Process. Syst. , vol. 33, pp. 20 378–20 389, 2020
2020
Earlier work this paper cites.
J. Liu, A. Moreau, M. Preuss, J. Rapin, B. Roziere, F. Teytaud, and O. Teytaud, “Versatile black-box optimization,” in Proc. of the 2020 Genet. and Evolut. Comput. Conf. , 2020, pp. 620–628
2020
Earlier work this paper cites.
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut, “Albert: A lite bert for self-supervised learning of language representations,” in Proc. Int. Conf. Learn. Representations , 2020
2020
Earlier work this paper cites.
M. Artetxe, S. Ruder, and D. Yogatama, “On the cross-lingual transferability of monolingual representations,” in Proc. Annu. Meeting Assoc. Comput. Linguistics , 2020, pp. 4623–4637
2020
Earlier work this paper cites.
X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,” in Proc. Annu. Meeting Assoc. Comput. Linguistics, Int. Joint Conf. Natural Lang. Process. , 2021, pp. 4582–4597
2021
Earlier work this paper cites.
A. Rücklé, G. Geigle, M. Glockner, T. Beck, J. Pfeiffer, N. Reimers, and I. Gurevych, “AdapterDrop: On the efficiency of adapters in transformers,” in Proc. Conf. Empir. Methods Natural Lang. Process. , 2021, pp. 7930–7946
2021
Earlier work this paper cites.
J. Pfeiffer, A. Kamath, A. Rücklé, K. Cho, and I. Gurevych, “AdapterFusion: Non-destructive task composition for transfer learning,” in Proc. Conf. Eur. Chapter Assoc. Comput. Linguistics , 2021, pp. 487–503
2021
Cited alongside, same era.
R. Karimi Mahabadi, S. Ruder, M. Dehghani, and J. Henderson, “Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks,” in Proc. Annu. Meeting Assoc. Comput. Linguistics, Int. Joint Conf. Natural Lang. Process. , 2021, pp. 565–576
2021
Cited alongside, same era.
K. Hambardzumyan, H. Khachatrian, and J. May, “WARP: Word-level Adversarial ReProgramming,” in Proc. Annu. Meeting Assoc. Comput. Linguistics, Int. Joint Conf. Natural Lang. Process. , 2021, pp. 4921–4933
2021
Cited alongside, same era.
B. Lester, R. Al-Rfou, and N. Constant, “The power of scale for parameter-efficient prompt tuning,” in Proc. Conf. Empir. Methods Natural Lang. Process. , 2021, pp. 3045–3059
2021
Cited alongside, same era.
2023
Closest in time.
E. Almazrouei, H. Alobeidli, A. Alshamsi, A. Cappelli, R. Cojocaru, M. Alhammadi, M. Daniele, D. Heslow, J. Launay, Q. Malartic et al. , “The falcon series of language models: Towards open frontier models,” Hugging Face repository , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
Y.-L. Sung, V. Nair, and C. Raffel, “Training neural networks with fixed sparse masks,” in Proc. Adv. Neural Inf. Process. Syst. , 2021
2021
Cited alongside, same era.
R. Xu, F. Luo, Z. Zhang, C. Tan, B. Chang, S. Huang, and F. Huang, “Raise a child in large language model: Towards effective and generalizable fine-tuning,” in Proc. Conf. Empir. Methods Natural Lang. Process. , 2021, pp. 9514–9528
2021
Cited alongside, same era.
D. Guo, A. Rush, and Y. Kim, “Parameter-efficient transfer learning with diff pruning,” in Proc. Annu. Meeting Assoc. Comput. Linguistics, Int. Joint Conf. Natural Lang. Process. , 2021, pp. 4884–4896
2021
Cited alongside, same era.
A. Aghajanyan, S. Gupta, and L. Zettlemoyer, “Intrinsic dimensionality explains the effectiveness of language model fine-tuning,” in Proc. Annu. Meeting Assoc. Comput. Linguistics, Int. Joint Conf. Natural Lang. Process. , 2021, pp. 7319–7328
2021
Cited alongside, same era.
R. Karimi Mahabadi, J. Henderson, and S. Ruder, “Compacter: Efficient low-rank hypercomplex adapter layers,” Proc. Adv. Neural Inf. Process. Syst. , vol. 34, pp. 1022–1035, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
A. Zhang, Y. Tay, S. Zhang, A. Chan, A. T. Luu, S. Hui, and J. Fu, “Beyond fully-connected layers with quaternions: Parameterization of hypercomplex multiplications with $1/n$ parameters,” in Proc. Int. Conf. Learn. Representations , 2021
2021
Cited alongside, same era.
2023
Closest in time.
Z. Wang, R. Panda, L. Karlinsky, R. Feris, H. Sun, and Y. Kim, “Multitask prompt tuning enables parameter-efficient transfer learning,” in Proc. Int. Conf. Learn. Representations , 2023
2023
Closest in time.
X. Yang, J. Y. Huang, W. Zhou, and M. Chen, “Parameter-efficient tuning with special token adaptation,” in Proc. Conf. Eur. Chapter Assoc. Comput. Linguistics , 2023, pp. 865–872
2023
Closest in time.
Y. Chen, Q. Fu, G. Fan, L. Du, J.-G. Lou, S. Han, D. Zhang, Z. Li, and Y. Xiao, “Hadamard adapter: An extreme parameter-efficient adapter tuning method for pre-trained language models,” in Proc. 32nd ACM Int. Conf. Inf. Knowl. Manage. , 2023, pp. 276–285
2023
Closest in time.
N. Lawton, A. Kumar, G. Thattai, A. Galstyan, and G. Ver Steeg, “Neural architecture search for parameter-efficient fine-tuning of large pre-trained language models,” in Proc. Findings Assoc. Comput. Linguistics , 2023, pp. 8506–8515
2023
Closest in time.
Z. Fu, H. Yang, A. M.-C. So, W. Lam, L. Bing, and N. Collier, “On the effectiveness of parameter-efficient fine-tuning,” in Proc. AAAI Conf. Artif. Intell. , vol. 37, no. 11, 2023, pp. 12 799–12 807
2023
Closest in time.
M. Valipour, M. Rezagholizadeh, I. Kobyzev, and A. Ghodsi, “DyLoRA: Parameter-efficient tuning of pre-trained models using dynamic search-free low-rank adaptation,” in Proc. Conf. Eur. Chapter Assoc. Comput. Linguistics , 2023, pp. 3274–3287
2023
Closest in time.
Q. Zhang, M. Chen, A. Bukharin, P. He, Y. Cheng, W. Chen, and T. Zhao, “Adaptive budget allocation for parameter-efficient fine-tuning,” in Proc. Int. Conf. Learn. Representations , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
J. Chen, A. Zhang, X. Shi, M. Li, A. Smola, and D. Yang, “Parameter-efficient fine-tuning design spaces,” in Proc. Int. Conf. Learn. Representations , 2023
2023
Closest in time.
G. Zeng, P. Zhang, and W. Lu, “One network, many masks: Towards more parameter-efficient transfer learning,” in Proc. Annu. Meeting Assoc. Comput. Linguistics , 2023, pp. 7564–7580
2023
Closest in time.
L. Xu and W. Wang, “Improving aspect-based sentiment analysis with contrastive learning,” Natural Language Processing Journal , vol. 3, p. 100009, 2023
2023
Closest in time.
M. T. Hosseini, A. Ghaffari, M. S. Tahaei, M. Rezagholizadeh, M. Asgharian, and V. P. Nia, “Towards fine-tuning pre-trained language models with integer forward and backward propagation,” in Proc. Findings Assoc. Comput. Linguistics , 2023, pp. 1867–1876
2023
Closest in time.
L. Xu, H. Xie, Z. Li, F. L. Wang, W. Wang, and Q. Li, “Contrastive learning models for sentence representations,” ACM Trans. Intel. Syst. Tec. , vol. 14, no. 4, pp. 1–34, 2023
2023
Closest in time.
G. Ilharco, M. T. Ribeiro, M. Wortsman, L. Schmidt, H. Hajishirzi, and A. Farhadi, “Editing models with task arithmetic,” in Proc. Int. Conf. Learn. Representations , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto, “Stanford alpaca: An instruction-following llama model,” 2023
2023
Closest in time.
N. Gu, P. Fu, X. Liu, Z. Liu, Z. Lin, and W. Wang, “A gradient control method for backdoor attacks on parameter-efficient tuning,” in Proc. Annu. Meeting Assoc. Comput. Linguistics , 2023, pp. 3508–3520
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
X. He, C. Li, P. Zhang, J. Yang, and X. E. Wang, “Parameter-efficient model adaptation for vision transformers,” in Proc. AAAI Conf. Artif. Intell. , vol. 37, no. 1, 2023, pp. 817–825
2023
Closest in time.
Z. Xu, Z. Chen, Y. Zhang, Y. Song, X. Wan, and G. Li, “Bridging vision and language encoders: Parameter-efficient tuning for referring image segmentation,” in IEEE Int. Conf. Comput. Vis. , 2023, pp. 17 503–17 512
2023
Closest in time.
A. Chronopoulou, M. Peters, A. Fraser, and J. Dodge, “AdapterSoup: Weight averaging to improve generalization of pretrained language models,” in Proc. Findings Assoc. Comput. Linguistics , 2023, pp. 2054–2063
2063
Closest in time.