Fetching the paper…
Reading the bibliography…
Fine-tuning large language models (LLM) can be costly.
The pascal recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2005
Earlier work this paper cites.
The mnist database of handwritten digits
LeCun, Y. and Cortes, C · 2005
Earlier work this paper cites.
The second pascal recognising textual entailment challenge
Haim, R. B., Dagan, I., Dolan, B., Ferro, L., Giampiccolo, D., Magnini, B., and Szpektor, I · 2006
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
Giampiccolo, D., Magnini, B., Dagan, I., and Dolan, W. B · 2007
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Bentivogli, L., Clark, P., Dagan, I., and Giampiccolo, D · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., Ng, A. Y., et al · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Roemmele, M., Bejan, C. A., and Gordon, A. S · 2011
Earlier work this paper cites.
The german traffic sign recognition benchmark: a multi-class classification competition
Stallkamp, J., Schlipsing, M., Salmen, J., and Igel, C · 2011
Earlier work this paper cites.
The winograd schema challenge
Levesque, H., Davis, E., and Morgenstern, L · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Bubeck, S. et al · 2015
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Remote sensing image scene classification: Benchmark and state of the art
Cheng, G., Han, J., and Lu, X · 2017
Earlier work this paper cites.
Learning efficient convolutional networks through network slimming
Liu, Z., Li, J., Shen, Z., Huang, G., Yan, S., and Zhang, C · 2017
Earlier work this paper cites.
Variational dropout sparsifies deep neural networks
Molchanov, D., Ashukha, A., and Vetrov, D · 2017
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Khashabi, D., Chaturvedi, S., Roth, M., Upadhyay, S., and Roth, D · 2018
Earlier work this paper cites.
Wic: the word-in-context dataset for evaluating context-sensitive meaning representations
Pilehvar, M. T. and Camacho-Collados, J · 2018
Earlier work this paper cites.
Record: Bridging the gap between human and machine commonsense reading comprehension
Zhang, S., Liu, X., Liu, J., Gao, J., Duh, K., and Van Durme, B · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
The commitmentbank: Investigating projection in naturally occurring discourse
De Marneffe, M.-C., Simons, M., and Tonhauser, J · 2019
Earlier work this paper cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M · 2019
Earlier work this paper cites.
Visualizing and understanding the effectiveness of bert
Hao, Y., Dong, L., Wei, F., and Xu, K · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
High-dimensional statistics: A non-asymptotic viewpoint , volume 48
Wainwright, M. J · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2019
Cited alongside, same era.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Aghajanyan, A., Zettlemoyer, L., and Gupta, S · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Sparse GPU kernels for deep learning
Gale, T., Zaharia, M., Young, C., and Elsen, E · 2020
Cited alongside, same era.
Convolutional bypasses are better vision transformer adapters
Jie, S. and Deng, Z.-H · 2022
Later among the works it cites.
Parameter-efficient sparsity for large language models fine-tuning
Li, Y., Luo, F., Tan, C., Wang, M., Huang, S., Li, S., and Bai, J · 2022
Later among the works it cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C. A · 2022
Later among the works it cites.
Vl-adapter: Parameter-efficient transfer learning for vision-and-language tasks
Sung, Y.-L., Cho, J., and Bansal, M · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guo, D., Rush, A. M., and Kim, Y · 2020
Cited alongside, same era.
Adapterfusion: Non-destructive task composition for transfer learning
Pfeiffer, J., Kamath, A., Rücklé, A., Cho, K., and Gurevych, I · 2020
Cited alongside, same era.
Adapterdrop: On the efficiency of adapters in transformers
Rücklé, A., Geigle, G., Glockner, M., Beck, T., Pfeiffer, J., Reimers, N., and Gurevych, I · 2020
Cited alongside, same era.
Masking as an efficient alternative to finetuning for pretrained language models
Zhao, M., Lin, T., Mi, F., Jaggi, M., and Schütze, H · 2020
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Compacter: Efficient low-rank hypercomplex adapter layers
Karimi Mahabadi, R., Henderson, J., and Ruder, S · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Cited alongside, same era.
Later among the works it cites.
Parameter-efficient fine-tuning design spaces
Chen, J., Zhang, A., Shi, X., Li, M., Smola, A., and Yang, D · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Later among the works it cites.
Sparse low-rank adaptation of pre-trained language models
Ding, N., Lv, X., Wang, Q., Chen, Y., Zhou, B., Liu, Z., and Sun, M · 2023
Later among the works it cites.
Llama-adapter v2: Parameter-efficient visual instruction model
Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., et al · 2023
Later among the works it cites.
Nola: Networks as linear combination of low rank random basis
Koohpayegani, S. A., Navaneet, K., Nooralinejad, P., Kolouri, S., and Pirsiavash, H · 2023
Later among the works it cites.
Vera: Vector-based random matrix adaptation
Kopiczko, D. J., Blankevoort, T., and Asano, Y. M · 2023
Later among the works it cites.
Scaling down to scale up: A guide to parameter-efficient fine-tuning
Lialin, V., Deshpande, V., and Rumshisky, A · 2023
Later among the works it cites.
Vision transformers are parameter-efficient audio-visual learners
Lin, Y.-B., Sung, Y.-L., Lei, J., Bansal, M., and Bertasius, G · 2023
Later among the works it cites.
Gpt understands, too
Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., and Tang, J · 2023
Later among the works it cites.
Full parameter fine-tuning for large language models with limited resources
Lv, K., Yang, Y., Liu, T., Gao, Q., Guo, Q., and Qiu, X · 2023
Later among the works it cites.
Towards efficient fine-tuning of pre-trained code models: An experimental study and beyond
Shi, E., Wang, Y., Zhang, H., Du, L., Han, S., Zhang, D., and Sun, H · 2023
Later among the works it cites.
Exploring the impact of model scaling on parameter-efficient tuning
Su, Y., Chan, C.-M., Cheng, J., Qin, Y., Lin, Y., Hu, S., Yang, Z., Ding, N., Sun, X., Xie, G., et al · 2023
Later among the works it cites.
Topics in random matrix theory , volume 132
Tao, T · 2023
Later among the works it cites.
Qa-lora: Quantization-aware low-rank adaptation of large language models
Xu, Y., Xie, L., Gu, X., Chen, X., Chang, H., Zhang, H., Chen, Z., Zhang, X., and Tian, Q · 2023
Later among the works it cites.
Zelikman, E., Huang, Q., Liang, P., Haber, N., and Goodman, N. D · 2023
Later among the works it cites.
Autopeft: Automatic configuration search for parameter-efficient fine-tuning
Zhou, H., Wan, X., Vulić, I., and Korhonen, A · 2023
Later among the works it cites.
Delta-lora: Fine-tuning high-rank parameters with the delta of low-rank matrices
Zi, B., Qi, X., Wang, L., Wang, J., Wong, K.-F., and Zhang, L · 2023
Later among the works it cites.
Rosa: Accurate parameter-efficient fine-tuning via robust adaptation
Nikdan, M., Tabesh, S., Crnčević, E., and Alistarh, D · 2024
Closest in time.
Task arithmetic in the tangent space: Improved editing of pre-trained models
Ortiz-Jimenez, G., Favero, A., and Frossard, P · 2024
Closest in time.