Fetching the paper…
Reading the bibliography…
With the increasingly powerful performances and enormous scales of pretrained models, promoting parameter efficiency in fine-tuning has become a crucial need for effective and efficient adaptation to various downstream tasks.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C. A · 1965
Earlier work this paper cites.
Solution of sparse linear least squares problems using givens rotations
George, A. and Heath, M. T · 1980
Earlier work this paper cites.
Fast givens rotations for orthogonal similarity transformations
Rath, W · 1982
Earlier work this paper cites.
Numerical recipes 3rd edition: The art of scientific computing
Press, W. H · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S · 2018
Earlier work this paper cites.
Learning towards minimum hyperspherical energy
Liu, W., Lin, R., Liu, Z., Liu, L., Yu, Z., Dai, B., and Song, L · 2018
Earlier work this paper cites.
Approximating orthogonal matrices with effective givens factorization
Frerix, T. and Bruna, J · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Earlier work this paper cites.
A large-scale study of representation learning with the visual task adaptation benchmark
Zhai, X., Puigcerver, J., Kolesnikov, A., Ruyssen, P., Riquelme, C., Lucic, M., Djolonga, J., Pinto, A. S., Neumann, M., Dosovitskiy, A., et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Adapterfusion: Non-destructive task composition for transfer learning
Pfeiffer, J., Kamath, A., Rücklé, A., Cho, K., and Gurevych, I · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Aghajanyan, A., Gupta, S., and Zettlemoyer, L · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2021
Cited alongside, same era.
He, P., Gao, J., and Chen, W · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Zaken, E. B., Goldberg, Y., and Ravfogel, S · 2022
Later among the works it cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Later among the works it cites.
Parameter-efficient fine-tuning design spaces
Chen, J., Zhang, A., Shi, X., Li, M., Smola, A., and Yang, D · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Cited alongside, same era.
Pixelated butterfly: Simple and efficient sparse training for neural network models
Chen, B., Dao, T., Liang, K., Yang, J., Song, Z., Rudra, A., and Re, C · 2022
Cited alongside, same era.
Krona: Parameter efficient tuning with kronecker adapter
Edalati, A., Tahaei, M., Kobyzev, I., Nia, V. P., Clark, J. J., and Rezagholizadeh, M · 2022
Cited alongside, same era.
Sparseadapter: An easy approach for improving the parameter-efficiency of adapters
He, S., Ding, L., Dong, D., Zhang, J., and Tao, D · 2022
Cited alongside, same era.
Domain generalization through the lens of angular invariance
Jin, Y., Chu, X., Wang, Y., and Zhu, W · 2022
Cited alongside, same era.
Patient health representation learning via correlational sparse prior of medical features
Ma, X., Wang, Y., Chu, X., Ma, L., Tang, W., Zhao, J., Yuan, Y., and Wang, G · 2022
Cited alongside, same era.
Later among the works it cites.
Llama factory
hiyouga · 2023
Later among the works it cites.
Fact: Factor-tuning for lightweight adaptation on vision transformer
Jie, S. and Deng, Z.-H · 2023
Later among the works it cites.
Distributional correlation–aware knowledge distillation for stock trading volume prediction
Li, L., Zhang, Z., Bao, R., Harimoto, K., and Sun, X · 2023
Later among the works it cites.
Scaling down to scale up: A guide to parameter-efficient fine-tuning
Lialin, V., Deshpande, V., and Rumshisky, A · 2023
Later among the works it cites.
Controlling text-to-image diffusion by orthogonal finetuning
Qiu, Z., Liu, W., Feng, H., Xue, Y., Feng, Y., Liu, Z., Zhang, D., Weller, A., and Schölkopf, B · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Seqcare: Sequential training with external medical knowledge graph for diagnosis prediction in healthcare data
Xu, Y., Chu, X., Yang, K., Wang, Z., Zou, P., Ding, H., Zhao, J., Wang, Y., and Xie, B · 2023
Later among the works it cites.
Adaptive budget allocation for parameter-efficient fine-tuning
Zhang, Q., Chen, M., Bukharin, A., He, P., Cheng, Y., Chen, W., and Zhao, T · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.
Lora dropout as a sparsity regularizer for overfitting control
Lin, Y., Ma, X., Chu, X., Jin, Y., Yang, Z., Wang, Y., and Mei, H · 2024
Closest in time.