Fetching the paper…
Reading the bibliography…
Parameter-efficient fine-tuning optimizes large, pre-trained foundation models by updating a subset of parameters; in this class, Low-Rank Adaptation (LoRA) is particularly effective.
Matrix correlation
Ramsay, J., ten Berge, J., and Styan, G · 1984
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
A thorough examination of the cnn/daily mail reading comprehension task
Chen, D., Bolton, J., and Manning, C. D · 2016
Earlier work this paper cites.
Information-theoretic analysis of generalization capability of learning algorithms
Xu, A. and Raginsky, M · 2017
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks, 2018
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Narayan, S., Cohen, S. B., and Lapata, M · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Earlier work this paper cites.
Similarity of neural network representations revisited
Kornblith, S., Norouzi, M., Lee, H., and Hinton, G · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2020
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2020
Earlier work this paper cites.
In search of lost domain generalization, 2020
Gulrajani, I. and Lopez-Paz, D · 2020
Cited alongside, same era.
Measuring massive multitask language understanding, 2020
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Cited alongside, same era.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2020
Cited alongside, same era.
What’s hidden in a randomly weighted neural network?
Ramanujan, V., Wortsman, M., Kembhavi, A., Farhadi, A., and Rastegari, M · 2020
Cited alongside, same era.
High-dimensional probability
Vershynin, R · 2020
Cited alongside, same era.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Aghajanyan, A., Gupta, S., and Zettlemoyer, L · 2021
Svdiff: Compact parameter space for diffusion fine-tuning
Han, L., Li, Y., Zhang, H., Milanfar, P., Metaxas, D., and Yang, F · 2023
Later among the works it cites.
Nola: Networks as linear combination of low rank random basis, 2023
Koohpayegani, S. A., Navaneet, K., Nooralinejad, P., Kolouri, S., and Pirsiavash, H · 2023
Later among the works it cites.
Scaling down to scale up: A guide to parameter-efficient fine-tuning
Lialin, V., Deshpande, V., and Rumshisky, A · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A survey of quantization methods for efficient neural network inference, 2021
Gholami, A., Kim, S., Dong, Z., Yao, Z., Mahoney, M. W., and Keutzer, K · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, J. E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., and Chen, W · 2021
Cited alongside, same era.
Fine-tuning can distort pretrained features and underperform out-of-distribution, 2022
Kumar, A., Raghunathan, A., Jones, R., Ma, T., and Liang, P · 2022
Cited alongside, same era.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C. A · 2022
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Cited alongside, same era.
Lora-fa: Memory-efficient low-rank adaptation for large language models fine-tuning, 2023a
Zhang, L., Zhang, L., Shi, S., Chu, X., and Li, B
Cited in the paper.
Later among the works it cites.
How far can camels go? exploring the state of instruction tuning on open resources
Wang, Y., Ivison, H., Dasigi, P., Hessel, J., Khot, T., Chandu, K. R., Wadden, D., MacMillan, K., Smith, N. A., Beltagy, I., et al · 2023
Later among the works it cites.
Yadav, P., Choshen, L., Raffel, C., and Bansal, M · 2023
Later among the works it cites.
The expressive power of low-rank adaptation, 2023
Zeng, Y. and Lee, K · 2023
Later among the works it cites.
Prilora: Pruned and rank-increasing low-rank adaptation
Benedek, N. and Wolf, L · 2024
Closest in time.
Lq-lora: Low-rank plus quantized matrix decomposition for efficient language model finetuning, 2024
Guo, H., Greengard, P., Xing, E. P., and Kim, Y · 2024
Closest in time.
Vera: Vector-based random matrix adaptation, 2024
Kopiczko, D. J., Blankevoort, T., and Asano, Y. M · 2024
Closest in time.
Parameter-efficient orthogonal finetuning via butterfly factorization
Liu, W., Qiu, Z., Feng, Y., Xiu, Y., Xue, Y., Yu, L., Feng, H., Liu, Z., Heo, J., Peng, S., Wen, Y., Black, M. J., Weller, A., and Schölkopf, B · 2024
Closest in time.