Fetching the paper…
Reading the bibliography…
We investigate parameter-efficient fine-tuning (PEFT) methods that can provide good accuracy under limited computational and memory budgets in the context of large language models (LLMs).
Robust estimates, residuals, and outlier detection with multiresponse data
Gnanadesikan, R. and Kettenring, J. R · 1972
Earlier work this paper cites.
Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography
Fischler, M. A. and Bolles, R. C · 1981
Earlier work this paper cites.
A framework for robust subspace learning
De La Torre, F. and Black, M. J · 2003
Earlier work this paper cites.
Robust statistics , volume 523
Huber, P. J · 2004
Earlier work this paper cites.
Robust l/sub 1/norm factorization in the presence of outliers and missing data by alternative convex programming
Ke, Q. and Kanade, T · 2005
Earlier work this paper cites.
Robust principal component analysis: Exact recovery of corrupted low-rank matrices via convex optimization
Wright, J., Ganesh, A., Rao, S., Peng, Y., and Ma, Y · 2009
Earlier work this paper cites.
Robust principal component analysis?
Candès, E. J., Li, X., Ma, Y., and Wright, J · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., et al · 2012
Earlier work this paper cites.
Gpu kernels for block-sparse weights
Gray, S., Radford, A., and Kingma, D. P · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Seq2sql: Generating structured queries from natural language using reinforcement learning
Zhong, V., Xiong, C., and Socher, R · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S., et al · 2018
Earlier work this paper cites.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S · 2019
Earlier work this paper cites.
ViGGO: A video game corpus for data-to-text generation in open-domain conversation
Juraska, J., Bowden, K., and Walker, M · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Earlier work this paper cites.
Rigging the lottery: Making all tickets winners
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Earlier work this paper cites.
Sparse GPU kernels for deep learning
Gale, T., Zaharia, M., Young, C., and Elsen, E · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V., Wolf, T., and Rush, A · 2020
Earlier work this paper cites.
Woodfisher: Efficient second-order approximation for neural network compression
Singh, S. P. and Alistarh, D · 2020
Earlier work this paper cites.
Raft: A real-world few-shot text classification benchmark
Alex, N., Lifland, E., Tunstall, L., Thakur, A., Maham, P., Riedel, C. J., Hine, E., Ashurst, C., Sedille, P., Carlier, A., et al · 2021
Cited alongside, same era.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Cited alongside, same era.
Dsee: Dually sparsity-embedded efficient tuning of pre-trained language models
Chen, X., Chen, T., Chen, W., Awadallah, A. H., Wang, Z., and Cheng, Y · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Cited alongside, same era.
Sten: An interface for efficient sparsity in pytorch
Ivanov, A., Dryden, N., and Hoefler, T · 2022
Later among the works it cites.
Exposing and exploiting fine-grained block structures for fast and accurate sparse training
Jiang, P., Hu, L., and Song, S · 2022
Later among the works it cites.
The optimal bert surgeon: Scalable and accurate second-order pruning for large language models
Kurtic, E., Campos, D., Nguyen, T., Frantar, E., Kurtz, M., Fineran, B., Goin, M., and Alistarh, D · 2022
Later among the works it cites.
Efficient quantized sparse matrix operations on tensor cores
Li, S., Osawa, K., and Hoefler, T · 2022
Later among the works it cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C. A · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks
Hoefler, T., Alistarh, D., Ben-Nun, T., Dryden, N., and Peste, A · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Accelerated sparse neural training: A provable and efficient method to find n: m transposable masks
Hubara, I., Chmiel, B., Island, M., Banner, R., Naor, J., and Soudry, D · 2021
Cited alongside, same era.
Fedpara: Low-rank hadamard product for communication-efficient federated learning
Hyeon-Woo, N., Ye-Bin, M., and Oh, T.-H · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Cited alongside, same era.
P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks
Liu, X., Ji, K., Fu, Y., Tam, W. L., Du, Z., Yang, Z., and Tang, J · 2021
Cited alongside, same era.
Metaicl: Learning to learn in context
Min, S., Lewis, M., Zettlemoyer, L., and Hajishirzi, H · 2021
Cited alongside, same era.
Peft: State-of-the-art parameter-efficient fine-tuning methods
Mangrulkar, S., Gugger, S., Debut, L., Belkada, Y., Paul, S., and Bossan, B · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks
Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Naik, A., Ashok, A., Dhanasekaran, A. S., Arunkumar, A., Stap, D., et al · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.
Greedy bilateral sketch, completion & smoothing
Zhou, T. and Tao, D · 2022
Later among the works it cites.
Venom: A vectorized n: M format for unleashing the power of sparse tensor cores
Castro, R. L., Ivanov, A., Andrade, D., Ben-Nun, T., Fraguela, B. B., and Hoefler, T · 2023
Later among the works it cites.
Sparse finetuning for inference acceleration of large language models
Kurtic, E., Kuznedelev, D., Frantar, E., Goin, M., and Alistarh, D · 2023
Later among the works it cites.
Platypus: Quick, cheap, and powerful refinement of llms
Lee, A. N., Hunter, C. J., and Ruiz, N · 2023
Later among the works it cites.
Loftq: Lora-fine-tuning-aware quantization for large language models
Li, Y., Yu, Y., Liang, C., He, P., Karampatziakis, N., Chen, W., and Zhao, T · 2023
Later among the works it cites.
Gpt understands, too
Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., and Tang, J · 2023
Later among the works it cites.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023b
MosaicML · 2023
Later among the works it cites.
Fine-Tuning LLMs: LoRA or Full-Parameter?, 2023
Niederfahrenhorst, A., Hakhamaneshi, K., and Ahmad, R · 2023
Later among the works it cites.
Sparseprop: Efficient sparse backpropagation for faster training of neural networks at the edge
Nikdan, M., Pegolotti, T., Iofinova, E., Kurtic, E., and Alistarh, D · 2023
Later among the works it cites.
Controlling text-to-image diffusion by orthogonal finetuning
Qiu, Z., Liu, W., Feng, H., Xue, Y., Feng, Y., Liu, Z., Zhang, D., Weller, A., and Schölkopf, B · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Adaptive budget allocation for parameter-efficient fine-tuning
Zhang, Q., Chen, M., Bukharin, A., He, P., Cheng, Y., Chen, W., and Zhao, T · 2023
Later among the works it cites.