Fetching the paper…
Reading the bibliography…
Parameter-efficient fine-tuning methods, represented by LoRA, play an essential role in adapting large-scale pre-trained models to downstream tasks.
Stability and generalization
Bousquet, O. and Elisseeff, A · 2002
Earlier work this paper cites.
Stability of randomized learning algorithms
Elisseeff, A., Evgeniou, T., Pontil, M., and Kaelbing, L. P · 2005
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. R · 2012
Earlier work this paper cites.
Understanding dropout
Baldi, P. and Sadowski, P. J · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
Wan, L., Zeiler, M., Zhang, S., Le Cun, Y., and Fergus, R · 2013
Earlier work this paper cites.
An empirical analysis of dropout in piecewise linear networks
Warde-Farley, D., Goodfellow, I. J., Courville, A., and Bengio, Y · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Efficient object localization using convolutional networks
Tompson, J., Goroshin, R., Jain, A., LeCun, Y., and Bregler, C · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Recurrent dropout without memory loss
Semeniuta, S., Severyn, A., and Barth, E · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Stability and generalization of learning algorithms that converge to global optima
Charles, Z. and Papailiopoulos, D · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S · 2018
Earlier work this paper cites.
Data-dependent stability of stochastic gradient descent
Kuzborskij, I. and Lampert, C · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Rajpurkar, P., Jia, R., and Liang, P · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2019
Cited alongside, same era.
Deberta: Decoding-enhanced bert with disentangled attention
Peft: State-of-the-art parameter-efficient fine-tuning methods
Mangrulkar, S., Gugger, S., Debut, L., Belkada, Y., Paul, S., and Bossan, B · 2022
Later among the works it cites.
Uncertainty quantification with pre-trained language models: A large-scale empirical analysis
Xiao, Y., Liang, P. P., Bhatt, U., Neiswanger, W., Salakhutdinov, R., and Morency, L.-P · 2022
Later among the works it cites.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
Zaken, E. B., Goldberg, Y., and Ravfogel, S · 2022
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Later among the works it cites.
On the effectiveness of parameter-efficient fine-tuning
Fu, Z., Yang, H., So, A. M.-C., Lam, W., Bing, L., and Collier, N · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
He, P., Liu, X., Gao, J., and Chen, W · 2020
Cited alongside, same era.
Being bayesian, even just a bit, fixes overconfidence in relu networks
Kristiadi, A., Hein, M., and Hennig, P · 2020
Cited alongside, same era.
Adapterfusion: Non-destructive task composition for transfer learning
Pfeiffer, J., Kamath, A., Rücklé, A., Cho, K., and Gurevych, I · 2020
Cited alongside, same era.
He, P., Gao, J., and Chen, W · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
How can we know when language models know? on the calibration of language models for question answering
Jiang, Z., Araki, J., Ding, H., and Neubig, G · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Cited alongside, same era.
Preserving pre-trained features helps calibrate fine-tuned language models
He, G., Chen, J., and Zhu, J · 2023
Later among the works it cites.
Llama factory
hiyouga · 2023
Later among the works it cites.
Scaling down to scale up: A guide to parameter-efficient fine-tuning
Lialin, V., Deshpande, V., and Rumshisky, A · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., and Manning, C. D · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Xu, L., Xie, H., Qin, S.-Z. J., Tao, X., and Wang, F. L · 2023
Later among the works it cites.
Adaptive budget allocation for parameter-efficient fine-tuning
Zhang, Q., Chen, M., Bukharin, A., He, P., Cheng, Y., Chen, W., and Zhao, T · 2023
Later among the works it cites.
On the calibration of large language models and alignment
Zhu, C., Xu, B., Wang, Q., Zhang, Y., and Mao, Z · 2023
Later among the works it cites.
Delta-lora: Fine-tuning high-rank parameters with the delta of low-rank matrices
Zi, B., Qi, X., Wang, L., Wang, J., Wong, K.-F., and Zhang, L · 2023
Later among the works it cites.
Parameter efficient quasi-orthogonal fine-tuning via givens rotation
Ma, X., Chu, X., Yang, Z., Lin, Y., Gao, X., and Zhao, J · 2024
Closest in time.