Fetching the paper…
Reading the bibliography…
Fine-tuning is the primary methodology for tailoring pre-trained large language models to specific tasks.
An algorithm for quadratic programming
Frank, M., Wolfe, P., et al · 1956
Earlier work this paper cites.
Sparse approximate solutions to semidefinite programs
Hazan, E · 2008
Earlier work this paper cites.
Projection-free online learning
Hazan, E. and Kale, S · 2012
Earlier work this paper cites.
Revisiting frank-wolfe: Projection-free sparse convex optimization
Jaggi, M · 2013
Earlier work this paper cites.
A linearly convergent variant of the conditional gradient algorithm under strong convexity, with applications to online and stochastic optimization
Garber, D. and Hazan, E · 2016
Earlier work this paper cites.
Convergence rate of frank-wolfe for non-convex objectives
Lacoste-Julien, S · 2016
Earlier work this paper cites.
Stochastic frank-wolfe methods for nonconvex optimization
Reddi, S. J., Sra, S., Póczos, B., and Smola, A · 2016
Earlier work this paper cites.
Linear convergence of a frank-wolfe type algorithm over trace-norm balls
Allen-Zhu, Z., Hazan, E., Hu, W., and Li, Y · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Simple, scalable adaptation for neural machine translation
Bapna, A. and Firat, O · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
BERT and PALs: Projected attention layers for efficient adaptation in multi-task learning
Stickland, A. C. and Murray, I · 2019
Cited alongside, same era.
Adapterfusion: Non-destructive task composition for transfer learning
Pfeiffer, J., Kamath, A., Rücklé, A., Cho, K., and Gurevych, I · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2020
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Later among the works it cites.
Longlora: Efficient fine-tuning of long-context large language models
Chen, Y., Qian, S., Tang, H., Lai, X., Liu, Z., Han, S., and Jia, J · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Later among the works it cites.
Fine-tuning language models with just forward passes
Malladi, S., Gao, T., Nichani, E., Damian, A., Lee, J. D., Chen, D., and Arora, S · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards a unified view of parameter-efficient transfer learning
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., and Neubig, G · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Cited alongside, same era.
Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks
Mahabadi, R. K., Ruder, S., Dehghani, M., and Henderson, J · 2021
Cited alongside, same era.
Wang, Y., Wang, W., Joty, S., and Hoi, S. C · 2021
Cited alongside, same era.
Qin, C., Xia, W., Jiao, F., and Joty, S · 2023
Later among the works it cites.
Tied-lora: Enhacing parameter efficiency of lora with weight tying
Renduchintala, A., Konuk, T., and Kuchaiev, O · 2023
Later among the works it cites.
S-lora: Serving thousands of concurrent lora adapters
Sheng, Y., Cao, S., Li, D., Hooper, C., Lee, N., Yang, S., Chou, C., Zhu, B., Zheng, L., Keutzer, K., et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Multilora: Democratizing lora for better multi-task learning
Wang, Y., Lin, Y., Zeng, X., and Zhang, G · 2023
Later among the works it cites.
Adaptive budget allocation for parameter-efficient fine-tuning
Zhang, Q., Chen, M., Bukharin, A., He, P., Cheng, Y., Chen, W., and Zhao, T · 2023
Later among the works it cites.