Fetching the paper…
Reading the bibliography…
Current parameter-efficient fine-tuning (PEFT) methods build adapters widely agnostic of the context of downstream task to learn, or the context of important knowledge to maintain.
Building a large annotated corpus of english: The penn treebank
M. Marcus, B. Santorini, and M. A. Marcinkiewicz · 1993
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
J. Berant, A. Chou, R. Frostig, and P. Liang · 2013
Earlier work this paper cites.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
I. J. Goodfellow, M. Mirza, D. Xiao, A. Courville, and Y. Bengio · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Learning without forgetting
Z. Li and D. Hoiem · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
D. Lopez-Paz and M. Ranzato · 2017
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
S.-A. Rebuffi, A. Kolesnikov, G. Sperl, and C. H. Lampert · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Measuring the intrinsic dimension of objective landscapes
C. Li, H. Farkhoor, R. Liu, and J. Yosinski · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Comet: Commonsense transformers for automatic knowledge graph construction
A. Bosselut, H. Rashkin, M. Sap, C. Malaviya, A. Celikyilmaz, and Y. Choi · 2019
Earlier work this paper cites.
Learning a unified classifier incrementally via rebalancing
S. Hou, X. Pan, C. C. Loy, Z. Wang, and D. Lin · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, et al · 2019
Earlier work this paper cites.
Latent retrieval for weakly supervised open domain question answering
K. Lee, M.-W. Chang, and K. Toutanova · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Earlier work this paper cites.
Continual unsupervised representation learning
D. Rao, F. Visin, A. Rusu, R. Pascanu, Y. W. Teh, and R. Hadsell · 2019
Earlier work this paper cites.
Learning to learn without forgetting by maximizing transfer and minimizing interference
M. Riemer, I. Cases, R. Ajemian, M. Liu, I. Rish, Y. Tu, , and G. Tesauro · 2019
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
A. Aghajanyan, L. Zettlemoyer, and S. Gupta · 2020
Earlier work this paper cites.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Compacter: Efficient low-rank hypercomplex adapter layers
R. Karimi Mahabadi, J. Henderson, and S. Ruder · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
B. Lester, R. Al-Rfou, and N. Constant · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
X. L. Li and P. Liang · 2021
Cited alongside, same era.
Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks
R. K. Mahabadi, S. Ruder, M. Dehghani, and J. Henderson · 2021
Cited alongside, same era.
Adapterfusion: Non-destructive task composition for transfer learning
J. Pfeiffer, A. Kamath, A. Rücklé, K. Cho, and I. Gurevych · 2021
Cited alongside, same era.
Der: Dynamically expandable representation for class incremental learning
S. Yan, J. Xie, and X. He · 2021
Cited alongside, same era.
Towards a unified view of parameter-efficient transfer learning
J. He, C. Zhou, X. Ma, T. Berg-Kirkpatrick, and G. Neubig · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Asvd: Activation-aware singular value decomposition for compressing large language models
Z. Yuan, Y. Shang, Y. Song, Q. Wu, Y. Yan, and G. Sun · 2023
Later among the works it cites.
Increlora: Incremental parameter allocation method for parameter-efficient fine-tuning
F. Zhang, L. Li, J. Chen, Z. Jiang, B. Wang, and Y. Qian · 2023
Later among the works it cites.
Composing parameter-efficient modules with arithmetic operation
J. Zhang, J. Liu, J. He, et al · 2023
Later among the works it cites.
Pruning meets low-rank parameter-efficient fine-tuning
M. Zhang, C. Shen, Z. Yang, L. Ou, X. Yu, B. Zhuang, et al · 2023
Later among the works it cites.
Adaptive budget allocation for parameter-efficient fine-tuning
Q. Zhang, M. Chen, A. Bukharin, P. He, Y. Cheng, W. Chen, and T. Zhao · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fine-tuned language models are continual learners
T. Scialom, T. Chakrabarty, and S. Muresan · 2022
Cited alongside, same era.
M. Valipour, M. Rezagholizadeh, I. Kobyzev, and A. Ghodsi · 2022
Cited alongside, same era.
Towards theoretically inspired neural initialization optimization
Y. Yang, H. Wang, H. Yuan, and Z. Lin · 2022
Cited alongside, same era.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Cited alongside, same era.
Renaissance: A survey into ai text-to-image generation in the era of large model
F. Bie, Y. Yang, Z. Zhou, A. Ghanem, M. Zhang, Z. Yao, X. Wu, C. Holmes, P. Golnari, D. A. Clifton, et al · 2023
Cited alongside, same era.
One-for-all: Generalized lora for parameter-efficient fine-tuning
A. Chavan, Z. Liu, D. Gupta, E. Xing, and Z. Shen · 2023
Cited alongside, same era.
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2023
Later among the works it cites.
Judging LLM-as-a-judge with MT-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Later among the works it cites.
SPT: Learning to selectively insert prompts for better prompt tuning
W. Zhu and M. Tan · 2023
Later among the works it cites.
Simple and scalable strategies to continually pre-train large language models
A. Ibrahim, B. Thérien, K. Gupta, M. L. Richter, Q. Anthony, T. Lesort, E. Belilovsky, and I. Rish · 2024
Closest in time.
Vera: Vector-based random matrix adaptation
D. J. Kopiczko, T. Blankevoort, and Y. M. Asano · 2024
Closest in time.
Owq: Outlier-aware weight quantization for efficient fine-tuning and inference of large language models
C. Lee, J. Jin, T. Kim, H. Kim, and E. Park · 2024
Closest in time.
Genview: Enhancing view quality with pretrained generative model for self-supervised learning
X. Li, Y. Yang, X. Li, J. Wu, Y. Yu, B. Ghanem, and M. Zhang · 2024
Closest in time.
X. Li, Y. Yang, J. Wu, B. Ghanem, L. Nie, and M. Zhang · 2024
Closest in time.
Loftq: LoRA-fine-tuning-aware quantization for large language models
Y. Li, Y. Yu, C. Liang, N. Karampatziakis, P. He, W. Chen, and T. Zhao · 2024
Closest in time.
Awq: Activation-aware weight quantization for llm compression and acceleration
J. Lin, J. Tang, H. Tang, S. Yang, W.-M. Chen, W.-C. Wang, G. Xiao, X. Dang, C. Gan, and S. Han · 2024
Closest in time.
Dora: Weight-decomposed low-rank adaptation
S.-Y. Liu, C.-Y. Wang, H. Yin, P. Molchanov, Y.-C. F. Wang, K.-T. Cheng, and M.-H. Chen · 2024
Closest in time.
Pissa: Principal singular values and singular vectors adaptation of large language models
F. Meng, Z. Wang, and M. Zhang · 2024
Closest in time.
Llama pro: Progressive llama with block expansion
C. Wu, Y. Gan, Y. Ge, Z. Lu, J. Wang, Y. Feng, P. Luo, and Y. Shan · 2024
Closest in time.
Continual learning for large language models: A survey
T. Wu, L. Luo, Y.-F. Li, S. Pan, T.-T. Vu, and G. Haffari · 2024
Closest in time.
WizardLM: Empowering large pre-trained language models to follow complex instructions
C. Xu, Q. Sun, K. Zheng, X. Geng, P. Zhao, J. Feng, C. Tao, Q. Lin, and D. Jiang · 2024
Closest in time.
QA-loRA: Quantization-aware low-rank adaptation of large language models
Y. Xu, L. Xie, X. Gu, X. Chen, H. Chang, H. Zhang, Z. Chen, X. ZHANG, and Q. Tian · 2024
Closest in time.
Towards interpretable deep local learning with successive gradient reconciliation
Y. Yang, X. Li, M. Alfarra, H. A. A. K. Hammoud, A. Bibi, P. Torr, and B. Ghanem · 2024
Closest in time.
Metamath: Bootstrap your own mathematical questions for large language models
L. Yu, W. Jiang, H. Shi, J. YU, Z. Liu, Y. Zhang, J. Kwok, Z. Li, A. Weller, and W. Liu · 2024
Closest in time.
Language models are super mario: Absorbing abilities from homologous models as a free lunch
L. Yu, B. Yu, H. Yu, F. Huang, and Y. Li · 2024
Closest in time.
Investigating the catastrophic forgetting in multimodal large language model fine-tuning
Y. Zhai, S. Tong, X. Li, M. Cai, Q. Qu, Y. J. Lee, and Y. Ma · 2024
Closest in time.
Galore: Memory-efficient llm training by gradient low-rank projection
J. Zhao, Z. Zhang, B. Chen, Z. Wang, A. Anandkumar, and Y. Tian · 2024
Closest in time.
Opencodeinterpreter: Integrating code generation with execution and refinement
T. Zheng, G. Zhang, T. Shen, X. Liu, B. Y. Lin, J. Fu, W. Chen, and X. Yue · 2024
Closest in time.
Model tailor: Mitigating catastrophic forgetting in multi-modal large language models
D. Zhu, Z. Sun, Z. Li, T. Shen, K. Yan, S. Ding, K. Kuang, and C. Wu · 2024
Closest in time.