Fetching the paper…
Reading the bibliography…
Parameter-efficient tuning (PEFT) techniques like low-rank adaptation (LoRA) offer training efficiency on Large Language Models, but their impact on model performance remains limited.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. and Brockett, C · 2005
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Riemannian walk for incremental learning: Understanding forgetting and intransigence
Chaudhry, A., Dokania, P. K., Ajanthan, T., and Torr, P. H · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Talmor, A., Herzig, J., Lourie, N., and Berant, J · 2019
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2019
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. and Liang, P · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., et al · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E., Shen, Y., Wallis, P., Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks
Liu, X., Ji, K., Fu, Y., Tam, W., Du, Z., Yang, Z., and Tang, J · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-w., Zhu, S.-C., Tafjord, O., Clark, P., and Kalyan, A · 2022
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2023
Later among the works it cites.
Dou, S., Zhou, E., Liu, Y., Gao, S., Zhao, J., Shen, W., Zhou, Y., Xi, Z., Wang, X., Fan, X., Pu, S., Zhu, J., Zheng, R., Gui, T., Zhang, Q., and Huang, X · 2023
Later among the works it cites.
Openorca: An open dataset of gpt augmented flan reasoning traces
Lian, W., Goodson, B., Pentland, E., Cook, A., Vong, C., and ”Teknium” · 2023
Later among the works it cites.
Moelora: An moe-based parameter efficient fine-tuning method for multi-task medical applications
Liu, Q., Wu, X., Zhao, X., Zhu, Y., Xu, D., Tian, F., and Zheng, Y · 2023
Later among the works it cites.
Full parameter fine-tuning for large language models with limited resources
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cross-task generalization via natural language crowdsourcing instructions
Mishra, S., Khashabi, D., Baral, C., and Hajishirzi, H · 2022
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Le Scao, T., Raja, A., et al · 2022
Cited alongside, same era.
Large language models encode clinical knowledge
Singhal, K., Azizi, S., Tu, T., Mahdavi, S., Wei, J., Chung, H., Scales, N., Tanwani, A., et al · 2022
Cited alongside, same era.
St-moe: Designing stable and transferable sparse expert models
Zoph, B., Bello, I., Kumar, S., Du, N., Huang, Y., Dean, J., Shazeer, N., and Fedus, W · 2022
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A., Hauth, A., et al · 2023
Cited alongside, same era.
Sparse moe as the new dropout: Scaling dense and self-slimmable transformer
Chen, T., Zhang, Z., Jaiwal, A., Liu, S., and Wang, Z · 2023
Cited alongside, same era.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C
Cited in the paper.
Lv, K., Yang, Y., Liu, T., Gao, Q., Guo, Q., and Qiu, X · 2023
Later among the works it cites.
Orca: Progressive learning from complex explanation traces of gpt-4, 2023
Mukherjee, S., Mitra, A., Jawahar, G., Agarwal, S., Palangi, H., and Awadallah, A · 2023
Later among the works it cites.
What do llamas really think? revealing preference biases in language model representations
Tang, R., Zhang, X., Lin, J., and Ture, F · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Pushing mixture of experts to the limit: extremely parameter efficient moe for instruction tuning
Zadouri, T., Ustun, A., Ahmadian, A., Ermis, B., Locatelli, A., and Hooker, S · 2023
Later among the works it cites.
Judgelm: Fine-tuned large language models are scalable judges
Zhu, L., Wang, X., and Wang, X · 2023
Later among the works it cites.
Jiang, A., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D., Casas, D. d. l., et al · 2024
Closest in time.