Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have the ability to solve a variety of tasks, such as text summarization and mathematical questions, just out of the box, but they are often trained with a single task in mind.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Gururangan, S.; Marasović, A.; Swayamdipta, S.; Lo, K.; Beltagy, I.; Downey, D.; and Smith, N. A. 2020 · 2004
Earlier work this paper cites.
Learning under concept drift: an overview
Žliobaitė, I. 2010 · 2010
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G. S.; and Dean, J. 2013 · 2013
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, B.; Moore, E.; Ramage, D.; Hampson, S.; and y Arcas, B. A. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
AdapterSoup: Weight Averaging to Improve Generalization of Pretrained Language Models
Chronopoulou, A.; Peters, M. E.; Fraser, A.; and Dodge, J. 2023 · 2018
Earlier work this paper cites.
Generating Wikipedia by Summarizing Long Sequences
Liu, P. J.; Saleh, M.; Pot, E.; Goodrich, B.; Sepassi, R.; Kaiser, L.; and Shazeer, N. 2018 · 2018
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
SCAFFOLD: Stochastic controlled averaging for federated learning
Karimireddy, S. P.; Kale, S.; Mohri, M.; Reddi, S.; Stich, S.; and Suresh, A. T. 2020 · 2020
Earlier work this paper cites.
Federated optimization in heterogeneous networks
Li, T.; Sahu, A. K.; Zaheer, M.; Sanjabi, M.; Talwalkar, A.; and Smith, V. 2020 · 2020
Earlier work this paper cites.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
Bender, E. M.; Gebru, T.; McMillan-Major, A.; and Shmitchell, S. 2021 · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Bommasani, R.; Hudson, D. A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M. S.; Bohg, J.; Bosselut, A.; Brunskill, E.; et al. 2021 · 2021
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W.; et al. 2021 · 2021
Earlier work this paper cites.
Kenton, Z.; Everitt, T.; Weidinger, L.; Gabriel, I.; Mikulik, V.; and Irving, G. 2021 · 2021
Earlier work this paper cites.
I-BERT: Integer-only BERT quantization
Kim, S.; Gholami, A.; Yao, Z.; Mahoney, M. W.; and Keutzer, K. 2021 · 2021
Earlier work this paper cites.
The Power of Scale for Parameter-Efficient Prompt Tuning
Lester, B.; Al-Rfou, R.; and Constant, N. 2021 · 2021
Earlier work this paper cites.
Model-contrastive federated learning
Li, Q.; He, B.; and Song, D. 2021 · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li, X. L.; and Liang, P. 2021 · 2021
Earlier work this paper cites.
Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks
Mahabadi, R. K.; Ruder, S.; Dehghani, M.; and Henderson, J. 2021 · 2021
Earlier work this paper cites.
AdapterFusion: Non-Destructive Task Composition for Transfer Learning
Pfeiffer, J.; Kamath, A.; Rücklé, A.; Cho, K.; and Gurevych, I. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models
Tamkin, A.; Brundage, M.; Clark, J.; and Ganguli, D. 2021 · 2021
Cited alongside, same era.
Git Re-Basin: Merging Models modulo Permutation Symmetries
Ainsworth, S.; Hayase, J.; and Srinivasa, S. 2022 · 2022
Cited alongside, same era.
Attentional mixtures of soft prompt tuning for parameter-efficient multi-task knowledge sharing
Asai, A.; Salehi, M.; Peters, M. E.; and Hajishirzi, H. 2022 · 2022
Cited alongside, same era.
Parameter-efficient fine-tuning of large-scale pre-trained language models
Ding, N.; Qin, Y.; Yang, G.; Wei, F.; Yang, Z.; Su, Y.; Hu, S.; Chen, Y.; Chan, C.-M.; Chen, W.; et al. 2023 · 2023
Closest in time.
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Frantar, E.; and Alistarh, D. 2023 · 2023
Closest in time.
LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models
Hu, Z.; Lan, Y.; Wang, L.; Xu, W.; Lim, E.-P.; Lee, R. K.-W.; Bing, L.; and Poria, S. 2023 · 2023
Closest in time.
Language is not all you need: Aligning perception with language models
Huang, S.; Dong, L.; Wang, W.; Hao, Y.; Singhal, S.; Ma, S.; Lv, T.; Cui, L.; Mohammed, O. K.; Liu, Q.; et al. 2023 · 2023
Closest in time.
The Emergence of Essential Sparsity in Large Pre-trained Models: The Weights that Matter
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Dettmers, T.; Lewis, M.; Belkada, Y.; and Zettlemoyer, L. 2022 · 2022
Cited alongside, same era.
GPTQ: Accurate post-training quantization for generative pre-trained transformers
Frantar, E.; Ashkboos, S.; Hoefler, T.; and Alistarh, D. 2022 · 2022
Cited alongside, same era.
Dataless knowledge fusion by merging weights of language models
Jin, X.; Ren, X.; Preotiuc-Pietro, D.; and Cheng, P. 2022 · 2022
Cited alongside, same era.
PEFT: State-of-the-art Parameter-Efficient Fine-Tuning methods
Mangrulkar, S.; Gugger, S.; Debut, L.; Belkada, Y.; Paul, S.; and Bossan, B. 2022 · 2022
Cited alongside, same era.
Merging models with Fisher-weighted averaging
Matena, M. S.; and Raffel, C. A. 2022 · 2022
Cited alongside, same era.
Cross-task generalization via natural language crowdsourcing instructions
Mishra, S.; Khashabi, D.; Baral, C.; and Hajishirzi, H. 2022 · 2022
Cited alongside, same era.
Is a modular architecture enough?
Mittal, S.; Bengio, Y.; and Lajoie, G. 2022 · 2022
Cited alongside, same era.
Jaiswal, A.; Liu, S.; Chen, T.; and Wang, Z. 2023 · 2023
Closest in time.
Pruning large language models via accuracy predictor
Ji, Y.; Cao, Y.; and Liu, J. 2023 · 2023
Closest in time.
FDAPT: Federated Domain-adaptive Pre-training for Language Models
Jiang, L.; Svoboda, F.; and Lane, N. D. 2023 · 2023
Closest in time.
Memory-Efficient Fine-Tuning of Compressed Large Language Models via sub-4-bit Integer Quantization
Kim, J.; Lee, J. H.; Kim, S.; Park, J.; Yoo, K. M.; Kwon, S. J.; and Lee, D. 2023 · 2023
Closest in time.
Textbooks are all you need II: phi-1.5 technical report
Li, Y.; Bubeck, S.; Eldan, R.; Del Giorno, A.; Gunasekar, S.; and Lee, Y. T. 2023 · 2023
Closest in time.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Lin, J.; Tang, J.; Tang, H.; Yang, S.; Dang, X.; and Han, S. 2023 · 2023
Closest in time.
LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Liu, Z.; Oguz, B.; Zhao, C.; Chang, E.; Stock, P.; Mehdad, Y.; Shi, Y.; Krishnamoorthi, R.; and Chandra, V. 2023 · 2023
Closest in time.
LLM-Pruner: On the Structural Pruning of Large Language Models
Ma, X.; Fang, G.; and Wang, X. 2023 · 2023
Closest in time.
Soft Merging of Experts with Adaptive Routing
Muqeeth, M.; Liu, H.; and Raffel, C. 2023 · 2023
Closest in time.
From Sparse to Soft Mixtures of Experts
Puigcerver, J.; Riquelme, C.; Mustafa, B.; and Houlsby, N. 2023 · 2023
Closest in time.
Mixture of Prompt Experts for Generalizable and Interpretable Question Answering
Si, C.; Shi, W.; Zhao, C.; Zettlemoyer, L.; and Boyd-Graber, J. 2023 · 2023
Closest in time.
A Simple and Effective Pruning Approach for Large Language Models
Sun, M.; Liu, Z.; Bair, A.; and Kolter, J. Z. 2023 · 2023
Closest in time.
Universality and Limitations of Prompt Tuning
Wang, Y.; Chauhan, J.; Wang, W.; and Hsieh, C.-J. 2023 · 2023
Closest in time.
Parameter efficient multi-task fine-tuning by learning to transfer token-wise prompts
Wu, M.; Liu, W.; Xu, J.; Lv, C.; Ling, Z.; Li, T.; Huang, L.; Zheng, X.; and Huang, X.-J. 2023 · 2023
Closest in time.
Xu, Z.; Liu, Z.; Chen, B.; Tang, Y.; Wang, J.; Zhou, K.; Hu, X.; and Shrivastava, A. 2023 · 2023
Closest in time.
FedPrompt: Communication-Efficient and Privacy Preserving Prompt Tuning in Federated Learning
Zhao, H.; Du, W.; Li, F.; Li, P.; and Liu, G. 2023 · 2023
Closest in time.