Fetching the paper…
Reading the bibliography…
Fine-tuning is the most effective way of adapting pre-trained large language models (LLMs) to downstream applications.
Random forests
L. Breiman · 2001
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
J. Berant, A. Chou, R. Frostig, and P. Liang · 2013
Earlier work this paper cites.
Tensorflow: learning functions at scale
M. Abadi · 2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Professor forcing: A new algorithm for training recurrent networks
A. M. Lamb, A. G. ALIAS PARTH GOYAL, Y. Zhang, S. Zhang, A. C. Courville, and Y. Bengio · 2016
Earlier work this paper cites.
Pruning filters for efficient convnets
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf · 2016
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
R. Nallapati, B. Zhou, C. Gulcehre, B. Xiang, et al · 2016
Earlier work this paper cites.
Sparse communication for distributed gradient descent
A. F. Aji and K. Heafield · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
M. Sundararajan, A. Taly, and Q. Yan · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Snip: Single-shot network pruning based on connection sensitivity
N. Lee, T. Ajanthan, and P. H. Torr · 2018
Earlier work this paper cites.
Samsum corpus: A human-annotated dialogue dataset for abstractive summarization
B. Gliwa, I. Mochol, M. Biesek, and A. Wawer · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Cited alongside, same era.
Huggingface’s transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, et al · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, J. Gao, Y. Choi, et al · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models
E. B. Zaken, S. Ravfogel, and Y. Goldberg · 2021
Later among the works it cites.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2022
Later among the works it cites.
On-device training under 256kb memory
J. Lin, L. Zhu, W.-M. Chen, W.-C. Wang, C. Gan, and S. Han · 2022
Later among the works it cites.
P-tuning: Prompt tuning can be comparable to fine-tuning across scales and tasks
X. Liu, K. Ji, Y. Fu, W. Tam, Z. Du, Z. Yang, and J. Tang · 2022
Later among the works it cites.
Crosslingual generalization through multitask finetuning
N. Muennighoff, T. Wang, L. Sutawika, A. Roberts, S. Biderman, T. L. Scao, M. S. Bari, S. Shen, Z.-X. Yong, H. Schoelkopf, et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Cachola, K. Lo, A. Cohan, and D. S. Weld · 2020
Cited alongside, same era.
How does weight correlation affect generalisation ability of deep neural networks?
G. Jin, X. Yi, L. Zhang, L. Zhang, S. Schewe, and X. Huang · 2020
Cited alongside, same era.
Layer-adaptive sparsity for the magnitude-based pruning
J. Lee, S. Park, S. Mo, S. Ahn, and J. Shin · 2020
Cited alongside, same era.
Green ai
R. Schwartz, J. Dodge, N. A. Smith, and O. Etzioni · 2020
Cited alongside, same era.
Dialogsum: A real-life scenario dialogue summarization dataset
Y. Chen, Y. Liu, L. Chen, and Y. Zhang · 2021
Cited alongside, same era.
Fast axiomatic attribution for neural networks
R. Hesse, S. Schaub-Meyer, and S. Roth · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Cited alongside, same era.
Later among the works it cites.
Fine-tuned language models are continual learners
T. Scialom, T. Chakrabarty, and S. Muresan · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Later among the works it cites.
https://aiindex.stanford.edu/report/ , 2023
2023 AI index report · 2023
Closest in time.
h2ogpt: Democratizing large language models
A. Candel, J. McKinney, P. Singer, P. Pfeiffer, M. Jeblick, P. Prabhu, J. Gambera, M. Landry, S. Bansal, R. Chesler, et al · 2023
Closest in time.
Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models
Z. Hu, Y. Lan, L. Wang, W. Xu, E.-P. Lim, R. K.-W. Lee, L. Bing, and S. Poria · 2023
Closest in time.
Tinytrain: Deep neural network training at the extreme edge
Y. D. Kwon, R. Li, S. I. Venieris, J. Chauhan, N. D. Lane, and C. Mascolo · 2023
Closest in time.
Make your pre-trained model reversible: From parameter to memory efficient fine-tuning
B. Liao, S. Tan, and C. Monz · 2023
Closest in time.
Fine-tuning language models with just forward passes
S. Malladi, T. Gao, E. Nichani, A. Damian, J. D. Lee, D. Chen, and S. Arora · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
H. Wang and W. Gao · 2023
Closest in time.
Adaptive budget allocation for parameter-efficient fine-tuning
Q. Zhang, M. Chen, A. Bukharin, P. He, Y. Cheng, W. Chen, and T. Zhao · 2023
Closest in time.