Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have made remarkable advancements in the field of natural language processing.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning,
H. Liu, D. Tam, M. Muqeeth, J. Mohta, T. Huang, M. Bansal, C. A. Raffel, · 1965
Earlier work this paper cites.
Decoupled weight decay regularization,
I. Loshchilov, F. Hutter, · 2017
Earlier work this paper cites.
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, · 2017
Earlier work this paper cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training,
Y. Lin, S. Han, H. Mao, Y. Wang, W. J. Dally, · 2017
Earlier work this paper cites.
Interactive supercomputing on 40,000 cores for machine learning and data analysis,
A. Reuther, J. Kepner, C. Byun, S. Samsi, W. Arcand, D. Bestor, B. Bergeron, V. Gadepally, M. Houle, M. Hubbell, M. Jones, A. Klein, L. Milechin, J. Mullen, A. Prout, A. Rosa, C. Yee, P. Michaleas, · 2018
Earlier work this paper cites.
Energy and policy considerations for deep learning in nlp,
E. Strubell, A. Ganesh, A. McCallum, · 2019
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering,
Q. Jin, B. Dhingra, Z. Liu, W. W. Cohen, X. Lu, · 2019
Earlier work this paper cites.
N. Poerner, U. Waltinger, H. Schütze, · 2020
Earlier work this paper cites.
Squeezebert: What can computer vision teach nlp about efficient neural networks?,
F. N. Iandola, A. E. Shaw, R. Krishna, K. W. Keutzer, · 2020
Earlier work this paper cites.
Program synthesis with large language models,
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al., · 2021
Cited alongside, same era.
Evaluating large language models trained on code,
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. d. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al., · 2021
Cited alongside, same era.
Domain-specific language model pretraining for biomedical natural language processing,
Y. Gu, R. Tinn, H. Cheng, M. Lucas, N. Usuyama, X. Liu, T. Naumann, J. Gao, H. Poon, · 2021
Cited alongside, same era.
Few-shot question answering by pretraining span selection,
O. Ram, Y. Kirstain, J. Berant, A. Globerson, O. Levy, · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation,
Towards parameter-efficient automation of data wrangling tasks with prefix-tuning,
D. Vos, T. Döhmen, S. Schelter, · 2022
Later among the works it cites.
Biogpt: generative pre-trained transformer for biomedical text generation and mining,
R. Luo, L. Sun, Y. Xia, T. Qin, S. Zhang, H. Poon, T.-Y. Liu, · 2022
Later among the works it cites.
Z. Guo, P. Wang, Y. Wang, S. Yu, Dr. LLaMA: Improving small language models in domain-specific qa via generative data augmentation, 2023. URL: https://github.com/zguo0525/Dr.LLaMA
2023
Closest in time.
A survey of large language models,
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al., · 2023
Closest in time.
A short survey of viewing large language models in legal aspect,
Z. Sun, · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. L. Li, P. Liang, · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models,
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, · 2021
Cited alongside, same era.
Guiding generative language models for data augmentation in few-shot text classification,
A. Edwards, A. Ushio, J. Camacho-Collados, H. de Ribaupierre, A. Preece, · 2021
Cited alongside, same era.
Training compute-optimal large language models,
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al., · 2022
Cited alongside, same era.
Closest in time.
Capabilities of gpt-4 on medical challenge problems,
H. Nori, N. King, S. M. McKinney, D. Carignan, E. Horvitz, · 2023
Closest in time.
Llama: Open and efficient foundation language models,
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al., · 2023
Closest in time.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, T. B. Hashimoto, Stanford alpaca: An instruction-following llama model, https://github.com/tatsu-lab/stanford_alpaca , 2023
2023
Closest in time.