Fetching the paper…
Reading the bibliography…
This study presents a comprehensive analysis and comparison of two predominant fine-tuning methodologies - full-parameter fine-tuning and parameter-efficient tuning - within the context of medical Large Language Models (LLMs).
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Large language models in medicine
Thirunavukarasu, A. J.; Ting, D. S. J.; Elangovan, K.; Gutierrez, L.; Tan, T. F.; and Ting, D. S. W. 2023 · 1940
Earlier work this paper cites.
Jin, D.; Pan, E.; Oufattole, N.; Weng, W.-H.; Fang, H.; and Szolovits, P. 2020 · 2009
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
A Question-Entailment Approach to Question Answering
Ben Abacha, A.; and Demner-Fushman, D. 2019 · 2019
Earlier work this paper cites.
Parameter-Efficient Transfer Learning for NLP
Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; de Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019 · 2019
Earlier work this paper cites.
PubMedQA: A Dataset for Biomedical Research Question Answering
Jin, Q.; Dhingra, B.; Liu, Z.; Cohen, W.; and Lu, X. 2019 · 2019
Earlier work this paper cites.
HEAD-QA: A Healthcare Dataset for Complex Reasoning
Vilares, D.; and Gómez-Rodríguez, C. 2019 · 2019
Earlier work this paper cites.
Exploring Versatile Generative Language Model Via Parameter-Efficient Transfer Learning
Lin, Z.; Madotto, A.; and Fung, P. 2020 · 2020
Earlier work this paper cites.
Chatbots in the fight against the COVID-19 pandemic
Miner, A. S.; Laranjo, L.; and Kocaballi, A. B. 2020 · 2020
Earlier work this paper cites.
CORD-19: The COVID-19 Open Research Dataset
Wang, L. L.; Lo, K.; Chandrasekhar, Y.; Reas, R.; Yang, J.; Burdick, D.; Eide, D.; Funk, K.; Katsis, Y.; Kinney, R. M.; Li, Y.; Liu, Z.; Merrill, W.; Mooney, P.; Murdick, D. A.; Rishi, D.; Sheehan, J.; Shen, Z.; Stilson, B.; Wade, A. D.; Wang, K.; Wang, N. X. R.; Wilhelm, C.; Xie, B.; Raymond, D. M.; Weld, D. S.; Etzioni, O.; and Kohlmeier, S. 2020 · 2020
Earlier work this paper cites.
Large Language Models for Information Retrieval: A Survey
Zhu, Y.; Yuan, H.; Wang, S.; Liu, J.; Liu, W.; Deng, C.; Dou, Z.; and Wen, J.-R. 2023 · 2020
Earlier work this paper cites.
A framework for few-shot language model evaluation
Gao, L.; Tow, J.; Biderman, S.; Black, S.; DiPofi, A.; Foster, C.; Golding, L.; Hsu, J.; McDonell, K.; Muennighoff, N.; Phang, J.; Reynolds, L.; Tang, E.; Thite, A.; Wang, B.; Wang, K.; and Zou, A. 2021 · 2021
Earlier work this paper cites.
Towards a unified view of parameter-efficient transfer learning
He, J.; Zhou, C.; Ma, X.; Berg-Kirkpatrick, T.; and Neubig, G. 2021 · 2021
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
The Power of Scale for Parameter-Efficient Prompt Tuning
Lester, B.; Al-Rfou, R.; and Constant, N. 2021 · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
Li, X. L.; and Liang, P. 2021 · 2021
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022 · 2022
Earlier work this paper cites.
MedMCQA: A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering
Pal, A.; Umapathi, L. K.; and Sankarasubbu, M. 2022 · 2022
Cited alongside, same era.
Multitask Prompted Training Enables Zero-Shot Task Generalization
Sanh, V.; Webson, A.; Raffel, C.; Bach, S. H.; Sutawika, L.; Alyafeai, Z.; Chaffin, A.; Stiegler, A.; Scao, T. L.; Raja, A.; Dey, M.; Bari, M. S.; Xu, C.; Thakker, U.; Sharma, S. S.; Szczechla, E.; Kim, T.; Chhablani, G.; Nayak, N.; Datta, D.; Chang, J.; Jiang, M. T.-J.; Wang, H.; Manica, M.; Shen, S.; Yong, Z. X.; Pandey, H.; Bawden, R.; Wang, T.; Neeraj, T.; Rozen, J.; Sharma, A.; Santilli, A.; Fevry, T.; Fries, J. A.; Teehan, R.; Bers, T.; Biderman, S.; Gao, L.; Wolf, T.; and Rush, A. M. 2022 · 2022
Cited alongside, same era.
Large Language Models Encode Clinical Knowledge
Singhal, K.; Azizi, S.; Tu, T.; Sara Mahdavi, S.; Wei, J.; Chung, H. W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S.; Payne, P.; Seneviratne, M.; Gamble, P.; Kelly, C.; Scharli, N.; Chowdhery, A.; Mansfield, P.; Aguera y Arcas, B.; Webster, D.; Corrado, G. S.; Matias, Y.; Chou, K.; Gottweis, J.; Tomasev, N.; Liu, Y.; Rajkomar, A.; Barral, J.; Semturs, C.; Karthikesalingam, A.; and Natarajan, V. 2022 · 2022
Cited alongside, same era.
Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks
HuggingFace H4 Stack Exchange Preference Dataset
Lambert, N.; Tunstall, L.; Rajani, N.; and Thrush, T. 2023 · 2023
Later among the works it cites.
A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets
Laskar, M. T. R.; Bari, M. S.; Rahman, M.; Bhuiyan, M. A. H.; Joty, S.; and Huang, J. X. 2023 · 2023
Later among the works it cites.
Platypus: Quick, cheap, and powerful refinement of llms
Lee, A. N.; Hunter, C. J.; and Ruiz, N. 2023 · 2023
Later among the works it cites.
OpenOrca: An Open Dataset of GPT Augmented FLAN Reasoning Traces
Lian, W.; Goodson, B.; Pentland, E.; Cook, A.; Vong, C.; and ”Teknium”. 2023 · 2023
Later among the works it cites.
Parameter-Efficient Fine-Tuning without Introducing New Latency
Liao, B.; Meng, Y.; and Monz, C. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, Y.; Mishra, S.; Alipoormolabashi, P.; Kordi, Y.; Mirzaei, A.; Arunkumar, A.; Ashok, A.; Dhanasekaran, A. S.; Naik, A.; Stap, D.; Pathak, E.; Karamanolakis, G.; Lai, H. G.; Purohit, I.; Mondal, I.; Anderson, J.; Kuznia, K.; Doshi, K.; Patel, M.; Pal, K. K.; Moradshahi, M.; Parmar, M.; Purohit, M.; Varshney, N.; Kaza, P. R.; Verma, P.; Puri, R. S.; Karia, R.; Sampat, S. K.; Doshi, S.; Mishra, S.; Reddy, S.; Patro, S.; Dixit, T.; Shen, X.; Baral, C.; Choi, Y.; Smith, N. A.; Hajishirzi, H.; and Khashabi, D. 2022 · 2022
Cited alongside, same era.
Finetuned Language Models Are Zero-Shot Learners
Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2022 · 2022
Cited alongside, same era.
A large language model for electronic health records
Yang, X.; Chen, A.; PourNejatian, N.; Shin, H. C.; Smith, K. E.; Parisien, C.; Compas, C.; Martin, C.; Costa, A. B.; Flores, M. G.; Zhang, Y.; Magoc, T.; Harle, C. A.; Lipori, G.; Mitchell, D. A.; Hogan, W. R.; Shenkman, E. A.; Bian, J.; and Wu, Y. 2022 · 2022
Cited alongside, same era.
Using ChatGPT to write patient clinic letters
Ali, S. R.; Dobbs, T. D.; Hutchings, H. A.; and Whitaker, I. S. 2023 · 2023
Cited alongside, same era.
A systematic review and meta-analysis on ChatGPT and its utilization in medical and dental research
Bagde, H.; Dhopte, A.; Alam, M. K.; and Basri, R. 2023 · 2023
Cited alongside, same era.
Evaluating the Feasibility of ChatGPT in Healthcare: An Analysis of Multiple Clinical and Research Scenarios
Cascella, M.; Montomoli, J.; Bellini, V.; and Bignami, E. 2023 · 2023
Cited alongside, same era.
MEDITRON-70B: Scaling Medical Pretraining for Large Language Models
Chen, Z.; Cano, A. H.; Romanou, A.; Bonnet, A.; Matoba, K.; Salvi, F.; Pagliardini, M.; Fan, S.; Köpf, A.; Mohtashami, A.; Sallinen, A.; Sakhaeirad, A.; Swamy, V.; Krawczuk, I.; Bayazit, D.; Marmet, A.; Montariol, S.; Hartley, M.-A.; Jaggi, M.; and Bosselut, A. 2023 · 2023
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms
Dettmers, T.; Pagnoni, A.; Holtzman, A.; and Zettlemoyer, L. 2023 · 2023
Cited alongside, same era.
Parameter-efficient fine-tuning of large-scale pre-trained language models
Ding, N.; Qin, Y.; Yang, G.; Wei, F.; Yang, Z.; Su, Y.; Hu, S.; Chen, Y.; Chan, C.-M.; Chen, W.; Yi, J.; Zhao, W.; Wang, X.; Liu, Z.; Zheng, H.-T.; Chen, J.; Liu, Y.; Tang, J.; Li, J.; and Sun, M. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
GPT understands, too
Liu, X.; Zheng, Y.; Du, Z.; Ding, M.; Qian, Y.; Yang, Z.; and Tang, J. 2023 · 2023
Later among the works it cites.
The Flan Collection: Designing Data and Methods for Effective Instruction Tuning
Longpre, S.; Hou, L.; Vu, T.; Webson, A.; Chung, H. W.; Tay, Y.; Zhou, D.; Le, Q. V.; Zoph, B.; Wei, J.; et al. 2023 · 2023
Later among the works it cites.
Capabilities of GPT-4 on Medical Challenge Problems
Nori, H.; King, N.; McKinney, S. M.; Carignan, D.; and Horvitz, E. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
A study of generative large language model for medical research and healthcare
Peng, C.; Yang, X.; Chen, A.; Smith, K. E.; PourNejatian, N.; Costa, A. B.; Martin, C.; Flores, M. G.; Zhang, Y.; Magoc, T.; Lipori, G.; Mitchell, D. A.; Ospina, N. S.; Ahmed, M. M.; Hogan, W. R.; Shenkman, E. A.; Guo, Y.; Bian, J.; and Wu, Y. 2023 · 2023
Later among the works it cites.
Towards Expert-Level Medical Question Answering with Large Language Models
Singhal, K.; Tu, T.; Gottweis, J.; Sayres, R.; Wulczyn, E.; Hou, L.; Clark, K.; Pfohl, S.; Cole-Lewis, H.; Neal, D.; Schaekermann, M.; Wang, A.; Amin, M.; Lachgar, S.; Mansfield, P.; Prakash, S.; Green, B.; Dominowska, E.; y Arcas, B. A.; Tomasev, N.; Liu, Y.; Wong, R.; Semturs, C.; Mahdavi, S. S.; Barral, J.; Webster, D.; Corrado, G. S.; Matias, Y.; Azizi, S.; Karthikesalingam, A.; and Natarajan, V. 2023 · 2023
Later among the works it cites.
ChatGPT and medicine: a potential threat to science or a step towards the future?
Souza, L. L. d.; Fonseca, F. P.; Martins, M. D.; de Almeida, O. P.; Pontes, H. A. R.; Coracin, F. L.; Lopes, M. A.; Khurram, S. A.; Santos-Silva, A. R.; Hagag, A.; and Vargas, P. A. 2023 · 2023
Later among the works it cites.
Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding
Toma, A.; Lawler, P. R.; Ba, J.; Krishnan, R. G.; Rubin, B. B.; and Wang, B. 2023 · 2023
Later among the works it cites.
ChatGPT, enhanced with clinical practice guidelines, is a superior decision support tool
Wang, Y.; Visweswaran, S.; Kappor, S.; Kooragayalu, S.; and Wu, X. 2023 · 2023
Later among the works it cites.
PMC-LLaMA: Towards Building Open-source Language Models for Medicine
Wu, C.; Lin, W.; Zhang, X.; Zhang, Y.; Wang, Y.; and Xie, W. 2023 · 2023
Later among the works it cites.
A Survey of Large Language Models
Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; Du, Y.; Yang, C.; Chen, Y.; Chen, Z.; Jiang, J.; Ren, R.; Li, Y.; Tang, X.; Liu, Z.; Liu, P.; Nie, J.-Y.; and Wen, J.-R. 2023 · 2023
Later among the works it cites.
A Survey of Large Language Models in Medicine: Principles, Applications, and Challenges
Zhou, H.; Liu, F.; Gu, B.; Zou, X.; Huang, J.; Wu, J.; Li, Y.; Chen, S. S.; Zhou, P.; Liu, J.; Hua, Y.; Mao, C.; Wu, X.; Zheng, Y.; Clifton, L.; Li, Z.; Luo, J.; and Clifton, D. A. 2023 · 2023
Later among the works it cites.