Fetching the paper…
Reading the bibliography…
As opposed to evaluating computation and logic-based reasoning, current benchmarks for evaluating large language models (LLMs) in medicine are primarily focused on question-answering involving domain knowledge and descriptive reasoning.
Prediction of creatinine clearance from serum creatinine
D. W. Cockcroft and H. Gault · 1976
Earlier work this paper cites.
A more accurate method to estimate glomerular filtration rate from serum creatinine: a new prediction equation
A. S. Levey, J. P. Bosch, J. B. Lewis, T. Greene, N. Rogers, D. Roth, and M. of Diet in Renal Disease Study Group* · 1999
Earlier work this paper cites.
Validation of clinical classification schemes for predicting stroke: results from the national registry of atrial fibrillation
B. F. Gage, A. D. Waterman, W. Shannon, M. Boechler, M. W. Rich, and M. J. Radford · 2001
Earlier work this paper cites.
Chest pain in the emergency room: value of the heart score
A. Six, B. Backus, and J. Kelder · 2008
Earlier work this paper cites.
2010 rheumatoid arthritis classification criteria
C. Initiative · 2010
Earlier work this paper cites.
Overview of the trec 2014 clinical decision support track
M. S. Simpson, E. M. Voorhees, and W. R. Hersh · 2014
Earlier work this paper cites.
Overview of the trec 2015 clinical decision support track
K. Roberts, M. S. Simpson, E. M. Voorhees, and W. R. Hersh · 2015
Earlier work this paper cites.
Clinical calculators in hospital medicine: availability, classification, and needs
M. A. Dziadzko, O. Gajic, B. W. Pickering, and V. Herasevich · 2016
Earlier work this paper cites.
Mdcalc medical calculator app review
A. Elovic and A. Pourmand · 2019
Earlier work this paper cites.
Medical calculators: prevalence, and barriers to use
T. A. Green, S. Whitt, J. L. Belden, S. Erdelez, and C.-R. Shyu · 2019
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Q. Jin, B. Dhingra, Z. Liu, W. Cohen, and X. Lu · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt · 2021
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
D. Jin, E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P. Szolovits · 2021
Cited alongside, same era.
Overview of the trec 2021 clinical trials track
K. Roberts, D. Demner-Fushman, E. M. Voorhees, S. Bedrick, and W. R. Hersh · 2021
Cited alongside, same era.
Bigbio: A framework for data-centric biomedical natural language processing
J. Fries, L. Weber, N. Seelam, G. Altay, D. Datta, S. Garda, S. Kang, R. Su, W. Kusa, S. Cahyawijaya, et al · 2022
Cited alongside, same era.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Cognitive architectures for language agents
T. R. Sumers, S. Yao, K. Narasimhan, and T. L. Griffiths · 2023
Later among the works it cites.
Evaluating large language models on medical evidence summarization
L. Tang, Z. Sun, B. Idnay, J. G. Nestor, A. Soroush, P. A. Elias, Z. Xu, Y. Ding, G. Durrett, J. F. Rousseau, et al · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
Large language models in medicine
A. J. Thirunavukarasu, D. S. J. Ting, K. Elangovan, L. Gutierrez, T. F. Tan, and D. S. W. Ting · 2023
Later among the works it cites.
Scaling clinical trial matching using large language models: A case study in oncology
C. Wong, S. Zhang, Y. Gu, C. Moung, J. Abel, N. Usuyama, R. Weerasinghe, B. Piening, T. Naumann, C. Bifulco, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
A. Pal, L. K. Umapathi, and M. Sankarasubbu · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, et al · 2023
Cited alongside, same era.
epfllm megatron-llm, 2023
A. H. Cano, M. Pagliardini, A. Köpf, K. Matoba, A. Mohtashami, X. Wang, O. S. Fan, A. Marmet, D. Bayazit, I. Krawczuk, Z. Chen, F. Salvi, A. Bosselut, and M. Jaggi · 2023
Cited alongside, same era.
Pal: Program-aided language models
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig · 2023
Cited alongside, same era.
Augmentation of chatgpt with clinician-informed tools improves performance on medical calculation tasks
A. J. Goodell, S. N. Chu, D. Rouholiman, and L. F. Chu · 2023
Cited alongside, same era.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Cited alongside, same era.
Later among the works it cites.
A large-scale dataset of patient summaries for retrieval-based clinical decision support systems
Z. Zhao, Q. Jin, F. Chen, T. Peng, and S. Yu · 2023
Later among the works it cites.
S. Zhuang, B. Koopman, and G. Zuccon · 2023
Later among the works it cites.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al · 2024
Closest in time.
Can large language models reason about medical questions?
V. Liévin, C. E. Hother, A. G. Motzfeldt, and O. Winther · 2024
Closest in time.
Augmenting large language models with chemistry tools
A. M. Bran, S. Cox, O. Schilter, C. Baldassari, A. D. White, and P. Schwaller · 2024
Closest in time.
Capabilities of gemini models in medicine
K. Saab, T. Tu, W.-H. Weng, R. Tanno, D. Stutz, E. Wulczyn, F. Zhang, T. Strother, C. Park, E. Vedadi, et al · 2024
Closest in time.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2024
Closest in time.
Opportunities and challenges for chatgpt and large language models in biomedicine and health
S. Tian, Q. Jin, L. Yeganova, P.-T. Lai, Q. Zhu, X. Chen, Y. Yang, Q. Chen, W. Kim, D. C. Comeau, et al · 2024
Closest in time.
Adapted large language models can outperform medical experts in clinical text summarization
D. Van Veen, C. Van Uden, L. Blankemeier, J.-B. Delbrouck, A. Aali, C. Bluethgen, A. Pareek, M. Polacin, E. P. Reis, A. Seehofnerová, et al · 2024
Closest in time.
Zero-shot clinical trial patient matching with llms
M. Wornow, A. Lozano, D. Dash, J. Jindal, K. W. Mahaffey, and N. H. Shah · 2024
Closest in time.
Pmc-llama: toward building open-source language models for medicine
C. Wu, W. Lin, X. Zhang, Y. Zhang, W. Xie, and Y. Wang · 2024
Closest in time.