Fetching the paper…
Reading the bibliography…
The ability of large language models (LLMs) to follow natural language instructions with human-level fluency suggests many opportunities in healthcare to reduce administrative burden and improve quality of care.
Rank correlation methods
M. G. Kendall · 1948
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Hiding in plain sight: use of realistic surrogates to reduce exposure of protected health information in clinical text
D. Carrell, B. Malin, J. Aberdeen, S. Bayer, C. Clark, B. Wellner, and L. Hirschman · 2013
Earlier work this paper cites.
Improvements to bm25 and language models examined
A. Trotman, A. Puurula, and B. Burgess · 2014
Earlier work this paper cites.
Feasibility and utility of applications of the common data model to multiple, disparate observational health databases
E. A. Voss, R. Makadia, A. Matcho, Q. Ma, C. Knoll, M. Schuemie, F. J. DeFalco, A. Londhe, V. Zhu, and P. B. Ryan · 2015
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Earlier work this paper cites.
chrf++: words helping character n-grams
M. Popović · 2017
Earlier work this paper cites.
emrQA: A large corpus for question answering on electronic medical records
A. Pampari, P. Raghavan, J. Liang, and J. Peng · 2018
Earlier work this paper cites.
Annotating electronic medical records for question answering
P. Raghavan, S. Patwardhan, J. J. Liang, and M. V. Devarakonda · 2018
Earlier work this paper cites.
Annotating and characterizing clinical sentences with explicit why-qa cues
J. Fan · 2019
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
2018 n2c2 shared task on adverse drug events and medication extraction in electronic health records
S. Henry, K. Buchan, M. Filannino, A. Stubbs, and O. Uzuner · 2020
Earlier work this paper cites.
The association between perceived electronic health record usability and professional burnout among us physicians
E. R. Melnick, L. N. Dyrbye, C. A. Sinsky, M. Trockel, C. P. West, L. Nedelec, M. A. Tutty, and T. Shanafelt · 2020
Earlier work this paper cites.
COMET: A neural framework for MT evaluation
R. Rei, C. Stewart, A. C. Farinha, and A. Lavie · 2020
Earlier work this paper cites.
Structured chart review: Assessment of a structured chart review methodology
A. Siems, R. Banks, R. Holubkov, K. L. Meert, C. Bauerfeld, D. Beyda, R. A. Berg, Y. Bulut, R. S. Burd, J. Carcillo, et al · 2020
Earlier work this paper cites.
How physicians spend their work time: an ecological momentary assessment
F. Toscano, E. O’Donnell, J. E. Broderick, M. May, P. Tucker, M. A. Unruh, G. Messina, and L. P. Casalino · 2020
Earlier work this paper cites.
Clinical reading comprehension: A thorough analysis of the emrQA dataset
X. Yue, B. J. Gutierrez, and H. Sun · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with bert
T. Zhang, V. Koshre, F. Wu, K. Weinberger, and Y. Artzi · 2020
Cited alongside, same era.
The skyrocketing volume of healthcare data makes privacy imperative
N. Culbertson · 2021
Cited alongside, same era.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
D. Jin, E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P. Szolovits · 2021
Cited alongside, same era.
Electronic health records and physician burnout: A scoping review
R. Muhiyaddin, A. H. ElFadl, E. Mohamed, Z. Shah, T. Alam, A. A. Abd-alrazaq, and M. S. Househ · 2021
Cited alongside, same era.
Experiments on portuguese clinical question answering
L. E. S. e. Oliveira, E. T. R. Schneider, Y. B. Gumiel, M. A. P. d. Luz, E. C. Paraiso, and C. Moro · 2021
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Closest in time.
Med42 - a clinical large language model, 2023
C. Christophe, A. Gupta, N. Hayat, P. Kanithi, A. Al-Mahrooqi, P. Munjal, M. Pimentel, T. Raha, R. Rajan, and S. Khan · 2023
Closest in time.
Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models
T. H. Kung, M. Cheatham, A. Medenilla, C. Sillos, L. De Leon, C. Elepaño, M. Madriaga, R. Aggabao, G. Diaz-Candido, J. Maningo, et al · 2023
Closest in time.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
N. T. MosaicML · 2023
Closest in time.
Dera: enhancing large language model completions with dialog-enabled resolving agents
V. Nair, E. Schumacher, G. Tso, and A. Kannan · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cliniqg4qa: Generating diverse questions for domain adaptation of clinical question answering
X. Yue, X. F. Zhang, Z. Yao, S. Lin, and H. Sun · 2021
Cited alongside, same era.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, E. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2022
Cited alongside, same era.
Flashattention: Fast and memory-efficient exact attention with io-awareness
T. Dao, D. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Cited alongside, same era.
Results of wmt22 metrics shared task: Stop using bleu–neural metrics are better and more robust
M. Freitag, R. Rei, N. Mathur, C.-k. Lo, C. Stewart, E. Avramidis, T. Kocmi, G. Foster, A. Lavie, and A. F. Martins · 2022
Cited alongside, same era.
Medical documentation burden among us office-based physicians in 2019: a national study
A. Gaffney, S. Woolhandler, C. Cai, D. Bor, J. Himmelstein, D. McCormick, and D. U. Himmelstein · 2022
Cited alongside, same era.
Learning to ask like a physician
E. Lehman, V. Lialin, K. E. Legaspi, A. J. Sy, P. T. Pile, N. R. Alberto, R. R. Ragasa, C. V. Puyat, M. K. Taliño, I. R. Alberto, P. G. Alfonso, D. Moukheiber, B. Wallace, A. Rumshisky, J. Liang, P. Raghavan, L. A. Celi, and P. Szolovits · 2022
Cited alongside, same era.
NLG evaluation metrics beyond correlation analysis: An empirical metric preference checklist
I. Nimah, M. Fang, V. Menkovski, and M. Pechenizkiy · 2023
Closest in time.
Capabilities of gpt-4 on medical challenge problems
H. Nori, N. King, S. M. McKinney, D. Carignan, and E. Horvitz · 2023
Closest in time.
Large language models propagate race-based medicine
J. A. Omiye, J. C. Lester, S. Spichak, V. Rotemberg, and R. Daneshjou · 2023
Closest in time.
Creation and adoption of large language models in medicine
N. H. Shah, D. Entwistle, and M. A. Pfeffer · 2023
Closest in time.
Large language models encode clinical knowledge
K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl, et al · 2023
Closest in time.
Large language models in medicine
A. J. Thirunavukarasu, D. S. J. Ting, K. Elangovan, L. Gutierrez, T. F. Tan, and D. S. W. Ting · 2023
Closest in time.
Clinical camel: An open expert-level medical language model with dialogue-based knowledge encoding, 2023
A. Toma, P. R. Lawler, J. Ba, R. G. Krishnan, B. B. Rubin, and B. Wang · 2023
Closest in time.
Creating large language model applications utilizing langchain: A primer on developing llm apps fast
O. Topsakal and T. C. Akinci · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Closest in time.
Pmc-llama: Further finetuning llama on medical papers
C. Wu, X. Zhang, Y. Zhang, Y. Wang, and W. Xie · 2023
Closest in time.
Alpacare:instruction-tuned large language models for medical application, 2023
X. Zhang, C. Tian, X. Yang, L. Chen, Z. Li, and L. R. Petzold · 2023
Closest in time.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al · 2023
Closest in time.