Fetching the paper…
Reading the bibliography…
Although recent advances in scaling large language models (LLMs) have resulted in improvements on many NLP tasks, it remains unclear whether these models trained primarily with general web text are the right tool in highly specialized, safety critical domains such as clinical text.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1907
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. (2019) · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N. M., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2020) · 1910
Earlier work this paper cites.
Pruning versus clipping in neural networks
Janowsky, S. A. (1989) · 1989
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T. J., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020) · 2001
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A. (2020) · 2004
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., and Smith, N. A. (2020) · 2004
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T. J., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 2005
Earlier work this paper cites.
Transition from paper to electronic inpatient physician notes
Payne, T. H., tenBroek, A. E., Fletcher, G. S., and Labuguen, M. C. (2010) · 2010
Earlier work this paper cites.
Annotating temporal information in clinical narratives
Sun, W., Rumshisky, A., and Uzuner, O. (2013) · 2012
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G. E., Vinyals, O., and Dean, J. (2015) · 2015
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Johnson, A. E., Pollard, T. J., Shen, L., Li-wei, H. L., Feng, M., Ghassemi, M., Moody, B., Szolovits, P., Celi, L. A., and Mark, R. G. (2016) · 2016
Earlier work this paper cites.
Fixing weight decay regularization in adam
Loshchilov, I. and Hutter, F. (2017) · 2017
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., and Song, D. X. (2018) · 2018
Earlier work this paper cites.
emrqa: A large corpus for question answering on electronic medical records
Pampari, A., Raghavan, P., Liang, J. J., and Peng, J. (2018) · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A. and Narasimhan, K. (2018) · 2018
Earlier work this paper cites.
Lessons from natural language inference in the clinical domain
Romanov, A. and Shivade, C. (2018) · 2018
Cited alongside, same era.
Publicly available clinical BERT embeddings
Alsentzer, E., Murphy, J., Boag, W., Weng, W.-H., Jindi, D., Naumann, T., and McDermott, M. (2019) · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 2019
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y. (2019) · 2019
Cited alongside, same era.
Relation extraction from clinical narratives using pre-trained language models
Wei, Q., Ji, Z., Si, Y., Du, J., Wang, J., Tiryaki, F., Wu, S., Tao, C., Roberts, K., and Qi, W. (2020) · 2019
Cited alongside, same era.
GPT-NeoX-20B: An open-source autoregressive language model
Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., Pieler, M., Prashanth, U. S., Purohit, S., Reynolds, L., Tow, J., Wang, B., and Weinbach, S. (2022) · 2022
Later among the works it cites.
Pubmed gpt: a domain-specific large language model for biomedical text
Bolton, E., Hall, D., Yasunaga, M., Liang, P., Carbin, M., Frankle, J., Venigalla, A., Manning, C., and Lee, T. (2022) · 2022
Later among the works it cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Narang, S., Mishra, G., Yu, A., Zhao, V., Huang, Y., Dai, A., Yu, H., Petrov, S., Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J. (2022) · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Does patient access to clinical notes change documentation?
Blease, C., Torous, J., and Hägglund, M. (2020) · 2020
Cited alongside, same era.
Pretrained language models for biomedical and clinical tasks: Understanding and extending the state-of-the-art
Lewis, P., Ott, M., Du, J., and Stoyanov, V. (2020) · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. (2021) · 2021
Cited alongside, same era.
Does bert pretrained on clinical notes reveal sensitive data?
Lehman, E. P., Jain, S., Pichotta, K., Goldberg, Y., and Wallace, B. C. (2021) · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P. (2021) · 2021
Cited alongside, same era.
Clip: A dataset for extracting action items for physicians from hospital discharge notes
Mullenbach, J., Pruksachatkun, Y., Adler, S., Seale, J. M., Swartz, J., McKelvey, T. G., Dai, H., Yang, Y., and Sontag, D. A. (2021) · 2021
Cited alongside, same era.
Structured prediction as translation between augmented natural languages
Paolini, G., Athiwaratkun, B., Krone, J., Ma, J., Achille, A., Anubhai, R., dos Santos, C. N., Xiang, B., and Soatto, S. (2021) · 2021
Cited alongside, same era.
Performance of chatgpt on usmle: Potential for ai-assisted medical education using large language models
Kung, T., Cheatham, M., Medenilla, A., Sillos, C., De Leon, L., Elepaño, C., Madriaga, M., Aggabao, R., Diaz-Candido, G., Maningo, J., Tseng, V., and ChatGPT (2022) · 2022
Later among the works it cites.
Clinical-longformer and clinical-bigbird: Transformers for long clinical sequences
Li, Y., Wehbe, R. M., Ahmad, F. S., Wang, H., and Luo, Y. (2022) · 2022
Later among the works it cites.
Can large language models reason about medical questions?
Li’evin, V., Hother, C. E., and Winther, O. (2022) · 2022
Later among the works it cites.
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Liu, H., Tam, D., Muqeeth, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C. (2022) · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L. E., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R. J. (2022) · 2022
Later among the works it cites.
Large language models encode clinical knowledge
Singhal, K., Azizi, S., Tu, T., Mahdavi, S., Wei, J. L. K., Chung, H. W., Scales, N., Tanwani, A. K., Cole-Lewis, H. J., Pfohl, S. J., Payne, P. A., Seneviratne, M. G., Gamble, P., Kelly, C., Scharli, N., Chowdhery, A., Mansfield, P. D., y Arcas, B. A., Webster, D. R., Corrado, G. S., Matias, Y., Chou, K. H.-L., Gottweis, J., Tomaev, N., Liu, Y., Rajkomar, A., Barral, J. K., Semturs, C., Karthikesalingam, A., and Natarajan, V. (2022) · 2022
Later among the works it cites.
RadQA: A question answering dataset to improve comprehension of radiology reports
Soni, S., Gudala, M., Pajouhi, A., and Roberts, K. (2022) · 2022
Later among the works it cites.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., hsin Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. (2022) · 2022
Later among the works it cites.
A large language model for electronic health records
Yang, X., Chen, A., PourNejatian, N., Shin, H. C., Smith, K. E., Parisien, C., Compas, C., Martin, C., Costa, A. B., Flores, M. G., and et al. (2022) · 2022
Later among the works it cites.
General availability of azure openai service expands access to large, advanced ai models with added enterprise benefits
Boyd, E. (2023) · 2023
Closest in time.
Author correction: Mimic-iv, a freely accessible electronic health record dataset
Johnson, A. E., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., Pollard, T. J., Moody, B., Gow, B., Lehman, L.-w. H., and et al. (2023) · 2023
Closest in time.