Fetching the paper…
Reading the bibliography…
Curated datasets for healthcare are often limited due to the need of human annotations from experts.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Mimic-cxr-jpg, a large publicly available database of labeled chest radiographs
Alistair EW Johnson, Tom J Pollard, Nathaniel R Greenbaum, Matthew P Lungren, Chih-ying Deng, Yifan Peng, Zhiyong Lu, Roger G Mark, Seth J Berkowitz, and Steven Horng. 2019 · 1901
Earlier work this paper cites.
Scibert: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 1903
Earlier work this paper cites.
Clinicalbert: Modeling clinical notes and predicting hospital readmission
Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. 2019 · 1904
Earlier work this paper cites.
Abstract meaning representation for sembanking
Laura Banarescu, Claire Bonial, Shu Cai, Madalina Georgescu, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Philipp Koehn, Martha Palmer, and Nathan Schneider. 2013 · 2013
Earlier work this paper cites.
An overview of the bioasq large-scale biomedical semantic indexing and question answering competition
George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, et al. 2015 · 2015
Earlier work this paper cites.
Preparing a collection of radiology examinations for distribution and retrieval
Dina Demner-Fushman, Marc D Kohli, Marc B Rosenman, Sonya E Shooshan, Laritza Rodriguez, Sameer Antani, George R Thoma, and Clement J McDonald. 2016 · 2016
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016 · 2016
Earlier work this paper cites.
Pubmed 200k rct: a dataset for sequential sentence classification in medical abstracts
Franck Dernoncourt and Ji Young Lee. 2017 · 2017
Earlier work this paper cites.
Construction of the literature graph in semantic scholar
Waleed Ammar, Dirk Groeneveld, Chandra Bhagavatula, Iz Beltagy, Miles Crawford, Doug Downey, Jason Dunkelberger, Ahmed Elgohary, Sergey Feldman, Vu Ha, et al. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Bias amplification in artificial intelligence systems
Kirsten Lloyd. 2018 · 2018
Earlier work this paper cites.
emrqa: A large corpus for question answering on electronic medical records
Anusri Pampari, Preethi Raghavan, Jennifer Liang, and Jian Peng. 2018 · 2018
Earlier work this paper cites.
Bioread: A new dataset for biomedical reading comprehension
Dimitris Pappas, Ion Androutsopoulos, and Harris Papageorgiou. 2018 · 2018
Earlier work this paper cites.
Microsoft’s politically correct chatbot is even worse than its racist one
Chloe Rose Stuart-Ulin. 2018 · 2018
Earlier work this paper cites.
Style transformer: Unpaired text style transfer without disentangled latent representation
Ning Dai, Jianze Liang, Xipeng Qiu, and Xuanjing Huang. 2019 · 2019
Earlier work this paper cites.
Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison
Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu, Silviana Ciurea-Ilcus, Chris Chute, Henrik Marklund, Behzad Haghgoo, Robyn Ball, Katie Shpanskaya, et al. 2019 · 2019
Cited alongside, same era.
Probing biomedical embeddings from language models
Qiao Jin, Bhuwan Dhingra, William W Cohen, and Xinghua Lu. 2019a · 2019
Cited alongside, same era.
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019b · 2019
Cited alongside, same era.
Feature-wise bias amplification
Klas Leino, Matt Fredrikson, Emily Black, Shayak Sen, and Anupam Datta. 2019 · 2019
Cited alongside, same era.
Transfer learning in biomedical natural language processing: An evaluation of bert and elmo on ten benchmarking datasets
Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019 · 2019
Cited alongside, same era.
Detecting beats in the photoplethysmogram: benchmarking open-source algorithms
Peter H Charlton, Kevin Kotzen, Elisa Mejía-Mejía, Philip J Aston, Karthik Budidha, Jonathan Mant, Callum Pettit, Joachim A Behar, and Panicos A Kyriacou. 2022 · 2022
Later among the works it cites.
Controlling bias exposure for fair interpretable predictions
Zexue He, Yu Wang, Julian McAuley, and Bodhisattwa Prasad Majumder. 2022 · 2022
Later among the works it cites.
BioGPT: generative pre-trained transformer for biomedical text generation and mining
Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022 · 2022
Later among the works it cites.
M3: Multi-level dataset for multi-document summarisation of medical studies
Julia Otmakhova, Karin Verspoor, Timothy Baldwin, Antonio Jimeno Yepes, and Jey Han Lau. 2022 · 2022
Later among the works it cites.
Leashing the inner demons: Self-detoxification for language models
Canwen Xu, Zexue He, Zhankui He, and Julian McAuley. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020 · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020 · 2020
Cited alongside, same era.
Biobert: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020 · 2020
Cited alongside, same era.
Effective transfer learning for identifying similar questions: matching user questions to covid-19 faqs
Clara H McCreery, Namit Katariya, Anitha Kannan, Manish Chablani, and Xavier Amatriain. 2020 · 2020
Cited alongside, same era.
Beyond accuracy: Behavioral testing of nlp models with checklist
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Cited alongside, same era.
Biomegatron: Larger biomedical domain language model
Hoo-Chang Shin, Yang Zhang, Evelina Bakhturina, Raul Puri, Mostofa Patwary, Mohammad Shoeybi, and Raghav Mani. 2020 · 2020
Cited alongside, same era.
Parrot: Paraphrase generation for NLU
Prithiviraj Damodaran. 2021 · 2021
Cited alongside, same era.
Radbert: Adapting transformer-based language models to radiology
An Yan, Julian McAuley, Xing Lu, Jiang Du, Eric Y Chang, Amilcare Gentili, and Chun-Nan Hsu. 2022 · 2022
Later among the works it cites.
Biobart: Pretraining and evaluation of a biomedical generative language model
Hongyi Yuan, Zheng Yuan, Ruyi Gan, Jiaxing Zhang, Yutao Xie, and Sheng Yu. 2022 · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022 · 2022
Later among the works it cites.
Chatgpt and the future of medical writing
Som Biswas. 2023 · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023 · 2023
Closest in time.
Automated medical coding on mimic-iii and mimic-iv: A critical review and replicability study
Joakim Edin, Alexander Junge, Jakob D Havtorn, Lasse Borgholt, Maria Maistro, Tuukka Ruotsalo, and Lars Maaløe. 2023 · 2023
Closest in time.
Dr. llama: Improving small language models in domain-specific qa via generative data augmentation
Zhen Guo, Peiqi Wang, Yanwei Wang, and Shangdi Yu. 2023 · 2023
Closest in time.
Zero-shot clinical entity recognition using chatgpt
Yan Hu, Iqra Ameer, Xu Zuo, Xueqing Peng, Yujia Zhou, Zehan Li, Yiming Li, Jianfu Li, Xiaoqian Jiang, and Hua Xu. 2023 · 2023
Closest in time.
Mimic-iv, a freely accessible electronic health record dataset
Alistair EW Johnson, Lucas Bulgarelli, Lu Shen, Alvin Gayles, Ayad Shammout, Steven Horng, Tom J Pollard, Benjamin Moody, Brian Gow, Li-wei H Lehman, et al. 2023 · 2023
Closest in time.
Does synthetic data generation of llms help clinical text mining?
Ruixiang Tang, Xiaotian Han, Xiaoqian Jiang, and Xia Hu. 2023 · 2023
Closest in time.
Pmc-llama: Towards building open-source language models for medicine
Chaoyi Wu, Weixiong Lin, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. 2023 · 2023
Closest in time.