Fetching the paper…
Reading the bibliography…
There has been an influx of biomedical domain-specific language models, showing language models pre-trained on biomedical text perform better on biomedical domain benchmarks than those trained on general domain text corpora such as Wikipedia and Books.
Scibert: Pretrained contextualized embeddings for scientific text
Iz Beltagy, Arman Cohan, and Kyle Lo. 2019 · 1903
Earlier work this paper cites.
Publicly available clinical bert embeddings
Emily Alsentzer, John R Murphy, Willie Boag, Wei-Hung Weng, Di Jin, Tristan Naumann, and Matthew McDermott. 2019 · 1904
Earlier work this paper cites.
Clinicalbert: Modeling clinical notes and predicting hospital readmission
Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. 2019 · 1904
Earlier work this paper cites.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019 · 1905
Earlier work this paper cites.
Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 1909
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 1910
Earlier work this paper cites.
Dice loss for data-imbalanced nlp tasks
Xiaoya Li, Xiaofei Sun, Yuxian Meng, Junjun Liang, Fei Wu, and Jiwei Li. 2019 · 1911
Earlier work this paper cites.
Text chunking using transformation-based learning
Lance A Ramshaw and Mitchell P Marcus. 1999 · 1999
Cited alongside, same era.
Smote: synthetic minority over-sampling technique
Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer. 2002 · 2002
Cited alongside, same era.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik F Sang and Fien De Meulder. 2003 · 2003
Cited alongside, same era.
Domain-specific language model pretraining for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2020 · 2007
Cited alongside, same era.
Microsoft research paraphrase corpus
Bill Dolan, Chris Brockett, and Chris Quirk. 2005 · 2008
Cited alongside, same era.
Ramoboost: ranked minority oversampling in boosting
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, H Lehman Li-wei, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. 2016 · 2016
Later among the works it cites.
Biocreative v cdr task corpus: a resource for chemical disease relation extraction
Jiao Li, Yueping Sun, Robin J Johnson, Daniela Sciaky, Chih-Hsuan Wei, Robert Leaman, Allan Peter Davis, Carolyn J Mattingly, Thomas C Wiegers, and Zhiyong Lu. 2016 · 2016
Later among the works it cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sheng Chen, Haibo He, and Edwardo A Garcia. 2010 · 2010
Cited alongside, same era.
Ncbi disease corpus: a resource for disease name recognition and concept normalization
Rezarta Islamaj Doğan, Robert Leaman, and Zhiyong Lu. 2014 · 2014
Cited alongside, same era.
The chemdner corpus of chemicals and drugs and its annotation principles
Martin Krallinger, Obdulia Rabal, Florian Leitner, Miguel Vazquez, David Salgado, Zhiyong Lu, Robert Leaman, Yanan Lu, Donghong Ji, Daniel M Lowe, et al. 2015 · 2015
Cited alongside, same era.
An overview of the bioasq large-scale biomedical semantic indexing and question answering competition
George Tsatsaronis, Georgios Balikas, Prodromos Malakasiotis, Ioannis Partalas, Matthias Zschunke, Michael R Alvers, Dirk Weissenborn, Anastasia Krithara, Sergios Petridis, Dimitris Polychronopoulos, et al. 2015 · 2015
Cited alongside, same era.
Trieu H. Trinh and Quoc V. Le. 2018 · 2018
Later among the works it cites.
Proceedings of the 18th bionlp workshop and shared task
Dina Demner-Fushman, K Bretonnel Cohen, Sophia Ananiadou, and Jun’ichi Tsujii. 2019 · 2019
Later among the works it cites.
Biobert: a pre-trained biomedical language representation model for biomedical text mining
J Lee, W Yoon, S Kim, D Kim, CH So, and J Kang. 2019 · 2019
Later among the works it cites.
Better language models and their implications
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.