Fetching the paper…
Reading the bibliography…
Recent studies on domain-specific BERT models show that effectiveness on downstream tasks can be improved when models are pretrained on in-domain data.
ClinicalBERT: Modeling clinical notes and predicting hospital readmission
Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. 2019 · 1904
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Levis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Language, context, and text: Aspects of language in a social-semiotic perspective
Michael A.K. Halliday and Ruqaiya Hasan. 1989 · 1989
Earlier work this paper cites.
Introduction to the bio-entity recognition task at JNLPBA
Jin-Dong Kim, Tomoko Ohta, Yoshimasa Tsuruoka, Yuka Tateisi, and Nigel Collier. 2004 · 2004
Earlier work this paper cites.
Bertweet: A pre-trained language model for English Tweets
Dat Quoc Nguyen, Thanh Vu, and Anh Tuan Nguyen. 2020 · 2005
Earlier work this paper cites.
Domain adaptation with structural correspondence learning
John Blitzer, Ryan McDonald, and Fernando Pereira. 2006 · 2006
Earlier work this paper cites.
Corpus-based and knowledge-based measures of text semantic similarity
Rada Mihalcea, Courtney Corley, and Carlo Strapparava. 2006 · 2006
Earlier work this paper cites.
Analysis of representations for domain adaptation
Shai Ben-David, John Blitzer, Koby Crammer, and Fernando Pereira. 2007 · 2007
Earlier work this paper cites.
2010 i2b2/va challenge on concepts, assertions, and relations in clinical text
Özlem Uzuner, Brett R South, Shuying Shen, and Scott L DuVall. 2011 · 2010
Earlier work this paper cites.
Using domain similarity for performance estimation
Vincent Van Asch and Walter Daelemans. 2010 · 2010
Earlier work this paper cites.
Kenlm: Faster and smaller language model queries
Kenneth Heafield. 2011 · 2011
Earlier work this paper cites.
How noisy social media text, how diffrnt social media sources?
Timothy Baldwin, Paul Cook, Marco Lui, Andrew MacKinlay, and Li Wang. 2013 · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
SemEval-2014 task 4: Aspect based sentiment analysis
Maria Pontiki, Dimitris Galanis, John Pavlopoulos, Harris Papageorgiou, Ion Androutsopoulos, and Suresh Manandhar. 2014 · 2014
Cited alongside, same era.
CADEC: A corpus of adverse drug event annotations
Sarvnaz Karimi, Alejandro Metke-Jimenez, Madonna Kemp, and Chen Wang. 2015 · 2015
Cited alongside, same era.
From word embeddings to document distances
Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger. 2015 · 2015
Cited alongside, same era.
PPDB 2.0: Better paraphrase ranking, fine-grained entailment relations, word embeddings, and style classification
Overview of the third social media mining for health (SMM4H) shared tasks at EMNLP 2018
Davy Weissenbacher, Abeed Sarker, Michael J. Paul, and Graciela Gonzalez-Hernandez. 2018 · 2018
Later among the works it cites.
Publicly available clinical BERT embeddings
Emily Alsentzer, John Murphy, William Boag, Wei-Hung Weng, Di Jindi, Tristan Naumann, and Matthew McDermott. 2019 · 2019
Later among the works it cites.
Cloze-driven pretraining of self-attention networks
Alexei Baevski, Sergey Edunov, Yinhan Liu, Luke Zettlemoyer, and Michael Auli. 2019 · 2019
Later among the works it cites.
SciBERT: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 2019
Later among the works it cites.
Using similarity measures to select pretraining data for NER
Xiang Dai, Sarvnaz Karimi, Ben Hachey, and Cecile Paris. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ellie Pavlick, Pushpendre Rastogi, Juri Ganitkevitch, Benjamin Van Durme, and Chris Callison-Burch. 2015 · 2015
Cited alongside, same era.
Broad twitter corpus: A diverse named entity recognition resource
Leon Derczynski, Kalina Bontcheva, and Ian Roberts. 2016 · 2016
Cited alongside, same era.
Learning to select data for transfer learning with bayesian optimization
Sebastian Ruder and Barbara Plank. 2017 · 2017
Cited alongside, same era.
Shot or not: Comparison of NLP approaches for vaccination behaviour detection
Aditya Joshi, Xiang Dai, Sarvnaz Karimi, Ross Sparks, Cecile Paris, and C Raina MacIntyre. 2018 · 2018
Cited alongside, same era.
A corpus with multi-level annotations of patients, interventions and outcomes to support language processing for medical literature
Benjamin Nye, Junyi Jessy Li, Roma Patel, Yinfei Yang, Iain Marshall, Ani Nenkova, and Byron Wallace. 2018 · 2018
Cited alongside, same era.
Improving language understanding with unsupervised learning
Alec Radford, Karthik Narasimhan, Tim Salimans, and Iyya Sutskever. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Iyya Sutskever. 2019 · 2019
Later among the works it cites.
Neural Transfer Learning for Natural Language Processing
Sebastian Ruder. 2019 · 2019
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Closest in time.