Fetching the paper…
Reading the bibliography…
Contextualized word embeddings such as ELMo and BERT provide a foundation for strong performance across a wide range of natural language processing tasks by pretraining on large corpora of unlabeled text.
BioBERT: pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2019 · 1901
Earlier work this paper cites.
Learning and evaluating general linguistic intelligence
Dani Yogatama, Cyprien de Masson d’Autume, Jerome Connor, Tomas Kocisky, Mike Chrzanowski, Lingpeng Kong, Angeliki Lazaridou, Wang Ling, Lei Yu, Chris Dyer, et al. 2019 · 1901
Earlier work this paper cites.
SciBERT: Pretrained contextualized embeddings for scientific text
Iz Beltagy, Arman Cohan, and Kyle Lo. 2019 · 1903
Earlier work this paper cites.
Unsupervised data augmentation
Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc V. Le. 2019 · 1904
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen. 1989 · 1989
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. 1993 · 1993
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Penn-Helsinki Parsed Corpus of Early Modern English
Anthony Kroch, Beatrice Santorini, and Ariel Diertani. 2004 · 2004
Earlier work this paper cites.
Becoming Wikipedian: transformation of participation in a collaborative online encyclopedia
Susan L Bryant, Andrea Forte, and Amy Bruckman. 2005 · 2005
Earlier work this paper cites.
Domain adaptation for statistical classifiers
Hal Daumé III and Daniel Marcu. 2006 · 2006
Earlier work this paper cites.
Part-of-speech tagging for middle English through alignment and projection of parallel diachronic texts
Taesun Moon and Jason Baldridge. 2007 · 2007
Earlier work this paper cites.
Vard2: A tool for dealing with spelling variation in historical corpora
Alistair Baron and Paul Rayson. 2008 · 2008
Earlier work this paper cites.
What’s being said near “Martha”? Exploring name entities in literary text collections
Romain Vuillemot, Tanya Clement, Catherine Plaisant, and Amit Kumar. 2009 · 2009
Earlier work this paper cites.
Quantitative analysis of culture using millions of digitized books
Jean-Baptiste Michel, Yuan Kui Shen, Aviva Presser Aiden, Adrian Veres, Matthew K Gray, Joseph P Pickett, Dale Hoiberg, Dan Clancy, Peter Norvig, Jon Orwant, et al. 2011 · 2011
Earlier work this paper cites.
Automatically constructing a normalisation dictionary for microblogs
Bo Han, Paul Cook, and Timothy Baldwin. 2012 · 2012
Earlier work this paper cites.
How noisy social media text, how diffrnt social media sources
Timothy Baldwin, Paul Cook, Marco Lui, Andrew MacKinlay, and Li Wang. 2013 · 2013
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013 · 2013
Cited alongside, same era.
What to do about bad language on the internet
Jacob Eisenstein. 2013 · 2013
Cited alongside, same era.
Supporting exploratory text analysis in literature study
Aditi Muralidharan and Marti A Hearst. 2013 · 2013
Cited alongside, same era.
Semi-supervised sequence learning
Andrew M Dai and Quoc V Le. 2015 · 2015
Cited alongside, same era.
Unsupervised multi-domain adaptation with feature embeddings
Yi Yang and Jacob Eisenstein. 2015 · 2015
Model transfer for tagging low-resource languages using a bilingual dictionary
Meng Fang and Trevor Cohn. 2017 · 2017
Later among the works it cites.
Adversarial adaptation of synthetic or stale data
Young-Bum Kim, Karl Stratos, and Dongchan Kim. 2017 · 2017
Later among the works it cites.
Variational recurrent adversarial deep domain adaptation
Sanjay Purushotham, Wilka Carvalho, Tanachat Nilanon, and Yan Liu. 2017 · 2017
Later among the works it cites.
Domain adaptation with adversarial training and graph embeddings
Firoj Alam, Shafiq Joty, and Muhammad Imran. 2018 · 2018
Later among the works it cites.
Stylistic variation over 200 years of court proceedings according to gender and social class
Stefania Degaetano-Ortlieb. 2018 · 2018
Later among the works it cites.
Word embeddings quantify 100 years of gender and ethnic stereotypes
Nikhil Garg, Londa Schiebinger, Dan Jurafsky, and James Zou. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Cited alongside, same era.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky. 2016 · 2016
Cited alongside, same era.
Quantitative approaches to diachronic corpus linguistics
Martin Hilpert and Stefan Th Gries. 2016 · 2016
Cited alongside, same era.
Bidirectional LSTM for named entity recognition in twitter messages
Nut Limsopatham and Nigel Collier. 2016 · 2016
Cited alongside, same era.
How transferable are neural networks in NLP applications?
Lili Mou, Zhao Meng, Rui Yan, Ge Li, Yan Xu, Lu Zhang, and Zhi Jin. 2016 · 2016
Cited alongside, same era.
Results of the wnut16 named entity recognition shared task
Benjamin Strauss, Bethany Toma, Alan Ritter, Marie-Catherine De Marneffe, and Wei Xu. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Later among the works it cites.
Neural adaptation layers for cross-domain named entity recognition
Bill Yuchen Lin and Wei Lu. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Evaluating historical text normalization systems: How well do they generalize?
Alexander Robertson and Sharon Goldwater. 2018 · 2018
Later among the works it cites.
Pivot based language modeling for improved neural domain adaptation
Yftah Ziser and Roi Reichart. 2018 · 2018
Later among the works it cites.
Using similarity measures to select pretraining data for NER
Xiang Dai, Sarvnaz Karimi, Ben Hachey, and Cecile Paris. 2019 · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
To tune or not to tune? adapting pretrained representations to diverse tasks
Matthew E. Peters, Sebastian Ruder, and Noah A. Smith. 2019 · 2019
Closest in time.