Fetching the paper…
Reading the bibliography…
Neural language representation models such as BERT pre-trained on large-scale corpora can well capture rich semantic patterns from plain text, and be fine-tuned to consistently improve the performance of various NLP tasks.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik F Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2006 · 2006
Earlier work this paper cites.
Proceedings of the fourth international workshop on semantic evaluations (semeval-2007)
Eneko Agirre, Llu’is M‘arquez, and Richard Wicentowski. 2007 · 2007
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Ronan Collobert and Jason Weston. 2008 · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. 2008 · 2008
Earlier work this paper cites.
Tagme: on-the-fly annotation of short text fragments (by wikipedia entities)
Paolo Ferragina and Ugo Scaiella. 2010 · 2010
Earlier work this paper cites.
Word representations: a simple and general method for semi-supervised learning
Joseph Turian, Lev Ratinov, and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
The Winograd schema challenge
Hector J Levesque, Ernest Davis, and Leora Morgenstern. 2011 · 2011
Earlier work this paper cites.
Representing general relational knowledge in conceptnet 5
Robert Speer and Catherine Havasi. 2012 · 2012
Earlier work this paper cites.
Translating embeddings for modeling multi-relational data
Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Knowledge graph and text jointly embedding
Zhen Wang, Jianwen Zhang, Jianlin Feng, and Zheng Chen. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Semi-supervised sequence learning
Andrew M Dai and Quoc V Le. 2015 · 2015
Earlier work this paper cites.
Design challenges for entity linking
Xiao Ling, Sameer Singh, and Daniel S Weld. 2015 · 2015
Cited alongside, same era.
Representing text for joint embedding of text and knowledge bases
Kristina Toutanova, Danqi Chen, Patrick Pantel, Hoifung Poon, Pallavi Choudhury, and Michael Gamon. 2015 · 2015
Cited alongside, same era.
Distant supervision for relation extraction via piecewise convolutional neural networks
Daojian Zeng, Kang Liu, Yubo Chen, and Jun Zhao. 2015 · 2015
Cited alongside, same era.
Joint representation learning of text and knowledge for knowledge graph completion
Xu Han, Zhiyuan Liu, and Maosong Sun. 2016 · 2016
Cited alongside, same era.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Cited alongside, same era.
Ultra-fine entity typing
Eunsol Choi, Omer Levy, Yejin Choi, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Later among the works it cites.
Mem2seq: Effectively incorporating knowledge bases into end-to-end task-oriented dialog systems
Andrea Madotto, Chien-Sheng Wu, and Pascale Fung. 2018 · 2018
Later among the works it cites.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018 · 2018
Later among the works it cites.
Knowledgeable reader: Enhancing cloze-style reading comprehension with external commonsense knowledge
Todor Mihaylov and Anette Frank. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural relation extraction with selective attention over instances
Yankai Lin, Shiqi Shen, Zhiyuan Liu, Huanbo Luan, and Maosong Sun. 2016 · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
An attentive neural architecture for fine-grained entity type classification
Sonse Shimaoka, Pontus Stenetorp, Kentaro Inui, and Sebastian Riedel. 2016 · 2016
Cited alongside, same era.
Joint learning of the embedding of words and entities for named entity disambiguation
Ikuya Yamada, Hiroyuki Shindo, Hideaki Takeda, and Yoshiyasu Takefuji. 2016 · 2016
Cited alongside, same era.
Bridge text and knowledge by learning multi-prototype entity mention embedding
Yixin Cao, Lifu Huang, Heng Ji, Xu Chen, and Juanzi Li. 2017 · 2017
Cited alongside, same era.
Semi-supervised sequence tagging with bidirectional language models
Matthew Peters, Waleed Ammar, Chandra Bhagavatula, and Russell Power. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amapreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
Put it back: Entity typing with language model enhancement
Ji Xin, Hao Zhu, Xu Han, Zhiyuan Liu, and Maosong Sun. 2018 · 2018
Later among the works it cites.
Adaptive knowledge sharing in multi-task learning: Improving low-resource neural machine translation
Poorya Zaremoodi, Wray Buntine, and Gholamreza Haffari. 2018 · 2018
Later among the works it cites.
Swag: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Later among the works it cites.
Graph convolution over pruned dependency trees improves relation extraction
Yuhao Zhang, Peng Qi, and Christopher D Manning. 2018 · 2018
Later among the works it cites.
Improving question answering by commonsense-based pre-training
Wanjun Zhong, Duyu Tang, Nan Duan, Ming Zhou, Jiahai Wang, and Jian Yin. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
Ernie: Enhanced representation through knowledge integration
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu. 2019 · 2019
Closest in time.