Fetching the paper…
Reading the bibliography…
We introduce a new type of deep contextualized word representation that models both (1) complex characteristics of word use (e.g., syntax and semantics), and (2) how these uses vary across linguistic contexts (i.e., to model polysemy).
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
Using a semantic concordance for sense identification
George A. Miller, Martin Chodorow, Shari Landes, Claudia Leacock, and Robert G. Thomas. 1994 · 1994
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando Pereira. 2001 · 2001
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
The proposition bank: An annotated corpus of semantic roles
Martha Palmer, Paul Kingsbury, and Daniel Gildea. 2005 · 2005
Earlier work this paper cites.
Word representations: A simple and general method for semi-supervised learning
Joseph P. Turian, Lev-Arie Ratinov, and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel P. Kuksa. 2011 · 2011
Earlier work this paper cites.
Conll-2012 shared task: Modeling multilingual unrestricted coreference in ontonotes
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Olga Uryupina, and Yuchen Zhang. 2012 · 2012
Earlier work this paper cites.
Adadelta: An adaptive learning rate method
Matthew D. Zeiler. 2012 · 2012
Earlier work this paper cites.
Easy victories and uphill battles in coreference resolution
Greg Durrett and Dan Klein. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Towards robust linguistic analysis using ontonotes
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Hwee Tou Ng, Anders Björkelund, Olga Uryupina, Yuchen Zhang, and Zhi Zhong. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Y Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2014 · 2014
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Efficient non-parametric estimation of multiple embeddings per word in vector space
Arvind Neelakantan, Jeevan Shankar, Alexandre Passos, and Andrew McCallum. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Semi-supervised sequence learning
Andrew M. Dai and Quoc V. Le. 2015 · 2015
Cited alongside, same era.
An empirical exploration of recurrent network architectures
Rafal Józefowicz, Wojciech Zaremba, and Ilya Sutskever. 2015 · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Finding function in form: Compositional character models for open vocabulary word representation
Wang Ling, Chris Dyer, Alan W. Black, Isabel Trancoso, Ramon Fermandez, Silvio Amir, Luís Marujo, and Tiago Luís. 2015 · 2015
Cited alongside, same era.
Training very deep networks
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber. 2015 · 2015
Cited alongside, same era.
End-to-end learning of semantic role labeling using recurrent neural networks
Jie Zhou and Wei Xu. 2015 · 2015
Cited alongside, same era.
Learning global features for coreference resolution
Sam Wiseman, Alexander M. Rush, and Stuart M. Shieber. 2016 · 2016
Later among the works it cites.
Text classification improved by integrating bidirectional lstm with two-dimensional max pooling
Peng Zhou, Zhenyu Qi, Suncong Zheng, Jiaming Xu, Hongyun Bao, and Bo Xu. 2016 · 2016
Later among the works it cites.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James R. Glass. 2017 · 2017
Later among the works it cites.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Later among the works it cites.
Enhanced lstm for natural language inference
Qian Chen, Xiao-Dan Zhu, Zhen-Hua Ling, Si Wei, Hui Jiang, and Diana Inkpen. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jimmy Ba, Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Cited alongside, same era.
Named entity recognition with bidirectional LSTM-CNNs
Jason Chiu and Eric Nichols. 2016 · 2016
Cited alongside, same era.
Deep reinforcement learning for mention-ranking coreference models
Kevin Clark and Christopher D. Manning. 2016 · 2016
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Cited alongside, same era.
Embeddings for word sense disambiguation: An evaluation study
Ignacio Iacobacci, Mohammad Taher Pilehvar, and Roberto Navigli. 2016 · 2016
Cited alongside, same era.
Exploring the limits of language modeling
Rafal Józefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Cited alongside, same era.
Christopher Clark and Matthew Gardner. 2017 · 2017
Later among the works it cites.
A joint many-task model: Growing a neural network for multiple nlp tasks
Kazuma Hashimoto, Caiming Xiong, Yoshimasa Tsuruoka, and Richard Socher. 2017 · 2017
Later among the works it cites.
Deep semantic role labeling: What works and what’s next
Luheng He, Kenton Lee, Mike Lewis, and Luke S. Zettlemoyer. 2017 · 2017
Later among the works it cites.
End-to-end neural coreference resolution
Kenton Lee, Luheng He, Mike Lewis, and Luke S. Zettlemoyer. 2017 · 2017
Later among the works it cites.
Stochastic answer networks for machine reading comprehension
Xiaodong Liu, Yelong Shen, Kevin Duh, and Jianfeng Gao. 2017 · 2017
Later among the works it cites.
Learned in translation: Contextualized word vectors
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher. 2017 · 2017
Later among the works it cites.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom. 2017 · 2017
Later among the works it cites.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2017 · 2017
Later among the works it cites.
Neural tree indexers for text understanding
Tsendsuren Munkhdalai and Hong Yu. 2017 · 2017
Later among the works it cites.
Semi-supervised sequence tagging with bidirectional language models
Matthew E. Peters, Waleed Ammar, Chandra Bhagavatula, and Russell Power. 2017 · 2017
Later among the works it cites.
Improving sequence to sequence learning with unlabeled data
Prajit Ramachandran, Peter Liu, and Quoc Le. 2017 · 2017
Later among the works it cites.
Bidirectional attention flow for machine comprehension
Min Joon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi. 2017 · 2017
Later among the works it cites.
Gated self-matching networks for reading comprehension and question answering
Wenhui Wang, Nan Yang, Furu Wei, Baobao Chang, and Ming Zhou. 2017 · 2017
Later among the works it cites.
Natural language inference over interaction space
Yichen Gong, Heng Luo, and Jian Zhang. 2018 · 2018
Closest in time.