Fetching the paper…
Reading the bibliography…
Pre-trained word embeddings learned from unlabeled text have become a standard component of neural network architectures for NLP tasks.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Combining labeled and unlabeled data with co-training
Avrim Blum and Tom Mitchell. 1998 · 1998
Earlier work this paper cites.
Text classification from labeled and unlabeled documents using em
Kamal Nigam, Andrew Kachites McCallum, Sebastian Thrun, and Tom Mitchell. 2000 · 2000
Earlier work this paper cites.
Introduction to the CoNLL-2000 shared task chunking
Erik F. Tjong Kim Sang and Sabine Buchholz. 2000 · 2000
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando Pereira. 2001 · 2001
Earlier work this paper cites.
Limitations of co-training for natural language learning from large datasets
David Pierce and Claire Cardie. 2001 · 2001
Earlier work this paper cites.
Discriminative training methods for hidden markov models: Theory and experiments with perceptron algorithms
Michael Collins. 2002 · 2002
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin. 2003 · 2003
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
A high-performance semi-supervised learning method for text chunking
Rie Kubota Ando and Tong Zhang. 2005 · 2005
Earlier work this paper cites.
Semi-supervised sequence modeling with syntactic topic models
Wei Li and Andrew McCallum. 2005 · 2005
Earlier work this paper cites.
Syntax-based semi-supervised named entity tagging
Behrang Mohit and Rebecca Hwa. 2005 · 2005
Earlier work this paper cites.
Semi-supervised structured output learning based on a hybrid generative and discriminative approach
Jun Suzuki, Akinori Fujino, and Hideki Isozaki. 2007 · 2007
Earlier work this paper cites.
Semi-supervised sequential labeling and segmentation using giga-word scale unlabeled data
Jun Suzuki and Hideki Isozaki. 2008 · 2008
Earlier work this paper cites.
Statistical machine translation
Philipp Koehn. 2009 · 2009
Earlier work this paper cites.
Design challenges and misconceptions in named entity recognition
Lev-Arie Ratinov and Dan Roth. 2009 · 2009
Cited alongside, same era.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur. 2010 · 2010
Cited alongside, same era.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel P. Kuksa. 2011 · 2011
Cited alongside, same era.
Bidirectional language model for handwriting recognition
Volkmar Frinken, Alicia Fornés, Josep Lladós, and Jean-Marc Ogier. 2012 · 2012
Cited alongside, same era.
Recurrent continuous translation models
Nal Kalchbrenner and Phil Blunsom. 2013 · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Later among the works it cites.
Skip-thought vectors
Jamie Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S. Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Later among the works it cites.
Joint entity recognition and disambiguation
Gang Luo, Xiaojiang Huang, Chin-Yew Lin, and Zaiqing Nie. 2015 · 2015
Later among the works it cites.
A bidirectional recurrent neural language model for machine translation
Álvaro Peris and Francisco Casacuberta. 2015 · 2015
Later among the works it cites.
Named entity recognition with bidirectional LSTM-CNNs
Jason Chiu and Eric Nichols. 2016 · 2016
Later among the works it cites.
A joint many-task model: Growing a neural network for multiple nlp tasks
Kazuma Hashimoto, Caiming Xiong, Yoshimasa Tsuruoka, and Richard Socher. 2016 · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Semi-supervised learning and domain adaptation in natural language processing
Anders Søgaard. 2013 · 2013
Cited alongside, same era.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, and Phillipp Koehn. 2014 · 2014
Cited alongside, same era.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Fast and robust neural network joint models for statistical machine translation
Jacob Devlin, Rabih Zbib, Zhongqiang Huang, Thomas Lamar, Richard M Schwartz, and John Makhoul. 2014 · 2014
Cited alongside, same era.
Distributed representations of sentences and documents
Quoc V. Le and Tomas Mikolov. 2014 · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Cited alongside, same era.
Later among the works it cites.
Learning distributed representations of sentences from unlabelled data
Felix Hill, Kyunghyun Cho, and Anna Korhonen. 2016 · 2016
Later among the works it cites.
Exploring the limits of language modeling
Rafal Józefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu. 2016 · 2016
Later among the works it cites.
Neural architectures for named entity recognition
Guillaume Lample, Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and Chris Dyer. 2016 · 2016
Later among the works it cites.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Later among the works it cites.
End-to-end sequence labeling via bi-directional LSTM-CNNs-CRF
Xuezhe Ma and Eduard H. Hovy. 2016 · 2016
Later among the works it cites.
context2vec: Learning generic context embedding with bidirectional lstm
Oren Melamud, Jacob Goldberger, and Ido Dagan. 2016 · 2016
Later among the works it cites.
Deep multi-task learning with low level tasks supervised at lower layers
Anders Søgaard and Yoav Goldberg. 2016 · 2016
Later among the works it cites.
The AI2 system at SemEval-2017 Task 10 (ScienceIE): semi-supervised end-to-end entity and relation extraction
Waleed Ammar, Matthew E. Peters, Chandra Bhagavatula, and Russell Power. 2017 · 2017
Closest in time.
Transfer learning for sequence tagging with hierarchical recurrent networks
Zhilin Yang, Ruslan Salakhutdinov, and William W. Cohen. 2017 · 2017
Closest in time.