Fetching the paper…
Reading the bibliography…
We propose a unified neural network architecture and learning algorithm that can be applied to various natural language processing tasks including: part-of-speech tagging, chunking, named entity recognition, and semantic role labeling.
Prediction and entropy of printed english
C. E. Shannon · 1951
Earlier work this paper cites.
Three models for the description of language
N. Chomsky · 1956
Earlier work this paper cites.
Mathematical Structures of Language
Z. S. Harris · 1968
Earlier work this paper cites.
Continuous speech recognition by statistical methods
F. Jelinek · 1976
Earlier work this paper cites.
A convergent gambling estimate of the entropy of english
T. Cover and R. King · 1978
Earlier work this paper cites.
An algorithm for suffix stripping
M. F. Porter · 1980
Earlier work this paper cites.
A learning scheme for asymmetric threshold networks
Y. LeCun · 1985
Earlier work this paper cites.
Learning internal representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Learning sets of filters using back-propagation
D. C. Plaut and G. E. Hinton · 1987
Earlier work this paper cites.
Learning quickly when irrelevant attributes abound: A new linear-threshold algorithm
N. Littlestone · 1988
Earlier work this paper cites.
Probabilistic Reasoning in Intelligent Systems
J. Pearl · 1988
Earlier work this paper cites.
Phoneme recognition using time-delay neural networks
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K.J. Lang · 1989
Earlier work this paper cites.
Probabilistic interpretation of feedforward classification network outputs, with relationships to statistical pattern recognition
J. S. Bridle · 1990
Earlier work this paper cites.
Stochastic gradient learning in neural networks
L. Bottou · 1991
Earlier work this paper cites.
A framework for the cooperation of learning algorithms
L. Bottou and P. Gallinari · 1991
Earlier work this paper cites.
Symbolic-neural systems and the use of hints for developing complex systems
S. C. Suddarth and A. D. C. Holden · 1991
Earlier work this paper cites.
Distributional part-of-speech tagging
H. Schütze · 1995
Earlier work this paper cites.
Bayesian Learning for Neural Networks
R. M. Neal · 1996
Earlier work this paper cites.
A maximum entropy model for part-of-speech tagging
A. Ratnaparkhi · 1996
Earlier work this paper cites.
The entropy of english using ppm-based models
W. J. Teahan and J. G. Cleary · 1996
Earlier work this paper cites.
Global training of document processing systems using graph transformer networks
L. Bottou, Y. LeCun, and Yoshua Bengio · 1997
Earlier work this paper cites.
Multitask Learning
R. Caruana · 1997
Earlier work this paper cites.
Online algorithms and stochastic approximations
L. Bottou · 1998
Earlier work this paper cites.
Learning to order things
W. W. Cohen, R. E. Schapire, and Y. Singer · 1998
Earlier work this paper cites.
Gradient based learning applied to document recognition
Y. Le Cun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 1998
Earlier work this paper cites.
Head-Driven Statistical Models for Natural Language Parsing
M. Collins · 1999
Earlier work this paper cites.
Transductive inference for text classification using support vector machines
T. Joachims · 1999
Earlier work this paper cites.
A maximum-entropy-inspired parser
E. Charniak · 2000
Earlier work this paper cites.
Use of support vector learning for chunk identification
T. Kudoh and Y. Matsumoto · 2000
Cited alongside, same era.
A novel use of statistical parsing to extract information from text
S. Miller, H. Fox, L. Ramshaw, and R. Weischedel · 2000
Cited alongside, same era.
A neural probabilistic language model
Y. Bengio and R. Ducharme · 2001
Cited alongside, same era.
Dependency networks for inference, collaborative filtering, and data visualization
D. Heckerman, D. M. Chickering, C. Meek, R. Rounthwaite, and C. Kadie · 2001
Cited alongside, same era.
Chunking with support vector machines
T. Kudo and Y. Matsumoto · 2001
Cited alongside, same era.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
J. Lafferty, A. McCallum, and F. Pereira · 2001
Cited alongside, same era.
Voting between multiple data representations for text chunking
H. Shen and A. Sarkar · 2005
Later among the works it cites.
Contrastive estimation: Training log-linear models on unlabeled data
N. A. Smith and J. Eisner · 2005
Later among the works it cites.
Joint parsing and semantic role labeling
C. Sutton and A. McCallum · 2005
Later among the works it cites.
Semi-Supervised Learning
O. Chapelle, B. Schölkopf, and A. Zien · 2006
Later among the works it cites.
A fast learning algorithm for deep belief nets
G. E. Hinton, S. Osindero, and Y.-W. Teh · 2006
Later among the works it cites.
Effective self-training for parsing
D. McClosky, E. Charniak, and M. Johnson · 2006
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Automatic labeling of semantic roles
D. Gildea and D. Jurafsky · 2002
Cited alongside, same era.
The necessity of parsing for predicate argument recognition
D. Gildea and M. Palmer · 2002
Cited alongside, same era.
Natural language grammar induction using a constituent-context model
D. Klein and C. D. Manning · 2002
Cited alongside, same era.
Connectionist language modeling for large vocabulary continuous speech recognition
H. Schwenk and J. L. Gauvain · 2002
Cited alongside, same era.
Named entity recognition with a maximum entropy approach
H. L. Chieu · 2003
Cited alongside, same era.
Named entity recognition through classifier combination
R. Florian, A. Ittycheriah, H. Jing, and T. Zhang · 2003
Cited alongside, same era.
G. Musillo and P. Merlo · 2006
Later among the works it cites.
The BellKor solution to the Netflix Prize
R. M. Bell, Y. Koren, and C. Volinsky · 2007
Later among the works it cites.
Greedy layer-wise training of deep networks
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle · 2007
Later among the works it cites.
Learning to rank with nonsmooth cost functions
C. J. C. Burges, R. Ragno, and Quoc Viet Le · 2007
Later among the works it cites.
Ranking the best instances
S. Clémençon and N. Vayatis · 2007
Later among the works it cites.
Three new graphical models for statistical language modelling
A Mnih and G. E. Hinton · 2007
Later among the works it cites.
A discriminative language model with pseudo-negative samples
D. Okanohara and J. Tsujii · 2007
Later among the works it cites.
Using Corpus Statistics on Entities to Improve Semi-supervised Relation Extraction from the Web
B. Rosenfeld and R. Feldman · 2007
Later among the works it cites.
Guided learning for bidirectional sequence classification
L. Shen, G. Satta, and A. K. Joshi · 2007
Later among the works it cites.
Dynamic Conditional Random Fields: Factorized Probabilistic Models for Labeling and Segmenting Sequence Data
C. Sutton, A. McCallum, and K. Rohanimanesh · 2007
Later among the works it cites.
Transductive learning for statistical machine translation
N. Ueffing, G. Haffari, and A. Sarkar · 2007
Later among the works it cites.
Simple semi-supervised dependency parsing
T. Koo, X. Carreras, and M. Collins · 2008
Later among the works it cites.
Structure compilation: trading structure for features
P. Liang, H. Daumé, III, and D. Klein · 2008
Later among the works it cites.
Modeling latent-dynamic in shallow parsing: a latent conditional model with improved inference
X. Sun, L.-P. Morency, D. Okanohara, and J. Tsujii · 2008
Later among the works it cites.
Semi-supervised sequential labeling and segmentation using giga-word scale unlabeled data
J. Suzuki and H. Isozaki · 2008
Later among the works it cites.
Deep learning via semi-supervised embedding
J. Weston, F. Ratle, and R. Collobert · 2008
Later among the works it cites.
Curriculum learning
Y. Bengio, J. Louradour, R. Collobert, and J. Weston · 2009
Later among the works it cites.
Distributional representations for handling sparsity in supervised sequence-labeling
F. Huang and A. Yates · 2009
Later among the works it cites.
Phrase clustering for discriminative learning
D. Lin and X. Wu · 2009
Later among the works it cites.
Design challenges and misconceptions in named entity recognition
L. Ratinov and D. Roth · 2009
Later among the works it cites.
Word representations: A simple and general method for semi-supervised learning
J. Turian, L. Ratinov, and Y. Bengio · 2010
Later among the works it cites.
The proposition bank: An annotated corpus of semantic roles
M. Palmer, D. Gildea, and P. Kingsbury · 2017
Closest in time.