Course in general linguistics
Ferdinand de Saussure. 1916 · 1916
Earlier work this paper cites.
The significance of letter position in word recognition
G.E. Rawlinson. 1976 · 1976
Earlier work this paper cites.
Multi-style training for robust isolated-word speech recognition
Richard Lippmann, Edward Martin, and D. Paul. 1987 · 1987
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Leon Bottou, Yoshua Bengio, and Patrick Haffner. 1998 · 1998
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira. 2001 · 2001
Earlier work this paper cites.
Raeding wrods with jubmled lettres there is a cost
K. Rayner, S.J. White, R.L. Johnson, and S.P. Liversedge. 2006 · 2006
Earlier work this paper cites.
Recurrent neural networks are universal approximators
Anton Maximilian Schäfer and Hans Georg Zimmermann. 2006 · 2006
Earlier work this paper cites.
Part-of-speech tagging for twitter: Annotation, features, and experiments
Kevin Gimpel, Nathan Schneider, Brendan O’Connor, Dipanjan Das, Daniel Mills, Jacob Eisenstein, Michael Heilman, Dani Yogatama, Jeffrey Flanigan, and Noah A. Smith. 2011 · 2011
Earlier work this paper cites.
The manifold tangent classifier
Salah Rifai, Yann Dauphin, Pascal Vincent, Yoshua Bengio, and Xavier Muller. 2011 · 2011
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Original
Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2012 · 2012
Earlier work this paper cites.