Fetching the paper…
Reading the bibliography…
Despite the tremendous empirical success of neural models in natural language processing, many of them lack the strong intuitions that accompany classical machine learning approaches.
On the definition of a family of automata
Marcel Paul Schützenberger. 1961 · 1961
Earlier work this paper cites.
Statistical inference for probabilistic functions of finite state Markov chains
Leonard E. Baum and Ted Petrie. 1966 · 1966
Earlier work this paper cites.
Semirings, Automata, Languages
Werner Kuich and Arto Salomaa, editors. 1986 · 1986
Earlier work this paper cites.
Rational Series and Their Languages
Jean Berstel, Jr. and Christophe Reutenauer. 1988 · 1988
Earlier work this paper cites.
Finite state automata and simple recurrent networks
Axel Cleeremans, David Servan-Schreiber, and James L. McClelland. 1989 · 1989
Earlier work this paper cites.
Serial order: A parallel, distributed processing approach
Michael I. Jordan. 1989 · 1989
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman. 1990 · 1990
Earlier work this paper cites.
An efficient gradient-based algorithm for online training of recurrent network trajectories
Ronald J. Williams and Jing Peng. 1990 · 1990
Earlier work this paper cites.
Learning and extracting finite state automata with second-order recurrent neural networks
C. Lee Giles, Clifford B Miller, Dong Chen, Hsing-Hen Chen, Guo-Zheng Sun, and Yee-Chun Lee. 1992 · 1992
Earlier work this paper cites.
Fool’s gold: Extracting finite state machines from recurrent network dynamics
John F. Kolen. 1993 · 1993
Earlier work this paper cites.
Multilayer feedforward networks with a nonpolynomial activation function can approximate any function
Moshe Leshno and Shimon Schocken. 1993 · 1993
Earlier work this paper cites.
Building a large annotated corpus of English: The Penn Treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. 1993 · 1993
Earlier work this paper cites.
On the computational power of neural nets
Hava T. Siegelmann and Eduardo D. Sontag. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Gradient-based Learning Applied to Document Recognition
Yann LeCun. 1998 · 1998
Earlier work this paper cites.
Hidden markov model interpretations of neural networks
Ingmar Visser, Maartje EJ Raijmakers, and Peter CM Molenaar. 2001 · 2001
Earlier work this paper cites.
Parameter estimation for probabilistic finite-state transducers
Jason Eisner. 2002 · 2002
Earlier work this paper cites.
Weighted finite-state transducers in speech recognition
Mehryar Mohri, Fernando Pereira, and Michael Riley. 2002 · 2002
Earlier work this paper cites.
A weighted finite state transducer implementation of the alignment template model for statistical machine translation
Shankar Kumar and William Byrne. 2003 · 2003
Earlier work this paper cites.
Rational kernels: Theory and algorithms
Corinna Cortes, Patrick Haffner, and Mehryar Mohri. 2004 · 2004
Earlier work this paper cites.
Mining and summarizing customer reviews
Minqing Hu and Bing Liu. 2004 · 2004
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
Bo Pang and Lillian Lee. 2004 · 2004
Earlier work this paper cites.
Modeling form for on-line following of musical performances
Bryan Pardo and William Birmingham. 2005 · 2005
Cited alongside, same era.
Juicer: A weighted finite-state transducer speech decoder
Darren Moore, John Dines, Mathew Magimai-Doss, Jithendra Vepa, Octavian Cheng, and Thomas Hain. 2006 · 2006
Cited alongside, same era.
Rational and recognisable power series
Jacques Sakarovitch. 2009 · 2009
Cited alongside, same era.
Enhanced sentiment learning using twitter hashtags and smileys
Dmitry Davidov, Oren Tsur, and Ari Rappoport. 2010 · 2010
Cited alongside, same era.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukás Burget, Jan Cernocký, and Sanjeev Khudanpur. 2010 · 2010
Cited alongside, same era.
A Non-parametric Model for the Discovery of Inflectional Paradigms from Plain Text Using Graphical Models over Strings
Markus Dreyer. 2011 · 2011
Semi-supervised question retrieval with gated convolutions
Tao Lei, Hrishikesh Joshi, Regina Barzilay, Tommi Jaakkola, Kateryna Tymoshenko, Alessandro Moschitti, and Lluís Màrquez. 2016 · 2016
Later among the works it cites.
Simplifying long short-term memory acoustic models for fast training and decoding
Yajie Miao, Jinyu Li, Yongqiang Wang, Shi-Xiong Zhang, and Yifan Gong. 2016 · 2016
Later among the works it cites.
Weighting finite-state transductions with neural context
Pushpendre Rastogi, Ryan Cotterell, and Jason Eisner. 2016 · 2016
Later among the works it cites.
Quasi-recurrent neural network
James Bradbury, Stephen Merity, Caiming Xiong, and Richard Socher. 2017 · 2017
Later among the works it cites.
Capacity and trainability in recurrent neural networks
Jasmine Collins, Jascha Sohl-Dickstein, and David Sussillo. 2017 · 2017
Later among the works it cites.
Intelligible language modeling with input switched affine networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Hidden factors and hidden topics: understanding rating dimensions with review text
Julian McAuley and Jure Leskovec. 2013 · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Learning longer memory in recurrent neural networks
Tomas Mikolov, Armand Joulin, Sumit Chopra, Michaël Mathieu, and Marc’Aurelio Ranzato. 2014 · 2014
Cited alongside, same era.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Cited alongside, same era.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals. 2014 · 2014
Cited alongside, same era.
Jakob N. Foerster, Justin Gilmer, Jan Chorowski, Jascha Sohl-Dickstein, and David Sussillo. 2017 · 2017
Later among the works it cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann Dauphin. 2017 · 2017
Later among the works it cites.
Hafez: an interactive poetry generation system
Marjan Ghazvininejad, Xing Shi, Jay Priyadarshi, and Kevin Knight. 2017 · 2017
Later among the works it cites.
Kenton Lee, Omer Levy, and Luke Zettlemoyer. 2017 · 2017
Later among the works it cites.
Adversarial ranking for language generation
Kevin Lin, Dianqi Li, Xiaodong He, Zhengyou Zhang, and Ming-Ting Sun. 2017 · 2017
Later among the works it cites.
Deep multitask learning for semantic dependency parsing
Hao Peng, Sam Thomson, and Noah A. Smith. 2017 · 2017
Later among the works it cites.
Using the output embedding to improve language models
Ofir Press and Lior Wolf. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V. Le. 2017 · 2017
Later among the works it cites.
Recurrent neural networks as weighted language recognizers
Yining Chen, Sorcha Gilroy, Kevin Knight, and Jonathan May. 2018 · 2018
Closest in time.
Simple recurrent units for highly parallelizable recurrence
Tao Lei, Yu Zhang, Sida I. Wang, Hui Dai, and Yoav Artzi. 2018 · 2018
Closest in time.
Independently recurrent neural network (IndRNN): Building A longer and deeper RNN
Shuai Li, Wanqing Li, Chris Cook, Ce Zhu, and Yanbo Gao. 2018 · 2018
Closest in time.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom. 2018 · 2018
Closest in time.
Backpropagating through structured argmax using a spigot
Hao Peng, Sam Thomson, and Noah A. Smith. 2018 · 2018
Closest in time.
SoPa: Bridging CNNs, RNNs, and weighted finite-state machines
Roy Schwartz, Sam Thomson, and Noah A. Smith. 2018 · 2018
Closest in time.
On the practical computational power of finite precision RNNs for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Closest in time.