Fetching the paper…
Reading the bibliography…
We empirically characterize the performance of discriminative and generative LSTM models for text classification.
On the probabilistic interpretation of neural network classifiers and discriminative training criteria
Ney, Hermann · 1995
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Discriminative vs informative learning
Rubinstein, Y. Dan and Hastie, Trevor · 1997
Earlier work this paper cites.
On discriminative vs. generative classifiers: A comparison of logistic regression and naive Bayes
Ng, Andrew Y. and Jordan, Michael I · 2001
Earlier work this paper cites.
Combining naive Bayes and n n -gram language models for text classification
Peng, Fuchun and Schuurmans, Dale · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Morin, Frederic and Bengio, Yoshua · 2005
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2012
Earlier work this paper cites.
A fast and simple algorithm for training neural probabilistic language models
Mnih, Andriy and Teh, Yee Whye · 2012
Cited alongside, same era.
Glove: Global vectors for word representation
Pennington, Jeffrey, Socher, Richard, and Manning, Christopher D · 2014
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
Jean, Sebastian, Cho, Kyunghyun, Memisevic, Roland, and Bengio, Yoshua · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
Zhang, Xiang, Zhao, Junbo, and LeCun, Yann · 2015
Cited alongside, same era.
Very deep convolutional networks for text classification
Conneau, Alexis, Schwenk, Holger, Barrault, Loic, and Lecun, Yann · 2016
Cited alongside, same era.
Bag of tricks for efficient text classification
Joulin, Armand, Grave, Edouard, Bojanowski, Piotr, and Mikolov, Tomas · 2016
Progressive neural networks
Rusu, Andrei A., Rabinowitz, Neil C., Desjardins, Guillaume, Soyer, Hubert, Kirkpatrick, James, Kavukcuoglu, Koray, Pascanu, Razvan, and Hadsell, Raia · 2016
Later among the works it cites.
Efficient character-level document classification by combining convolution and recurrent layers
Xiao, Yijun and Cho, Kyunghyun · 2016
Later among the works it cites.
Pathnet: Evolution channels gradient descent in super neural networks
Fernando, Chrisantha, Banarse, Dylan, Blundell, Charles, Zwols, Yori, Ha, David, Rusu, Andrei A., Pritzel, Alexander, and Wierstra, Daan · 2017
Closest in time.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, James, Pascanu, Razvan, Rabinowitz, Neil, Veness, Joel, Desjardins, Guillaume, Rusu, Andrei A., Milan, Kieran, Quan, John, Ramalhoa, Tiago, Grabska-Barwinska, Agnieszka, Hassabis, Demis, Clopath, Claudia, Kumaran, Dharshan, and Hadsell, Raia · 2017
Closest in time.
One-vs-each approximation to softmax for scalable estimation of probabilities
Titsias, Michalis K · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, Chiyuan, Bengio, Samy, Hardt, Moritz, Recht, Benjamin, and Vinyals, Oriol · 2017
Closest in time.