Fetching the paper…
Reading the bibliography…
Recurrent neural networks are a powerful tool for modeling sequential data, but the dependence of each timestep's computation on the previous timestep's output limits parallelism and makes RNNs unwieldy for very long sequences.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukás Burget, Jan Cernocký, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Multi-dimensional sentiment analysis with learned representations
Andrew L Maas, Andrew Y Ng, and Christopher Potts · 2011
Earlier work this paper cites.
ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Baselines and bigrams: Simple, good sentiment and topic classification
Sida Wang and Christopher D Manning · 2012
Earlier work this paper cites.
Effective use of word order for text categorization with convolutional neural networks
Rie Johnson and Tong Zhang · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Ensemble of generative and discriminative techniques for sentiment analysis of movie reviews
Grégoire Mesnil, Tomas Mikolov, Marc’Aurelio Ranzato, and Yoshua Bengio · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
M. T. Luong, H. Pham, and C. D. Manning · 2015
Cited alongside, same era.
Predicting polarities of tweets by composing word embeddings with long short-term memory
Xin Wang, Yuanchao Liu, Chengjie Sun, Baoxun Wang, and Xiaolong Wang · 2015
Cited alongside, same era.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Cited alongside, same era.
A C-LSTM neural network for text classification
Chunting Zhou, Chonglin Sun, Zhiyuan Liu, and Francis Lau · 2015
Cited alongside, same era.
Strongly-typed recurrent neural networks
David Balduzzi and Muhammad Ghifary · 2016
Cited alongside, same era.
MetaMind neural machine translation system for WMT 2016
Ask me anything: Dynamic memory networks for natural language processing
Ankit Kumar, Ozan Irsoy, Peter Ondruska, Mohit Iyyer, James Bradbury, Ishaan Gulrajani, Victor Zhong, Romain Paulus, and Richard Socher · 2016
Closest in time.
Fully character-level neural machine translation without explicit segmentation
Jason Lee, Kyunghyun Cho, and Thomas Hofmann · 2016
Closest in time.
A way out of the odyssey: Analyzing and combining recent insights for LSTMs
Shayne Longpre, Sabeek Pradhan, Caiming Xiong, and Richard Socher · 2016
Closest in time.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Closest in time.
Virtual adversarial training for semi-supervised text classification
Takeru Miyato, Andrew M Dai, and Ian Goodfellow · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
James Bradbury and Richard Socher · 2016
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, and Kilian Q Weinberger · 2016
Cited alongside, same era.
Neural machine translation in linear time
Nal Kalchbrenner, Lasse Espeholt, Karen Simonyan, Aaron van den Oord, Alex Graves, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Character-aware neural language models
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M. Rush · 2016
Cited alongside, same era.
Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Hugo Larochelle, Aaron Courville, et al · 2016
Cited alongside, same era.
Chainer: A next-generation open source framework for deep learning
Seiya Tokui, Kenta Oono, and Shohei Hido
Cited in the paper.
Closest in time.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2016
Closest in time.
Pixel recurrent neural networks
Aaron van den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu · 2016
Closest in time.
Sequence-to-sequence learning as beam-search optimization
Sam Wiseman and Alexander M Rush · 2016
Closest in time.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Closest in time.
Efficient character-level document classification by combining convolution and recurrent layers
Yijun Xiao and Kyunghyun Cho · 2016
Closest in time.
Dynamic memory networks for visual and textual question answering
Caiming Xiong, Stephen Merity, and Richard Socher · 2016
Closest in time.