Fetching the paper…
Reading the bibliography…
Common recurrent neural architectures scale poorly due to the intrinsic difficulty in parallelizing their state computations.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Learning question classifiers
Xin Li and Dan Roth. 2002 · 2002
Earlier work this paper cites.
Mining and summarizing customer reviews
Minqing Hu and Bing Liu. 2004 · 2004
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
Bo Pang and Lillian Lee. 2004 · 2004
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee. 2005 · 2005
Earlier work this paper cites.
Annotating expressions of opinions and emotions in language
Janyce Wiebe, Theresa Wilson, and Claire Cardie. 2005 · 2005
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Fast dropout training
Sida Wang and Christopher Manning. 2013 · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Opinion mining with deep recurrent neural networks
Ozan Irsoy and Claire Cardie. 2014 · 2014
Earlier work this paper cites.
A convolutional neural network for modelling sentences
Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
The DCU-ICTCAS MT system at WMT 2014 on german-english translation task
Liangyou Li, Xiaofeng Wu, Santiago Cortes Vaillo, Jun Xie, Andy Way, and Qun Liu. 2014 · 2014
Earlier work this paper cites.
The RWTH aachen german-english machine translation system for wmt 2014
Stephan Peitz, Joern Wuebker, Markus Freitag, and Hermann Ney. 2014 · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals. 2014 · 2014
Earlier work this paper cites.
Deep convolutional networks are hierarchical kernel machines
Fabio Anselmi, Lorenzo Rosasco, Cheston Tan, and Tomaso A. Poggio. 2015 · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
On using very large target vocabulary for neural machine translation
Sébastien Jean, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
A simple way to initialize recurrent networks of rectified linear units
Quoc V. Le, Navdeep Jaitly, and Geoffrey E. Hinton. 2015 · 2015
Earlier work this paper cites.
Molding cnns for text: non-linear, non-consecutive convolutions
Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2015 · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Training very deep networks
Rupesh K Srivastava, Klaus Greff, and Jürgen Schmidhuber. 2015 · 2015
Cited alongside, same era.
Self-adaptive hierarchical sentence model
Han Zhao, Zhengdong Lu, and Pascal Poupart. 2015 · 2015
Cited alongside, same era.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V. Le. 2016 · 2015
Cited alongside, same era.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann Dauphin. 2017 · 2017
Closest in time.
Accurate, large minibatch SGD: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Closest in time.
Lstm: A search space odyssey
Klaus Greff, Rupesh Kumar Srivastava, Jan Koutnx00EDk, Bas R. Steunebrink, and Jx00FCrgen Schmidhuber. 2017 · 2017
Closest in time.
Deep semantic role labeling: What works and what’s next
Luheng He, Kenton Lee, Mike Lewis, and Luke Zettlemoyer. 2017 · 2017
Closest in time.
Opennmt: Open-source toolkit for neural machine translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander Rush. 2017 · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jeremy Appleyard, Tomás Kociský, and Phil Blunsom. 2016 · 2016
Cited alongside, same era.
Strongly-typed recurrent neural networks
David Balduzzi and Muhammad Ghifary. 2016 · 2016
Cited alongside, same era.
Hierarchical multiscale recurrent neural networks
Junyoung Chung, Sungjin Ahn, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer. 2016 · 2016
Cited alongside, same era.
Persistent rnns: Stashing recurrent weights on-chip
Greg Diamos, Shubho Sengupta, Bryan Catanzaro, Mike Chrzanowski, Adam Coates, Erich Elsen, Jesse Engel, Awni Hannun, and Sanjeev Satheesh. 2016 · 2016
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Oleksii Kuchaiev and Boris Ginsburg. 2017 · 2017
Closest in time.
Kenton Lee, Omer Levy, and Luke S. Zettlemoyer. 2017 · 2017
Closest in time.
Deriving neural architectures from sequence and graph kernels
Tao Lei, Wengong Jin, Regina Barzilay, and Tommi Jaakkola. 2017 · 2017
Closest in time.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom. 2017 · 2017
Closest in time.
Mapping instructions and visual observations to actions with reinforcement learning
Dipendra Misra, John Langford, and Yoav Artzi. 2017 · 2017
Closest in time.
Fast-slow recurrent neural networks
Asier Mujika, Florian Meier, and Angelika Steger. 2017 · 2017
Closest in time.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Closest in time.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Closest in time.
Gated self-matching networks for reading comprehension and question answering
Wenhui Wang, Nan Yang, Furu Wei, Baobao Chang, and Ming Zhou. 2017 · 2017
Closest in time.
A sensitivity analysis of (and practitioners’ guide to) convolutional neural networks for sentence classification
Yingjie Zhang and Byron C. Wallace. 2017 · 2017
Closest in time.
Recurrent highway networks
Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber. 2017 · 2017
Closest in time.
Recurrent neural networks as weighted language recognizers
Yining Chen, Sorcha Gilroy, Kevin Knight, and Jonathan May. 2018 · 2018
Closest in time.
An analysis of neural language modeling at multiple scales
Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2018 · 2018
Closest in time.
Rational recurrences
Hao Peng, Roy Schwartz, Sam Thomson, and Noah A. Smith. 2018 · 2018
Closest in time.
Situated mapping of sequential instructions to actions with single-step reward observation
Alane Suhr and Yoav Artzi. 2018 · 2018
Closest in time.
Learning to map context-dependent sentences to executable formal queries
Alane Suhr, Srinivasan Iyer, and Yoav Artzi. 2018 · 2018
Closest in time.