Fetching the paper…
Reading the bibliography…
Recurrent neural network models with an attention mechanism have proven to be extremely effective on a wide variety of sequence-to-sequence problems.
Parallel prefix computation
Ladner, Richard E. and Fischer, Michael J · 1980
Earlier work this paper cites.
The design for the Wall Street Journal-based CSR corpus
Paul, Douglas B. and Baker, Janet M · 1992
Earlier work this paper cites.
DARPA TIMIT acoustic-phonetic continous speech corpus
Garofolo, John S., Lamel, Lori F., Fisher, William M., Fiscus, Jonathon G., and Pallett, David S · 1993
Earlier work this paper cites.
Continuous sigmoidal belief networks trained using slice sampling
Frey, Brendan J · 1997
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, Alex, Fernández, Santiago, Gomez, Faustino, and Schmidhuber, Jürgen · 2006
Earlier work this paper cites.
Semantic hashing
Salakhutdinov, Ruslan and Hinton, Geoffrey · 2009
Earlier work this paper cites.
Eigen v3
Guennebaud, Gaël, Jacob, Benoıt, Avery, Philip, Bachrach, Abraham, Barthelemy, Sebastien, et al · 2010
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Graves, Alex · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Graves, Alex · 2013
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Graves, Alex, Mohamed, Abdel-rahman, and Hinton, Geoffrey · 2013
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Yoshua, Léonard, Nicholas, and Courville, Aaron · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Cho, Kyunghyun, van Merriënboer, Bart, Gülçehre, Çağlar, Bahdanau, Dzmitry, Bougares, Fethi, Schwenk, Holger, and Bengio, Yoshua · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, Junyoung, Gulcehre, Caglar, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Earlier work this paper cites.
Graves, Alex, Wayne, Greg, and Danihelka, Ivo · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc V · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, Diederik and Ba, Jimmy · 2014
Earlier work this paper cites.
Dropout improves recurrent neural networks for handwriting recognition
Pham, Vu, Bluche, Théodore, Kermorvant, Christopher, and Louradour, Jérôme · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc V · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2015
Cited alongside, same era.
The IWSLT 2015 evaluation campaign
Cettolo, Mauro, Niehues, Jan, Stüker, Sebastian, Bentivogli, Luisa, Cattoni, Roldano, and Federico, Marcello · 2015
Cited alongside, same era.
Attention-based models for speech recognition
Chorowski, Jan, Bahdanau, Dzmitry, Serdyuk, Dmitriy, Cho, Kyunghyun, and Bengio, Yoshua · 2015
Cited alongside, same era.
Jaitly, Navdeep, Sussillo, David, Le, Quoc V., Vinyals, Oriol, Sutskever, Ilya, and Bengio, Samy · 2015
Cited alongside, same era.
Segmental recurrent neural networks
Kong, Lingpeng, Dyer, Chris, and Smith, Noah A · 2015
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Foerster, Jakob, Assael, Yannis M., de Freitas, Nando, and Whiteson, Shimon · 2016
Later among the works it cites.
Categorical reparameterization with gumbel-softmax
Jang, Eric, Gu, Shixiang, and Poole, Ben · 2016
Later among the works it cites.
Text summarization with TensorFlow
Liu, Peter J. and Pan, Xin · 2016
Later among the works it cites.
Learning online alignments with continuous rewards policy gradient
Luo, Yuping, Chiu, Chung-Cheng, Jaitly, Navdeep, and Sutskever, Ilya · 2016
Later among the works it cites.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, Chris J., Mnih, Andriy, and Teh, Yee Whye · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stanford neural machine translation systems for spoken language domain
Luong, Minh-Thang and Manning, Christopher D · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Luong, Minh-Thang, Pham, Hieu, and Manning, Christopher D · 2015
Cited alongside, same era.
A neural attention model for abstractive sentence summarization
Rush, Alexander M., Chopra, Sumit, and Weston, Jason · 2015
Cited alongside, same era.
End-to-end memory networks
Sukhbaatar, Sainbayar, Szlam, Arthur, Weston, Jason, and Fergus, Rob · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Xu, Kelvin, Ba, Jimmy, Kiros, Ryan, Cho, Kyunghyun, Courville, Aaron, Salakhudinov, Ruslan, Zemel, Rich, and Bengio, Yoshua · 2015
Cited alongside, same era.
Reinforcement learning neural turing machines
Zaremba, Wojciech and Sutskever, Ilya · 2015
Cited alongside, same era.
Learning to transduce with unbounded memory
Grefenstette, Edward, Hermann, Karl Moritz, Suleyman, Mustafa, and Blunsom, Phil · 2015
Cited alongside, same era.
Language as a latent variable: Discrete generative models for sentence compression
Miao, Yishu and Blunsom, Phil · 2016
Later among the works it cites.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
Nallapati, Ramesh, Zhou, Bowen, dos Santos, Cícero Nogueira, Gülçehre, Çaglar, and Xiang, Bing · 2016
Later among the works it cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, Tim and Kingma, Diederik P · 2016
Later among the works it cites.
Lookahead convolution layer for unidirectional recurrent neural networks
Wang, Chong, Yogatama, Dani, Coates, Adam, Han, Tony, Hannun, Awni, and Xiao, Bo · 2016
Later among the works it cites.
Efficient summarization with read-again and copy mechanism
Zeng, Wenyuan, Luo, Wenjie, Fidler, Sanja, and Urtasun, Raquel · 2016
Later among the works it cites.
Very deep convolutional networks for end-to-end speech recognition
Zhang, Yu, Chan, William, and Jaitly, Navdeep · 2016
Later among the works it cites.
TensorFlow: A system for large-scale machine learning
Abadi, Martin, Barham, Paul, Chen, Jianmin, Chen, Zhifeng, Davis, Andy, Dean, Jeffrey, Devin, Matthieu, Ghemawat, Sanjay, Irving, Geoffrey, Isard, Michael, Kudlur, Manjunath, Levenberg, Josh, Monga, Rajat, Moore, Sherry, Murray, Derek G., Steiner, Benoit, Tucker, Paul, Vasudevan, Vijay, Warden, Pete, Wicke, Martin, Yu, Yuan, and Zheng, Xiaoqiang · 2016
Later among the works it cites.
Adaptive computation time for recurrent neural networks
Graves, Alex · 2016
Later among the works it cites.
Text summarization with TensorFlow
Liu, Peter J. and Pan, Xin · 2016
Later among the works it cites.
Towards better decoding and language model integration in sequence to sequence models
Chorowski, Jan and Jaitly, Navdeep · 2017
Closest in time.
Kim, Yoon, Denton, Carl, Hoang, Luong, and Rush, Alexander M · 2017
Closest in time.
Training a subsampling mechanism in expectation
Raffel, Colin and Lawson, Dieterich · 2017
Closest in time.
Cutting-off redundant repeating generations for neural abstractive summarization
Suzuki, Jun and Nagata, Masaaki · 2017
Closest in time.
Towards better decoding and language model integration in sequence to sequence models
Chorowski, Jan and Jaitly, Navdeep · 2017
Closest in time.