Fetching the paper…
Reading the bibliography…
Sequence-to-sequence models with soft attention have been successfully applied to a wide variety of problems, but their decoding process incurs a quadratic time and space cost and is inapplicable to real-time sequence transduction.
Parallel prefix computation
Richard E. Ladner and Michael J. Fischer · 1980
Earlier work this paper cites.
The design for the Wall Street Journal-based CSR corpus
Douglas B. Paul and Janet M. Baker · 1992
Earlier work this paper cites.
Difference equations: an introduction with applications
Walter G. Kelley and Allan C. Peterson · 2001
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Predicting success in machine translation
Alexandra Birch, Miles Osborne, and Philipp Koehn · 2008
Earlier work this paper cites.
Eigen v3
Gaël Guennebaud, Benoıt Jacob, Philip Avery, Abraham Bachrach, Sebastien Barthelemy, et al · 2010
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves · 2012
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Çağlar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Navdeep Jaitly, David Sussillo, Quoc V. Le, Oriol Vinyals, Ilya Sutskever, and Samy Bengio · 2015
Cited alongside, same era.
Segmental recurrent neural networks
Lingpeng Kong, Chris Dyer, and Noah A. Smith · 2015
Cited alongside, same era.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D. Manning · 2015
Cited alongside, same era.
Feed-forward networks with attention can solve some long-term memory problems
Colin Raffel and Daniel P. W. Ellis · 2015
Cited alongside, same era.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
Ramesh Nallapati, Bowen Zhou, Cícero Nogueira dos Santos, Çaglar Gülçehre, and Bing Xiang · 2016
Later among the works it cites.
Lookahead convolution layer for unidirectional recurrent neural networks
Chong Wang, Dani Yogatama, Adam Coates, Tony Han, Awni Hannun, and Bo Xiao · 2016
Later among the works it cites.
Online segment to segment neural transduction
Lei Yu, Jan Buys, and Phil Blunsom · 2016
Later among the works it cites.
Very deep convolutional networks for end-to-end speech recognition
Yu Zhang, William Chan, and Navdeep Jaitly · 2016
Later among the works it cites.
Towards better decoding and language model integration in sequence to sequence models
Jan Chorowski and Navdeep Jaitly · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Cited alongside, same era.
Reinforcement learning neural turing machines
Wojciech Zaremba and Ilya Sutskever · 2015
Cited alongside, same era.
TensorFlow: A system for large-scale machine learning
Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Cited alongside, same era.
Sequence to sequence transduction with hard monotonic attention
Roee Aharoni and Yoav Goldberg · 2016
Cited alongside, same era.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals · 2016
Cited alongside, same era.
Learning online alignments with continuous rewards policy gradient
Yuping Luo, Chung-Cheng Chiu, Navdeep Jaitly, and Ilya Sutskever · 2016
Cited alongside, same era.
Advances in joint CTC-attention based end-to-end speech recognition with a deep CNN encoder and RNN-LM
Takaaki Hori, Shinji Watanabe, Yu Zhang, and William Chan · 2017
Closest in time.
Learning hard alignments with variational inference
Dieterich Lawson, George Tucker, Chung-Cheng Chiu, Colin Raffel, Kevin Swersky, and Navdeep Jaitly · 2017
Closest in time.
A comparison of sequence-to-sequence models for speech recognition
Rohit Prabhavalkar, Kanishka Rao, Tara Sainath, Bo Li, Leif Johnson, and Navdeep Jaitly · 2017
Closest in time.
Online and linear-time attention by enforcing monotonic alignments
Colin Raffel, Minh-Thang Luong, Peter J. Liu, Ron J. Weiss, and Douglas Eck · 2017
Closest in time.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning · 2017
Closest in time.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Sauros · 2017
Closest in time.