Fetching the paper…
Reading the bibliography…
Attention-based recurrent neural encoder-decoder models present an elegant solution to the automatic speech recognition problem.
“SWITCHBOARD: Telephone speech corpus for research and development,”
John J. Godfrey, Edward C. Holliman, and Jane McDaniel, · 1992
Earlier work this paper cites.
“Bidirectional recurrent neural networks,”
M. Schuster and K.K. Paliwal, · 1997
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“The Fisher Corpus: a Resource for the Next Generations of Speech-to-Text,”
Christopher Cieri David, David Miller, and Kevin Walker, · 2004
Earlier work this paper cites.
“The Kaldi Speech Recognition Toolkit,”
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, …, and Karel Vesely, · 2011
Earlier work this paper cites.
“Dropout improves recurrent neural networks for handwriting recognition,”
V. Pham, T. Bluche, C. Kermorvant, and J. Louradour, · 2014
Earlier work this paper cites.
“On Using Monolingual Corpora in Neural Machine Translation,”
Çaglar Gülçehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loïc Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding.,”
Yajie Miao, Mohammad Gowayyed, and Florian Metze, · 2015
Earlier work this paper cites.
“Scheduled sampling for sequence prediction with recurrent neural networks,”
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer, · 2015
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“TensorFlow: Large-Scale machine learning on heterogeneous systems,” 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, …, and Xiaoqiang Zheng, · 2015
Earlier work this paper cites.
“Fast and accurate recurrent neural network acoustic models for speech recognition,”
Hasim Sak, Andrew Senior, and Francoise Beaufays, · 2015
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals, · 2016
Cited alongside, same era.
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, …, and Jeffrey Dean, · 2016
Cited alongside, same era.
“Improving Neural Machine Translation Models with Monolingual Data,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Cited alongside, same era.
“Neural Machine Translation of Rare Words with Subword Units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Cited alongside, same era.
“Listening while Speaking: Speech Chain by Deep Learning,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2017
Later among the works it cites.
“Exploring Architectures, Data and Units For Streaming End-to-End Speech Recognition with RNN-Transducer,”
Kanishka Rao, Hasim Sak, and Rohit Prabhavalkar, · 2017
Later among the works it cites.
“In-Datacenter Performance Analysis of a Tensor Processing Unit,”
Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, …, and Doe Hyun Yoon, · 2017
Later among the works it cites.
“Improved training of end-to-end attention models for speech recognition,”
Albert Zeyer, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2018
Closest in time.
“State-of-the-art Speech Recognition with Sequence-to-Sequence Models,”
Chung-Cheng Chiu, Tara Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J. Weiss, Kanishka Rao, Katya Gonina, Navdeep Jaitly, Bo Li, Jan Chorowski, and Michiel Bacchiani, · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Rethinking the Inception Architecture for Computer Vision,”
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna, · 2016
Cited alongside, same era.
“Lower frame rate neural network acoustic models,”
Golan Pundak and Tara N. Sainath, · 2016
Cited alongside, same era.
“Towards better decoding and language model integration in sequence to sequence models,”
Jan Chorowski and Navdeep Jaitly, · 2017
Cited alongside, same era.
“Cold Fusion: Training Seq2seq Models Together with Language Models,”
Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh, and Adam Coates, · 2017
Cited alongside, same era.
“Unsupervised Pretraining for Sequence to Sequence Learning,”
Prajit Ramachandran, Peter J. Liu, and Quoc V. Le, · 2017
Cited alongside, same era.
“Exploring neural transducers for end-to-end speech recognition,”
Eric Battenberg, Jitong Chen, Rewon Child, Adam Coates, Yashesh Gaur Yi Li, Hairong Liu, Sanjeev Satheesh, Anuroop Sriram, and Zhenyao Zhu, · 2017
Cited alongside, same era.
“An Analysis of Incorporating an External Language Model into a Sequence-to-Sequence Model,”
Anjuli Kannan, Yonnghui Wu, Patrick Nguyen, Tara N. Sainath, Zhifeng Chen, and Rohit Prabhavalkar, · 2018
Closest in time.
“Multi-Modal Data Augmentation for End-to-End ASR,”
Adithya Renduchintala, Shuoyang Ding, Matthew Wiesner, and Shinji Watanabe, · 2018
Closest in time.
“Deep contextualized word representations,”
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer, · 2018
Closest in time.
“Building competitive direct acoustics-to-word models for English conversational speech recognition,”
Kartik Audhkhasi, Brian Kingsbury, Bhuvana Ramabhadran, George Saon, and Michael Picheny, · 2018
Closest in time.
“Minimum Word Error Rate Training for Attention-based Sequence-to-sequence Models,”
R. Prabhavalkar, T. N. Sainath, Y. Wu, P. Nguyen, Z. Chen, C. Chiu, and A. Kannan, · 2018
Closest in time.