Fetching the paper…
Reading the bibliography…
In the last few years, an emerging trend in automatic speech recognition research is the study of end-to-end (E2E) systems.
“Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Rectified linear units improve restricted boltzmann machines,”
Vinod Nair and Geoffrey E Hinton, · 2010
Earlier work this paper cites.
“Sequence Transduction with Recurrent Neural Networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Towards End-to-End Speech Recognition with Recurrent Neural Networks,”
A. Graves and N. Jaitley, · 2014
Earlier work this paper cites.
“Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation,”
K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, · 2014
Earlier work this paper cites.
“Empirical evaluation of gated recurrent neural networks on sequence modeling,”
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Learning Acoustic Frame Labeling for Speech Recognition with Recurrent Neural Networks,”
H. Sak, A. Senior, K. Rao, O. Irsoy, A. Graves, F. Beaufays, and J. Schalkwyk, · 2015
Earlier work this paper cites.
“EESEN: End-to-End Speech Recognition using Deep RNN Models and WFST-based Decoding,”
Y. Miao, M. Gowayyed, and F. Metze, · 2015
Earlier work this paper cites.
“Neural Machine Translation by Jointly Learning to Align and Translate,”
D. Bahdanau, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“Attention-Based Models for Speech Recognition,”
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“Convolutional, long short-term memory, fully connected deep neural networks,”
Tara N Sainath, Oriol Vinyals, Andrew Senior, and Haşim Sak, · 2015
Earlier work this paper cites.
“LSTM time and frequency recurrence for automatic speech recognition,”
J. Li, A. Mohamed, G. Zweig, and Y. Gong, · 2015
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Earlier work this paper cites.
“Neural Speech Recognizer: Acoustic-to-word LSTM Model for Large Vocabulary Speech Recognition,”
H. Soltau, H. Liao, and H. Sak, · 2016
Cited alongside, same era.
“Exploring multidimensional LSTMs for large vocabulary ASR,”
J. Li, A. Mohamed, G. Zweig, and Y. Gong, · 2016
Cited alongside, same era.
“Modeling time-frequency patterns with LSTM vs. convolutional architectures for LVCSR tasks,”
Tara N Sainath and Bo Li, · 2016
Cited alongside, same era.
“A prioritized grid long short-term memory RNN for speech recognition,”
Wei-Ning Hsu, Yu Zhang, and James Glass, · 2016
Cited alongside, same era.
“Multidimensional residual learning based on recurrent neural networks for acoustic modeling,”
Yuanyuan Zhao, Shuang Xu, and Bo Xu, · 2016
Cited alongside, same era.
“Residual LSTM: Design of a deep recurrent architecture for distant speech recognition,”
Jaeyoung Kim, Mostafa El-Khamy, and Jungwon Lee, · 2017
Later among the works it cites.
“Highway-LSTM and recurrent highway networks for speech recognition,”
Golan Pundak and Tara N Sainath, · 2017
Later among the works it cites.
“Towards Discriminatively-trained HMM-based End-to-end models for Automatic Speech Recognition,”
Hossein Hadian, Hossein Sameti, Daniel Povey, and Sanjeev Khudanpur, · 2018
Later among the works it cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Katya Gonina, et al., · 2018
Later among the works it cites.
“Improving the performance of online neural transducer models,”
Tara N Sainath, Chung-Cheng Chiu, Rohit Prabhavalkar, Anjuli Kannan, Yonghui Wu, Patrick Nguyen, and ZhiJeng Chen, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu Zhang, Guoguo Chen, Dong Yu, Kaisheng Yao, Sanjeev Khudanpur, and James Glass, · 2016
Cited alongside, same era.
“Deep speech 2: End-to-end speech recognition in English and Mandarin,”
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al., · 2016
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton, · 2016
Cited alongside, same era.
“Simplifying long short-term memory acoustic models for fast training and decoding,”
Y. Miao, J. Li, Y. Wang, S. Zhang, and Y. Gong, · 2016
Cited alongside, same era.
“A Comparison of Sequence-to-Sequence Models for Speech Recognition,”
R. Prabhavalkar, K. Rao, T. N. Sainath, B. Li, L. Johnson, and N. Jaitly, · 2017
Cited alongside, same era.
“Exploring Neural Transducers for End-to-End Speech Recognition,”
E. Battenberg, J. Chen, R. Child, A. Coates, Y. Gaur, Y. Li, H. Liu, S. Satheesh, D. Seetapun, A. Sriram, et al., · 2017
Cited alongside, same era.
“Recurrent neural aligner: An encoder-decoder neural network model for sequence to sequence mapping,”
Hasim Sak, Matt Shannon, Kanishka Rao, and Françoise Beaufays, · 2017
Cited alongside, same era.
“Advancing acoustic-to-word CTC model,”
Jinyu Li, Guoli Ye, Amit Das, Rui Zhao, and Yifan Gong, · 2018
Later among the works it cites.
“Advancing connectionist temporal classification with attention modeling,”
A. Das, J. Li, R. Zhao, and Y. Gong, · 2018
Later among the works it cites.
“Efficient implementation of recurrent neural network transducer in tensorflow,”
Tom Bagby, Kanishka Rao, and Khe Chai Sim, · 2018
Later among the works it cites.
“Layer trajectory LSTM,”
Jinyu Li, Changliang Liu, and Yifan Gong, · 2018
Later among the works it cites.
“Exploring layer trajectory LSTM with depth processing units and attention,”
Jinyu Li, Liang Lu, Changliang Liu, and Yifan Gong, · 2018
Later among the works it cites.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al., · 2019
Closest in time.
“Triggered attention for end-to-end speech recognition,”
Niko Moritz, Takaaki Hori, and Jonathan Le Roux, · 2019
Closest in time.
“Advancing acoustic-to-word CTC model with attention and mixed-units,”
Amit Das, Jinyu Li, Guoli Ye, Rui Zhao, and Yifan Gong, · 2019
Closest in time.
“Improving layer trajectory LSTM with future context frames,”
Jinyu Li, Liang Lu, Changliang Liu, and Yifan Gong, · 2019
Closest in time.