Fetching the paper…
Reading the bibliography…
In recent years, all-neural end-to-end approaches have obtained state-of-the-art results on several challenging automatic speech recognition (ASR) tasks.
“Long Short-Term Memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Practical variational inference for neural networks,”
A. Graves, · 2011
Earlier work this paper cites.
“Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups,”
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, · 2012
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Japanese and Korean Voice Search,”
M. Schuster and K. Nakajima, · 2012
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
A. Graves, A. r. Mohamed, and G. Hinton, · 2013
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
“Deep speech: Scaling Up End-to-End Speech Recognition,”
A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coates, et al., · 2014
Earlier work this paper cites.
“Attention-Based Models for Speech Recognition,”
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“Librispeech: An asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Acoustic modelling with cd-ctc-smbr lstm rnns,”
A. Senior, H. Sak, F. de Chaumont Quitry, T. Sainath, and K. Rao, · 2015
Cited alongside, same era.
“Librispeech: An ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Cited alongside, same era.
“Deep Speech 2: End-to-End Speech Recognition in English and Mandarin,”
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen, et al., · 2016
Cited alongside, same era.
“Listen, Attend and Spell: A Neural Network for Large Vocabulary Conversational Speech Recognition,”
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, · 2016
Cited alongside, same era.
“On the compression of recurrent neural networks with an application to lvcsr acoustic modeling for embedded speech recognition,”
Rohit Prabhavalkar, Ouais Alsharif, Antoine Bruguier, and Ian McGraw, · 2016
Cited alongside, same era.
“State-of-the-art speech recognition with sequence-to-sequence models,”
C-C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, E. Gonina, N. Jaitly, B. Li, J. Chorowski, and M. Bacchiani, · 2018
Later among the works it cites.
“Monotonic chunkwise attention,”
C.-C. Chiu and C. Raffel, · 2018
Later among the works it cites.
“Toward domain-invariant speech recognition via large scale training,”
A. Narayanan, A. Misra, K. C. Sim, G. Pundak, A. Tripathi, M. Elfeky, P. Haghani, T. Strohman, and M. Bacchiani, · 2018
Later among the works it cites.
“Compression of end-to-end models,”
R. Pang, T. Sainath, R. Prabhavalkar, S. Gupta, Y. Wu, S. Zhang, and C.-C. Chiu, · 2018
Later among the works it cites.
“SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,”
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, · 2019
Later among the works it cites.
“Rwth asr systems for librispeech: Hybrid vs attention,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Recent Advances in Google Real-time HMM-driven Unit Selection Synthesizer,”
X. Gonzalvo, S. Tazari, C.-A. Chan, M. Becker, A. Gutkin, and H. Silen, · 2016
Cited alongside, same era.
“Neural Speech Recognizer: Acoustic-to-Word LSTM Model for Large Vocabulary Speech Recognition,”
H. Soltau, H. Liao, and H. Sak, · 2017
Cited alongside, same era.
“Direct Acoustics-to-Word Models for English Conversational Speech Recognition,”
K. Audhkhasi, B. Ramabhadran, G. Saon, M. Picheny, and D. Nahamoo, · 2017
Cited alongside, same era.
“Hybrid ctc/attention architecture for end-to-end speech recognition,”
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, · 2017
Cited alongside, same era.
“Exploring Architectures, Data and Units for Streaming End-to-End Speech Recognition with RNN-Transducer,”
K. Rao, H. Sak, and R. Prabhavalkar, · 2017
Cited alongside, same era.
“Neural speech recognizer: Acoustic-to-word lstm model for large vocabulary speech recognition,”
H. Soltau, H. Liao, and H. Sak, · 2017
Cited alongside, same era.
C. Lüscher, E. Beck, K. Irie, M. Kitza, Michel W, A. Zeyer, R. Schlüter, and H. Ney, · 2019
Later among the works it cites.
“Recognizing Long-Form Speech Using Streaming End-to-End Models,”
A. Narayanan, R. Prabhavalkar, C.C. Chiu, D. Rybach, T.N. Sainath, and T. Strohman, · 2019
Later among the works it cites.
“A Comparison of End-to-end Models for Long-form Speech Recognition,”
C.-C. Chiu, W. Han, Y. Zhang, R. Pang, S. Kishchenko, P. Nguyen, A. Narayanan, H. Liao, S. Zhang, A. Kannan, R. Prabhavalkar, Z. Chen, T. Sainath, and Y. Wu, · 2019
Later among the works it cites.
“Lingvo: a modular and scalable framework for sequence-to-sequence modeling,” 2019
J. Shen, P. Nguyen, Y. Wu, Z. Chen, and et al., · 2019
Later among the works it cites.
“A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,”
Tara N. Sainath, Yanzhang He, Bo Li, Arun Narayanan, Ruoming Pang, Antoine Bruguier, Shuo-yiin Chang, Wei Li, Raziel Alvarez, Zhifeng Chen, and et al., · 2020
Closest in time.
“Specaugment on large scale datasets,”
Daniel S. Park, Yu Zhang, Chung-Cheng Chiu, Youzheng Chen, Bo Li, William Chan, Quoc V. Le, and Yonghui Wu, · 2020
Closest in time.