Fetching the paper…
Reading the bibliography…
End-to-end (E2E) automatic speech recognition (ASR) models, by now, have shown competitive performance on several benchmarks.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Efficient WFST-based one-pass decoding with on-the-fly hypothesis rescoring in extremely large vocabulary continuous speech recognition,”
T. Hori, C. Hori, Y. Minami, and A. Nakamura, · 2007
Earlier work this paper cites.
“Adam: A method for Stochastic Optimization,”
D. P. Kingma and J. Ba, · 2010
Earlier work this paper cites.
“Sequence Transduction with Recurrent Neural Networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Japanese and Korean Voice Search,”
M. Schuster and K. Nakajima, · 2012
Earlier work this paper cites.
“Attention-Based Models for Speech Recognition,”
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, · 2016
Earlier work this paper cites.
“Highway long short-term memory rnns for distant speech recognition,”
Y. Zhang, G. Chen, D. Yu, K. Yaco, S. Khudanpur, and J. Glass, · 2016
Earlier work this paper cites.
“Training deep bidirectional lstm acoustic model for lvcsr by a context-sensitive-chunk bptt approach,”
K. Chen and Q. Huo, · 2016
Earlier work this paper cites.
“Recent Advances in Google Real-time HMM-driven Unit Selection Synthesizer,”
X. Gonzalvo, S. Tazari, C.-A. Chan, M. Becker, A. Gutkin, and H. Silen, · 2016
Earlier work this paper cites.
“Tensorflow: A System for Large-scale Machine Learning,”
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, et al., · 2016
Earlier work this paper cites.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
S. Kim, T. Hori, and S. Watanabe, · 2017
Earlier work this paper cites.
“Deliberation networks: Sequence generation beyond one-pass decoding,”
Y. Xia, F. Tian, L. Wu, J. Lin, T. Qin, N. Yu, and T.-Y. Liu, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Cited alongside, same era.
“Minimum word error rate training for attention-based sequence-to-sequence models,”
R. Prabhavalkar, T. N. Sainath, Y. Wu, P. Nguyen, Z. Chen, C.-C. Chiu, and A. Kannan, · 2017
Cited alongside, same era.
“Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in Google Home,”
C. Kim, A. Misra, K. Chin, T. Hughes, A. Narayanan, et al., · 2017
Cited alongside, same era.
“In-datacenter Performance Analysis of a Tensor Processing Unit,”
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, et al., · 2017
Cited alongside, same era.
“Toward Domain-Invariant Speech Recognition via Large Scale Training,”
A. Narayanan, A. Misra, K. C. Sim, G. Pundak, A. Tripathi, et al., · 2018
Cited alongside, same era.
“Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling,”
J. Shen, P. Nguyen, Y. Wu, Z. Chen, M. X. Chen, et al., · 2019
Later among the works it cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, et al., · 2020
Closest in time.
“A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,”
T. N. Sainath, Y. He, B. Li, A. Narayanan, R. Pang, et al., · 2020
Closest in time.
“On the comparison of popular end-to-end models for large scale speech recognition,”
J. Li, Y. Wu, Y. Gaur, C. Wang, R. Zhao, and S. Liu, · 2020
Closest in time.
“Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Multi-dialect speech recognition with a single sequence-to-sequence model,”
B. Li, T. N. Sainath, K. C. Sim, M. Bacchiani, E. Weinstein, P. Nguyen, Z. Chen, Y. Wu, and K. Rao, · 2018
Cited alongside, same era.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, · 2019
Cited alongside, same era.
“Transformer-transducer: End-to-end speech recognition with self-attention,”
C.-F. Yeh, J. Mahadeokar, K. Kalgaonkar, Y. Wang, et al., · 2019
Cited alongside, same era.
“Two-pass end-to-end speech recognition,”
T. N. Sainath, R. Pang, D. Rybach, Y. He, R. Prabhavalkar, et al., · 2019
Cited alongside, same era.
“A comparison of end-to-end models for long-form speech recognition,”
C.-C. Chiu, W. Han, Y. Zhang, R. Pang, S. Kishchenko, et al., · 2019
Cited alongside, same era.
“Recurrent stacking of layers for compact neural machine translation models,”
R. Dabre and A. Fujita, · 2019
Cited alongside, same era.
“Recognizing long-form speech using streaming end-to-end models,”
A. Narayanan, R. Prabhavalkar, C.-C. Chiu, D. Rybach, et al., · 2019
Cited alongside, same era.
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, · 2020
Closest in time.
“Deliberation model based two-pass end-to-end speech recognition,”
K. Hu, T. N. Sainath, R. Pang, and R. Prabhavalkar, · 2020
Closest in time.
“Sequence to multi-sequence learning via conditional chain mapping for mixture signals,”
J. Shi, X. Chang, P. Guo, S. Watanabe, Y. Fujita, J. Xu, B. Xu, and L. Xie, · 2020
Closest in time.
“Universal ASR: Unify and improve streaming asr with full-context modeling,”
J. Yu, W. Han, A. Gulati, C.-C. Chiu, B. Li, T. N. Sainath, Y. Wu, and R. Pang, · 2020
Closest in time.
“Transformer transducer: One model unifying streaming and non-streaming speech recognition,”
A. Tripathi, J. Kim, Q. Zhang, H. Lu, and H. Sak, · 2020
Closest in time.
W. Han, Z Zhang, Y. Zhang, J. Yu, C.-C. Chiu, J. Qin, A. Gulati, R. Pang, and Y. Wu, · 2020
Closest in time.
“Towards fast and accurate streaming end-to-end asr,”
B. Li, S.-Y. Chang, T. N. Sainath, R. Pang, Y. He, T. Strohman, and Y. Wu, · 2020
Closest in time.
“FastEmit: Low-latency streaming asr with sequence-level emission regularization,”
J. Yu, C.-C. Chiu, B. Li, S.-Y. Chang, T. N. Sainath, et al., · 2021
Closest in time.