Fetching the paper…
Reading the bibliography…
End-to-end models have achieved state-of-the-art results on several automatic speech recognition tasks.
Discriminative training for large vocabulary speech recognition
D. Povey, · 2005
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
A. Graves, A. Mohamed, and G. Hinton, · 2013
Earlier work this paper cites.
“Improvements to the IBM speech activity detection system for the DARPA RATS program,”
S. Thomas, G. Saon, M. V. Segbroeck, and S. Narayanan, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D.P. Kingma and J. Ba, · 2015
Earlier work this paper cites.
“Deep learning-based telephony speech recognition in the wild,”
K. J. Han, S. Hahm, B. Kim, J. Kim, and I. Lane, · 2017
Earlier work this paper cites.
“Optimizing expected word error rate via sampling for speech recognition,”
M. Shannon, · 2017
Earlier work this paper cites.
“In-datacenter performance analysis of a tensor processing unit,”
N. P. Jouppi, C. Young, N. Patil, and et al., · 2017
Earlier work this paper cites.
“The Microsoft 2017 conversational speech recognition system,”
W. Xiong, L. Wu, F. Alleva, J. Droppo, X. Huang, and A. Stolcke, · 2018
Earlier work this paper cites.
“Minimum word error rate training for attention-based sequence-to-sequence models,”
R. Prabhavalkar, T. N. Sainath, Y. Wu, P. Nguyen, Z. Chen, C. Chiu, and A. Kannan, · 2018
Earlier work this paper cites.
“Recognizing long-form speech using streaming end-to-end models,”
A. Narayanan, R. Prabhavalkar, C. Chiu, D. Rybach, T. N. Sainath, and T. Strohman, · 2019
Earlier work this paper cites.
“Improving RNN transducer modeling for end-to-end speech recognition,”
J. Li, R. Zhao, H. Hu, and Y. Gong, · 2019
Cited alongside, same era.
“A comparison of end-to-end models for long-form speech recognition,”
C. Chiu, W. Han, Y. Zhang, R. Pang, S. Kishchenko, P. Nguyen, A. Narayanan, H. Liao, S. Zhang, A. Kannan, R. Prabhavalkar, Z. Chen, T. N. Sainath, and Y. Wu, · 2019
Cited alongside, same era.
“Lingvo: A modular and scalable framework for sequence-to-sequence modeling,”
J. Shen, P. Nguyen, et al., · 2019
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
D. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. Cubuk, and Q. Le, · 2019
Cited alongside, same era.
“Pushing the limits of semi-supervised learning for automatic speech recognition,”
Y. Zhang, J. Qin, D. Park, W. Han, C. Chiu, R. Pang, Q. Le, and Y. Wu, · 2020
Cited alongside, same era.
“Advancing RNN transducer technology for speech recognition,”
G. Saon, Z. Tüske, D. Bolanos, and B. Kingsbury, · 2021
Closest in time.
“Advanced long-context end-to-end speech recognition using context-expanded transformers,”
Takaaki Hori, Niko Moritz, Chiori Hori, and Jonathan Le Roux, · 2021
Closest in time.
Yu Zhang, Daniel S Park, et al., · 2021
Closest in time.
“On minimum word error rate training of the hybrid autoregressive transducer,”
L. Lu, Z. Meng, N. Kanda, J. Li, and Y. Gong, · 2021
Closest in time.
“Reducing exposure bias in training recurrent neural network transducers,”
X. Cui, B. Kingsbury, G. Saon, D. Haws, and Z. Tuske, · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Conformer: Convolution-augmented transformer for speech recognition,”
A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, · 2020
Cited alongside, same era.
“A new training pipeline for an improved neural transducer,”
A. Zeyer, A. Merboldt, R. Schlüter, and H. Ney, · 2020
Cited alongside, same era.
“Efficient minimum word error rate training of RNN-transducer for end-to-end speech recognition,”
J. Guo, G. Tiwari, J. Droppo, M. V. Segbroeck, C. Huang, A. Stolcke, and R. Maas, · 2020
Cited alongside, same era.
“Minimum bayes risk training of RNN-transducer for end-to-end speech recognition,”
C. Weng, C. Yu, J. Cui, C. Zhang, and D. Yu, · 2020
Cited alongside, same era.
“On the limit of English conversational speech recognition,”
Z. Tüske, G. Saon, and B. Kingsbury, · 2021
Cited alongside, same era.
“RNN-T models fail to generalize to out-of-domain audio: Causes and solutions,”
C. Chiu, A. Narayanan, W. Han, R. Prabhavalkar, Y. Zhang, N. Jaitly, R. Pang, T. N. Sainath, P. Nguyen, L. Cao, and Y. Wu, · 2021
Cited alongside, same era.
Q. Li, Y. Zhang, B. Li, L. Cao, and P. C. Woodland, · 2021
Closest in time.
“A lightweight framework for online voice activity detection in the wild,”
X. Xu, H. Dinkel, M. Wu, and K. Yu, · 2021
Closest in time.
“A better and faster end-to-end model for streaming ASR,”
B. Li, A. Gulati, J. Yu, T. N. Sainath, C. Chiu, A. Narayanan, S. Chang, R. Pang, Y. He, J. Qin, W. Han, Q. Liang, Y. Zhang, T. Strohman, and Y. Wu, · 2021
Closest in time.
“Bridging the gap between streaming and non-streaming asr systems by distilling ensembles of CTC and RNN-T models,”
T. Doutre, W. Han, C. Chi, R. Pang, O. Siohan, and L. Cao, · 2021
Closest in time.
“Improving streaming automatic speech recognition with non-streaming model distillation on unsupervised data,”
T. Doutre, W. Han, M. Ma, Z. Lu, C. Chiu, R. Pang, A. Narayanan, A. Misra, Y. Zhang, and L. Cao, · 2021
Closest in time.
“Less is more: Improved RNN-T decoding using limited label context and path merging,”
R. Prabhavalkar, Y. He, D. Rybach, S. Campbell, A. Narayanan, T. Strohman, and T. N. Sainath, · 2021
Closest in time.