Fetching the paper…
Reading the bibliography…
As one of the most popular sequence-to-sequence modeling approaches for speech recognition, the RNN-Transducer has achieved evolving performance with more and more sophisticated neural network models of growing size and increasing training epochs.
E. H. L. Aarts and J. H. M. Korst,
1990
Earlier work this paper cites.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone Speech Corpus for Research and Development,” in
1992
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,”
1997
Earlier work this paper cites.
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber, “Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks,” in
2006
Earlier work this paper cites.
R. Schlüter, I. Bezrukov, H. Wagner, and H. Ney, “Gammatone Features and Feature Combination for Large Vocabulary Speech Recognition,” in
2007
Earlier work this paper cites.
Y. Bengio, J. Louradour, R. Collobert, and J. Weston, “Curriculum learning,” in
2009
Earlier work this paper cites.
2012
Earlier work this paper cites.
K. Veselý, A. Ghoshal, L. Burget, and D. Povey, “Sequence-discriminative training of deep neural networks,” in
2013
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Estève, “Enhancing the TED-LIUM Corpus with Selected Data for Language Modeling and More TED Talks,” in
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” in
2015
Earlier work this paper cites.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-End Attention-based Large Vocabulary Speech Recognition,” in
2016
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, Attend and Spell: A Neural Network for Large Vocabulary Conversational Speech Recognition,” in
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the Inception Architecture for Computer Vision,” in
2016
Earlier work this paper cites.
T. Lin, P. Goyal, R. B. Girshick, K. He, and P. Dollár, “Focal Loss for Dense Object Detection,” in
2017
Cited alongside, same era.
A. Tripathi, H. Lu, H. Sak, and H. Soltau, “Monotonic Recurrent Neural Network Transducer and Decoding Strategies,” in
2019
Cited alongside, same era.
E. McDermott, H. Sak, and E. Variani, “A Density Ratio Approach to Language Model Fusion in End-to-End Automatic Speech Recognition,” in
2019
Cited alongside, same era.
L. N. Smith and N. Topin, “Super-Convergence: Very Fast Training of Neural Networks using Large Learning Rates,” in
2019
Cited alongside, same era.
B. Zoph, C.-C. Chiu, D. S. Park, E. D. Cubuk, Q. V. Le, W. Chan, and Y. Zhang, “SpecAugment: A Simple Augmentation Method for Automatic Speech Recognition,” in
2019
Cited alongside, same era.
W. Zhou, W. Michel, K. Irie, M. Kitza, R. Schlüter, and H. Ney, “The RWTH ASR system for TED-LIUM release 2: Improving Hybrid HMM with SpecAugment,” in
2020
Later among the works it cites.
D. S. Park et al., “Specaugment on Large Scale Datasets,” in
2020
Later among the works it cites.
Y. Wang et al., “Transformer-Based Acoustic Modeling for Hybrid Speech Recognition,” in
2020
Later among the works it cites.
F. Zhang, Y. Wang, X. Zhang, C. Liu, Y. Saraf, and G. Zweig, “Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces,” in
2020
Later among the works it cites.
Z. Tüske, G. Saon, K. Audhkhasi, and B. Kingsbury, “Single Headed Attention based Sequence-to-sequence Model for State-of-the-Art Results on Switchboard,” in
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Irie, A. Zeyer, R. Schlüter, and H. Ney, “Language Modeling with Deep Transformers,” in
2019
Cited alongside, same era.
K. Irie, A. Zeyer, R. Schlüter, and H. Ney, “Training Language Models for Long-Span Cross-Sentence Evaluation,” in
2019
Cited alongside, same era.
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, “Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss,” in
2020
Cited alongside, same era.
W. Han, Z. Zhang, Y. Zhang, J. Yu, C. Chiu, J. Qin, A. Gulati, R. Pang, and Y. Wu, “ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context,” in
2020
Cited alongside, same era.
A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for Speech Recognition,” in
2020
Cited alongside, same era.
M. Ghodsi, X. Liu, J. Apfel, R. Cabrera, and E. Weinstein, “Rnn-Transducer with Stateless Prediction Network,” in
2020
Cited alongside, same era.
J. Guo, G. Tiwari, J. Droppo, M. V. Segbroeck, C. Huang, A. Stolcke, and R. Maas, “Efficient Minimum Word Error Rate Training of RNN-Transducer for End-to-End Speech Recognition,” in
2020
Cited alongside, same era.
W. Zhou, S. Berger, R. Schlüter, and H. Ney, “Phoneme Based Neural Transducer for Large Vocabulary Speech Recognition,” in
2021
Later among the works it cites.
R. Prabhavalkar, Y. He, D. Rybach, S. Campbell, A. Narayanan, T. Strohman, and T. N. Sainath, “Less is More: Improved RNN-T Decoding Using Limited Label Context and Path Merging,” in
2021
Later among the works it cites.
Z. Meng, Y. Wu, N. Kanda, L. Lu, X. Chen, G. Ye, E. Sun, J. Li, and Y. Gong, “Minimum Word Error Rate Training with Language Model Fusion for End-to-End Speech Recognition,” in
2021
Later among the works it cites.
Z. Meng, S. Parthasarathy, E. Sun, Y. Gaur, N. Kanda, L. Lu, X. Chen, R. Zhao, J. Li, and Y. Gong, “Internal Language Model Estimation for Domain-Adaptive End-to-End Speech Recognition,” in
2021
Later among the works it cites.
G. Saon, Z. Tüske, D. Bolaños, and B. Kingsbury, “Advancing RNN Transducer Technology for Speech Recognition,” in
2021
Later among the works it cites.
W. Zhou, M. Zeineldeen, Z. Zheng, R. Schlüter, and H. Ney, “Acoustic Data-Driven Subword Modeling for End-to-End Speech Recognition,” in
2021
Later among the works it cites.
M. Zeineldeen, A. Glushko, W. Michel, A. Zeyer, R. Schlüter, and H. Ney, “Investigating Methods to Improve Language Model Integration for Attention-based Encoder-Decoder ASR Models,” in
2021
Later among the works it cites.
M. Zeineldeen, J. Xu, C. Lüscher, W. Michel, A. Gerstenberger, R. Schlüter, and H. Ney, “Conformer-based Hybrid ASR System for Switchboard Dataset ,” in
2022
Closest in time.
W. Zhou, Z. Zheng, R. Schlüter, and H. Ney, “On Language Model Integration for RNN Transducer based Speech Recognition,” in
2022
Closest in time.
Z. Tüske, G. Saon, and B. Kingsbury, “On the Limit of English Conversational Speech Recognition,” in
2066
Closest in time.