Fetching the paper…
Reading the bibliography…
Although end-to-end (E2E) automatic speech recognition (ASR) has shown state-of-the-art recognition accuracy, it tends to be implicitly biased towards the training data distribution which can degrade generalisation.
1990
Earlier work this paper cites.
S. J. Young, J. J. Odell, and P. C. Woodland, “Tree-based state tying for high accuracy modelling,” in Proc. HLT , 1994
1994
Earlier work this paper cites.
P. Gage, “A new algorithm for data compression,” The C Users Journal , vol. 12, no. 02, pp. 23–38, 1994
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in Proc. ICML , 2006
2006
Earlier work this paper cites.
G. E. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 1, pp. 30–42, 2012
2012
Earlier work this paper cites.
G. E. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. W. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition,” IEEE Signal Processing Magazine , vol. 29, p. 82, 2012
2012
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” ArXiv , vol. abs/1211.3711, 2012
2012
Earlier work this paper cites.
K. Veselý, A. Ghoshal, L. Burget, and D. Povey, “Sequence-discriminative training of deep neural networks,” in Proc. Interspeech , 2013
2013
Earlier work this paper cites.
A. Graves, A. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in Proc. ICASSP , 2013
2013
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Estève, “Enhancing the TED-LIUM corpus with selected data for language modeling and more TED talks,” in Proc. LREC , 2014
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in Proc. ICASSP , 2015
2015
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in Proc. NeurIPS , 2015
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. NeurIPS , 2017
2017
Earlier work this paper cites.
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid CTC/attention architecture for end-to-end speech recognition,” IEEE Journal of Selected Topics in Signal Processing , vol. 11, no. 8, pp. 1240–1253, 2017
2017
Earlier work this paper cites.
A. Sriram, H. Jun, S. Satheesh, and A. Coates, “Cold Fusion: Training Seq2Seq models together with language models,” in Proc. Interspeech , 2018
2018
Earlier work this paper cites.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. E. Y. Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai, “ESPnet: End-to-end speech processing toolkit,” in Proc. Interspeech , 2018
2018
Earlier work this paper cites.
D. Wang, X. Wang, and S. Lv, “An overview of end-to-end automatic speech recognition,” Symmetry , vol. 11, p. 1018, 2019
2019
Earlier work this paper cites.
H. Miao, G. Cheng, P. Zhang, T. Li, and Y. Yan, “Online hybrid CTC/attention architecture for end-to-end speech recognition,” in Proc. Interspeech , 2019
2019
Earlier work this paper cites.
J. Li, Y. Wu, Y. Gaur, C. Wang, R. Zhao, and S. Liu, “On the comparison of popular end-to-end models for large scale speech recognition,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
M. Ghodsi, X. Liu, J. Apfel, R. Cabrera, and E. Weinstein, “RNN-transducer with stateless prediction network,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
L. Dong and B. Xu, “CIF: Continuous integrate-and-fire for end-to-end speech recognition,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
C. Wang, Y. Wu, L. Lu, S. Liu, J. Li, G. Ye, and M. Zhou, “Low latency end-to-end streaming speech recognition with a scout network,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for speech recognition,” in Proc. Interspeech , 2020
W.-N. Hsu, A. Sriram, A. Baevski, T. Likhomanenko, Q. Xu, V. Pratap, J. Kahn, A. Lee, R. Collobert, G. Synnaeve, and M. Auli, “Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training,” in Proc. Interspeech , 2021
2021
Later among the works it cites.
Y. Shi, V. Nagaraja, C. Wu, J. Mahadeokar, D. Le, R. Prabhavalkar, A. Xiao, C.-F. Yeh, J. Chan, C. Fuegen, O. Kalinli, and M. L. Seltzer, “Dynamic encoder transducer: A flexible solution for trading off accuracy for latency,” in Proc. Interspeech , 2021
2021
Later among the works it cites.
J. Li, “Recent advances in end-to-end automatic speech recognition,” APSIPA Transactions on Signal and Information Processing , vol. 11, no. 1, 2022
2022
Later among the works it cites.
X. Chen, Z. Meng, S. Parthasarathy, and J. Li, “Factorized neural transducer for efficient language model adaptation,” in Proc. ICASSP , 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, “Transformer transducer: A streamable speech recognition model with Transformer encoders and RNN-T loss,” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
E. Variani, D. Rybach, C. Allauzen, and M. Riley, “Hybrid autoregressive transducer (HAT),” in Proc. ICASSP , 2020
2020
Cited alongside, same era.
W. Huang, W. Hu, Y. T. Yeung, and X. Chen, “Conv-Transformer transducer: Low latency, low frame rate, streamable end-to-end speech recognition,” in Proc. Interspeech , 2020
2020
Cited alongside, same era.
X. Chen, Y. Wu, Z. Wang, S. Liu, and J. Li, “Developing real-time streaming Transformer transducer for speech recognition on large-scale dataset,” in Proc. ICASSP , 2021
2021
Cited alongside, same era.
C. Yi, S. Zhou, and B. Xu, “Efficiently fusing pretrained acoustic and linguistic encoders for low-resource speech recognition,” IEEE Signal Process. Lett. , vol. 28, pp. 788–792, 2021
2021
Cited alongside, same era.
Y. Higuchi, N. Chen, Y. Fujita, H. Inaguma, T. Komatsu, J. Lee, J. Nozaki, T. Wang, and S. Watanabe, “A comparative study on non-autoregressive modelings for speech-to-text generation,” in Proc. ASRU , 2021
2021
Cited alongside, same era.
Z. Meng, N. Kanda, Y. Gaur, S. Parthasarathy, E. Sun, L. Lu, X. Chen, J. Li, and Y. Gong, “Internal language model training for domain-adaptive end-to-end speech recognition,” in Proc. ICASSP , 2021
2021
Cited alongside, same era.
Y. Higuchi, B. Yan, S. Arora, T. Ogawa, T. Kobayashi, and S. Watanabe, “BERT meets CTC: New formulation of end-to-end speech recognition with pre-trained masked language model,” in Proc. EMNLP (Findings) , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
K. Deng, S. Cao, Y. Zhang, L. Ma, G. Cheng, J. Xu, and P. Zhang, “Improving CTC-based speech recognition via knowledge transferring from pre-trained language models,” in Proc. ICASSP , 2022
2022
Later among the works it cites.
C. Choudhury, A. Gandhe, X. Ding, and I. Bulyko, “A likelihood ratio based domain adaptation method for E2E models,” in Proc. ICASSP , 2022
2022
Later among the works it cites.
W. Zhou, Z. Zheng, R. Schlüter, and H. Ney, “On language model integration for RNN transducer based speech recognition,” in Proc. ICASSP , 2022
2022
Later among the works it cites.
E. Tsunoo, Y. Kashiwagi, C. P. Narisetty, and S. Watanabe, “Residual language model for end-to-end speech recognition,” in Proc. Interspeech , 2022
2022
Later among the works it cites.
Z. Meng, Y. Gaur, N. Kanda, J. Li, X. Chen, Y. Wu, and Y. Gong, “Internal language model adaptation with text-only data for end-to-end speech recognition,” in Proc. Interspeech , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
X. Yang, Q. Li, and P. C. Woodland, “Knowledge distillation for neural transducers from large self-supervised pre-trained models,” in Proc. ICASSP , 2022
2022
Later among the works it cites.
D. Albesano, J. Andrés-Ferrer, N. Ferri, and P. Zhan, “On the prediction network architecture in RNN-T for ASR,” in Proc. Interspeech , 2022
2022
Later among the works it cites.
Q. Li, C. Zhang, and P. C. Woodland, “Combining hybrid DNN-HMM ASR systems with attention-based models using lattice rescoring,” Speech Communication , vol. 147, pp. 12–21, 2023
2023
Closest in time.
2023
Closest in time.
Z. Meng, T. Chen, R. Prabhavalkar, Y. Zhang, G. Wang, K. Audhkhasi, J. Emond, T. Strohman, B. Ramabhadran, W. R. Huang, E. Variani, Y. Huang, and P. J. Moreno, “Modular hybrid autoregressive transducer,” in Proc. SLT , 2023
2023
Closest in time.
Y. Sudo, S. Muhammad, Y. Peng, and S. Watanabe, “Time-synchronous one-pass beam search for parallel online and offline transducers with dynamic block training,” in Proc. Interspeech , 2023
2023
Closest in time.