Fetching the paper…
Reading the bibliography…
We present our transducer model on Librispeech.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
2005
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” Preprint arXiv:1211.3711, 2012
2012
Earlier work this paper cites.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in ICASSP , 2013
2013
Earlier work this paper cites.
L. Wan, M. Zeiler, S. Zhang, Y. L. Cun, and R. Fergus, “Regularization of neural networks using dropconnect,” in Proceedings of the 30th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, vol. 28, no. 3. Atlanta, Georgia, USA: PMLR, 17–19 Jun 2013, pp. 1058–1066. [Online]. Available: http://proceedings.mlr.press/v28/wan13.html
2013
Earlier work this paper cites.
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Proceedings of the International Conference on Learning Representations (ICLR) , San Diego, CA, 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: an ASR corpus based on public domain audio books,” in ICASSP . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
TensorFlow Development Team, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org. [Online]. Available: https://www.tensorflow.org/
2015
Earlier work this paper cites.
D. Krueger, T. Maharaj, J. Kramár, M. Pezeshki, N. Ballas, N. R. Ke, A. Goyal, Y. Bengio, A. Courville, and C. Pal, “Zoneout: Regularizing RNNs by randomly preserving hidden activations,” in ICLR , 2017
2017
Earlier work this paper cites.
Ç. Gülçehre, O. Firat, K. Xu, K. Cho, L. Barrault, H.-C. Lin, F. Bougares, H. Schwenk, and Y. Bengio, “On using monolingual corpora in neural machine translation,” Computer Speech & Language , vol. 45, pp. 137–148, Sep. 2017
2017
Cited alongside, same era.
A. Zeyer, K. Irie, R. Schlüter, and H. Ney, “Improved training of end-to-end attention models for speech recognition,” in Interspeech , Hyderabad, India, Sep. 2018
2018
Cited alongside, same era.
A. Zeyer, A. Merboldt, R. Schlüter, and H. Ney, “A comprehensive analysis on attention models,” in Interpretability and Robustness in Audio, Speech, and Language (IRASL) Workshop, NeurIPS , Montreal, Canada, Dec. 2018
2018
Cited alongside, same era.
A. Zeyer, T. Alkhouli, and H. Ney, “RETURNN as a generic flexible neural toolkit with application to translation and speech recognition,” in Annual Meeting of the Assoc. for Computational Linguistics , Melbourne, Australia, Jul. 2018
2018
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
E. Variani, D. Rybach, C. Allauzen, and M. Riley, “Hybrid autoregressive transducer (HAT),” in ICASSP , 2020
2020
Later among the works it cites.
A. Zeyer, A. Merboldt, R. Schlüter, and H. Ney, “A new training pipeline for an improved neural transducer,” in Interspeech , Shanghai, China, Oct. 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
E. McDermott, H. Sak, and E. Variani, “A density ratio approach to language model fusion in end-to-end automatic speech recognition,” in ASRU , 2019, pp. 434–441
2019
Cited alongside, same era.
C. Lüscher, E. Beck, K. Irie, M. Kitza, W. Michel, A. Zeyer, R. Schlüter, and H. Ney, “RWTH ASR systems for LibriSpeech: Hybrid vs attention,” in Interspeech , Graz, Austria, Sep. 2019, pp. 231–235
2019
Cited alongside, same era.
A. Zeyer, P. Bahar, K. Irie, R. Schlüter, and H. Ney, “A comparison of Transformer and LSTM encoder decoder models for ASR,” in ASRU , Sentosa, Singapore, Dec. 2019, pp. 8–15
2019
Cited alongside, same era.
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar, “Transformer transducer: A streamable speech recognition model with Transformer encoders and RNN-T loss,” in ICASSP . IEEE, 2020, pp. 7829–7833
2020
Cited alongside, same era.
2020
Later among the works it cites.
K. Irie, “Advancing neural language modeling in automatic speech recognition,” Ph.D. dissertation, RWTH Aachen University, Computer Science Department, RWTH Aachen University, Aachen, Germany, May 2020. [Online]. Available: http://publications.rwth-aachen.de/record/789081
2020
Later among the works it cites.
W. Zhou, S. Berger, R. Schlüter, and H. Ney, “Phoneme based neural transducer for large vocabulary speech recognition,” in IEEE International Conference on Acoustics, Speech, and Signal Processing , May 2021, submitted to
2021
Closest in time.
Z. Meng, S. Parthasarathy, E. Sun, Y. Gaur, N. Kanda, L. Lu, X. Chen, R. Zhao, J. Li, and Y. Gong, “Internal language model estimation for domain-adaptive end-to-end speech recognition,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 243–250
2021
Closest in time.
M. Zeineldeen, A. Glushko, W. Michel, A. Zeyer, R. Schlüter, and H. Ney, “Investigating methods to improve language model integration for attention-based encoder-decoder ASR models,” submitted to Interspeech 2021, 2021
2021
Closest in time.