Fetching the paper…
Reading the bibliography…
Attention-based encoder-decoder (AED) models learn an implicit internal language model (ILM) from the training transcriptions.
A. Krogh and J. Hertz, “A simple weight decay can improve generalization,” in NIPS , 1991
1991
Earlier work this paper cites.
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “Switchboard: Telephone speech corpus for research and development,” in ICASSP , 1992
1992
Earlier work this paper cites.
R. Schlüter, I. Bezrukov, H. Wagner, and H. Ney, “Gammatone features and feature combination for large vocabulary speech recognition,” in IEEE International Conference on Acoustics, Speech, and Signal Processing , Honolulu, HI, USA, Apr. 2007, pp. 649–652
2007
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” ArXiv , vol. abs/1211.3711, 2012
2012
Earlier work this paper cites.
A. Rousseau, P. Deléglise, and Y. Estève, “Enhancing the TED-LIUM corpus with selected data for language modeling and more TED talks,” in Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14) . Reykjavik, Iceland: European Language Resources Association (ELRA), May 2014, pp. 3935–3939. [Online]. Available: http://www.lrec-conf.org/proceedings/lrec2014/pdf/1104_Paper.pdf
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in ICASSP , 2015
2015
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in ICASSP , 2016. [Online]. Available: http://williamchan.ca/papers/wchan-icassp-2016.pdf
2016
Earlier work this paper cites.
2017
Cited alongside, same era.
S. Toshniwal, A. Kannan, C. Chiu, Y. Wu, T. N. Sainath, and K. Livescu, “A comparison of techniques for language model integration in encoder-decoder speech recognition,” in 2018 IEEE Spoken Language Technology Workshop (SLT) , 2018, pp. 369–375
2018
Cited alongside, same era.
A. Sriram, H. Jun, S. Satheesh, and A. Coates, “Cold fusion: Training seq2seq models together with language models,” in INTERSPEECH , 2018
2018
Cited alongside, same era.
F. Stahlberg, J. Cross, and V. Stoyanov, “Simple fusion: Return of the language model,” in Proceedings of the Third Conference on Machine Translation: Research Papers . Brussels, Belgium: Association for Computational Linguistics, Oct. 2018, pp. 204–211. [Online]. Available: https://www.aclweb.org/anthology/W18-6321
E. McDermott, H. Sak, and E. Variani, “A density ratio approach to language model fusion in end-to-end automatic speech recognition,” 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , pp. 434–441, 2019
2019
Later among the works it cites.
A. Zeyer, P. Bahar, K. Irie, R. Schlüter, and H. Ney, “A comparison of transformer and lstm encoder decoder models for asr,” in IEEE Automatic Speech Recognition and Understanding Workshop , Sentosa, Singapore, Dec. 2019, pp. 8–15
2019
Later among the works it cites.
K. Irie, A. Zeyer, R. Schlüter, and H. Ney, “Language modeling with deep transformers,” in INTERSPEECH , 2019
2019
Later among the works it cites.
J. Li, R. Zhao, Z. Meng, Y. Liu, W. Wei, S. Parthasarathy, V. Mazalov, Z. Wang, L. He, S. Zhao, and Y. Gong, “Developing rnn-t models surpassing high-performance hybrid models with customization capability,” in INTERSPEECH , 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
A. Zeyer, K. Irie, R. Schlüter, and H. Ney, “Improved training of end-to-end attention models for speech recognition,” in Interspeech , Hyderabad, India, Sep. 2018
2018
Cited alongside, same era.
A. Zeyer, T. Alkhouli, and H. Ney, “RETURNN as a generic flexible neural toolkit with application to translation and speech recognition,” in ACL , 2018
2018
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” in Interspeech , 2019
2019
Cited alongside, same era.
C. Shan, C. Weng, G. Wang, D. Su, M. Luo, D. Yu, and L. Xie, “Component fusion: Learning replaceable language model component for end-to-end speech recognition system,” in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2019, pp. 5361–5635
2019
Cited alongside, same era.
W. Michel, R. Schlüter, and H. Ney, “Early Stage LM Integration Using Local and Global Log-Linear Combination,” in Proc. Interspeech 2020 , 2020, pp. 3605–3609. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-2675
2020
Later among the works it cites.
E. Variani, D. Rybach, C. Allauzen, and M. Riley, “Hybrid autoregressive transducer (HAT),” in ICASSP , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
Z. Tüske, G. Saon, K. Audhkhasi, and B. Kingsbury, “Single headed attention based sequence-to-sequence model for state-of-the-art results on switchboard-300,” in INTERSPEECH , 2020
2020
Later among the works it cites.