Fetching the paper…
Reading the bibliography…
The external language models (LM) integration remains a challenging task for end-to-end (E2E) automatic speech recognition (ASR) which has no clear division between acoustic and language models.
“Learning small-size DNN with output-distribution-based criteria.,”
J. Li, R. Zhao, J.-T. Huang, and Y. Gong, · 1914
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, et al., · 2006
Earlier work this paper cites.
“Linear hidden transformations for adaptation of hybrid ann/hmm models,”
R. Gemello, F. Mana, S. Scanzio, et al., · 2007
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
G. Hinton, L. Deng, D. Yu, et al., · 2012
Earlier work this paper cites.
“A fast and simple algorithm for training neural probabilistic language models,”
A. Mnih and Y. W. Teh, · 2012
Earlier work this paper cites.
“Kl-divergence regularized deep neural network adaptation for improved large vocabulary speech recognition,”
D. Yu, K. Yao, H. Su, G. Li, and F. Seide, · 2013
Earlier work this paper cites.
“Speaker adaptation of context dependent deep neural networks,”
H. Liao, · 2013
Earlier work this paper cites.
“Fast speaker adaptation of hybrid nn/hmm model for speech recognition based on discriminative learning of speaker code,”
O. Abdel-Hamid and H. Jiang, · 2013
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
A. Graves and N. Jaitly, · 2014
Earlier work this paper cites.
“Deep speech: Scaling up end-to-end speech recognition,”
A. Hannun, C. Case, J. Casper, et al., · 2014
Earlier work this paper cites.
“Learning hidden unit contributions for unsupervised speaker adaptation of neural network acoustic models,”
P. Swietojanski and S. Renals, · 2014
Earlier work this paper cites.
“Long short-term memory recurrent neural network architectures for large scale acoustic modeling,”
F. Beaufays H. Sak, A. Senior, · 2014
Earlier work this paper cites.
“Dropout: a simple way to prevent neural networks from overfitting,”
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, et al., · 2014
Earlier work this paper cites.
“Attention-based models for speech recognition,”
J. K Chorowski, D. Bahdanau, D. Serdyuk, et al., · 2015
Earlier work this paper cites.
“Cluster adaptive training for deep neural network,”
T. Tan, Y. Qian, M. Yin, et al., · 2015
Earlier work this paper cites.
“On using monolingual corpora in neural machine translation,”
C. Gulcehre, O. Firat, K. Xu, et al., · 2015
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
R. Sennrich, B. Haddow, and A. Birch, · 2015
Earlier work this paper cites.
“Scheduled sampling for sequence prediction with recurrent neural networks,”
S. Bengio, O. Vinyals, N. Jaitly, et al., · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, et al., · 2015
Earlier work this paper cites.
“Neural speech recognizer: Acoustic-to-word lstm model for large vocabulary speech recognition,”
H. Soltau, H. Liao, and H. Sak, · 2016
Cited alongside, same era.
“End-to-end attention-based large vocabulary speech recognition,”
D. Bahdanau, J. Chorowski, D. Serdyuk, et al., · 2016
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, · 2016
Cited alongside, same era.
“Factorized hidden layer adaptation for deep neural network based acoustic modeling,”
L. Samarakoon and K. C. Sim, · 2016
Cited alongside, same era.
“Adversarial multi-task learning of deep neural networks for robust speech recognition.,”
Y. Shinohara, · 2016
Cited alongside, same era.
“Invariant representations for noisy speech recognition,”
“Speaker adaptation for multichannel end-to-end speech recognition,”
T. Ochiai, S. Watanabe, S. Katagiri, et al., · 2018
Later among the works it cites.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
A. Kannan, Y. Wu, P. Nguyen, et al., · 2018
Later among the works it cites.
“Cold fusion: Training seq2seq models together with language models,”
A. Sriram, H. Jun, S. Satheesh, et al., · 2018
Later among the works it cites.
“A comparison of techniques for language model integration in encoder-decoder speech recognition,”
S. Toshniwal, A. Kannan, C. Chiu, et al., · 2018
Later among the works it cites.
“Simple fusion: Return of the language model,”
F. Stahlberg, J. Cross, and V. Stoyanov, · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Serdyuk, K. Audhkhasi, P. Brakel, B. Ramabhadran, et al., · 2016
Cited alongside, same era.
“Towards better decoding and language model integration in sequence to sequence models,”
J. Chorowski and N. Jaitly, · 2016
Cited alongside, same era.
“Maximum a posteriori based decoding for ctc acoustic models.,”
N. Kanda, X. Lu, and H. Kawai, · 2016
Cited alongside, same era.
“Multi-channel speech recognition: Lstms all the way through,”
H. Erdogan, T. Hayashi, J. R. Hershey, et al., · 2016
Cited alongside, same era.
J. L. Ba, J. R. Kiros, and G. E Hinton, · 2016
Cited alongside, same era.
“Unsupervised adaptation with domain separation networks for robust speech recognition,”
Z. Meng, Z. Chen, V. Mazalov, J. Li, and Y. Gong, · 2017
Cited alongside, same era.
“Multi-level language modeling and decoding for open vocabulary end-to-end speech recognition,”
T. Hori, S. Watanabe, and J. Hershey, · 2017
Cited alongside, same era.
M. Jain, K. Schubert, J. Mahadeokar, et al., · 2019
Later among the works it cites.
“A comparative study on transformer vs RNN in speech applications,”
S. Karita, N. Chen, T. Hayashi, et al., · 2019
Later among the works it cites.
“Adversarial speaker adaptation,”
Z. Meng, J. Li, and Y. Gong, · 2019
Later among the works it cites.
“Conditional teacher-student learning,”
Z. Meng, J. Li, Y. Zhao, and Y. Gong, · 2019
Later among the works it cites.
“Speaker adaptation for attention-based end-to-end speech recognition,”
Z. Meng, Y. Gaur, J. Li, and Y. Gong, · 2019
Later among the works it cites.
“Domain adaptation via teacher-student learning for end-to-end speech recognition,”
Z. Meng, J. Li, Y. Gaur, and Y. Gong, · 2019
Later among the works it cites.
“Component fusion: Learning replaceable language model component for end-to-end speech recognition system,”
C. Shan, C. Weng, G. Wang, et al., · 2019
Later among the works it cites.
“A density ratio approach to language model fusion in end-to-end automatic speech recognition,”
E. McDermott, H. Sak, and E. Variani, · 2019
Later among the works it cites.
“Character-aware attention-based end-to-end speech recognition,”
Z. Meng, Y. Gaur, J. Li, and Y. Gong, · 2019
Later among the works it cites.
“Acoustic-to-phrase end-to-end speech recognition,”
Y. Gaur, J. Li, Z. Meng, and Y. Gong, · 2019
Later among the works it cites.
“A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,”
T. Sainath, Y. He, B. Li, et al., · 2020
Closest in time.
“Developing RNN-T models surpassing high-performance hybrid models with customization capability,”
J. Li, R. Zhao, Z. Meng, et al., · 2020
Closest in time.
“On the comparison of popular end-to-end models for large scale speech recognition,”
J. Li, Y. Wu, Y. Gaur, et al., · 2020
Closest in time.
“L-vector: Neural label embedding for domain adaptation,”
Z. Meng, H. Hu, J. Li, et al., · 2020
Closest in time.
“Hybrid autoregressive transducer (hat),”
E. Variani, D. Rybach, C. Allauzen, et al., · 2020
Closest in time.