Fetching the paper…
Reading the bibliography…
Text-only adaptation of an end-to-end (E2E) model remains a challenging task for automatic speech recognition (ASR).
S. Kullback and R. A. Leibler, “On information and sufficiency,” The annals of mathematical statistics , vol. 22, no. 1, pp. 79–86, 1951
1951
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez et al. , “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in ICML . ACM, 2006
2006
Earlier work this paper cites.
R. Gemello, F. Mana, S. Scanzio et al. , “Linear hidden transformations for adaptation of hybrid ANN/HMM models,” Speech Communication , vol. 49, no. 10, pp. 827 – 835, 2007
2007
Earlier work this paper cites.
T. Mikolov, M. Karafiát et al. , “Recurrent neural network based language model.” in Interspeech , 2010
2010
Earlier work this paper cites.
2012
Earlier work this paper cites.
D. Yu, K. Yao, H. Su et al. , “Kl-divergence regularized deep neural network adaptation for improved large vocabulary speech recognition,” in Proc. ICASSP , May 2013
2013
Earlier work this paper cites.
H. Liao, “Speaker adaptation of context dependent deep neural networks,” in Proc. ICASSP , May 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
P. Swietojanski and S. Renals, “Learning hidden unit contributions for unsupervised speaker adaptation of neural network acoustic models,” in Proc. SLT . IEEE, 2014, pp. 171–176
2014
Earlier work this paper cites.
J. Li, R. Zhao et al. , “Learning small-size DNN with output-distribution-based criteria.” in INTERSPEECH , 2014
2014
Earlier work this paper cites.
H. Sak, A. Senior, and F. Beaufays, “Long short-term memory recurrent neural network architectures for large scale acoustic modeling,” in Interspeech , 2014
2014
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk et al. , “Attention-based models for speech recognition,” in NIPS , 2015
2015
Earlier work this paper cites.
T. Tan, Y. Qian, M. Yin et al. , “Cluster adaptive training for deep neural network,” in Proc. ICASSP . IEEE, 2015
2015
Earlier work this paper cites.
C. Gulcehre, O. Firat, K. Xu et al. , “On using monolingual corpora in neural machine translation,” CoRR , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Y. Shinohara, “Adversarial multi-task learning of deep neural networks for robust speech recognition,” in INTERSPEECH , 2016
2016
Earlier work this paper cites.
D. Serdyuk, K. Audhkhasi et al. , “Invariant representations for noisy speech recognition,” in NIPS Workshop , 2016
2016
Earlier work this paper cites.
N. Kanda, X. Lu, and H. Kawai, “Maximum a posteriori based decoding for CTC acoustic models.” in Interspeech , 2016
2016
Earlier work this paper cites.
Z. Meng, Z. Chen, V. Mazalov, J. Li, and Y. Gong, “Unsupervised adaptation with domain separation networks for robust speech recognition,” in Proc. ASRU , 2017
2017
Earlier work this paper cites.
N. Kanda, X. Lu, and H. Kawai, “Minimum bayes risk training of CTC acoustic models in maximum a posteriori based decoding framework,” in Proc. ICASSP . IEEE, 2017
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar et al. , “Attention is all you need,” in Proc. NIPS , vol. 30, 2017
2017
Cited alongside, same era.
Z. Meng, S. Watanabe, J. R. Hershey et al. , “Deep long short-term memory adaptive beamforming networks for multichannel robust speech recognition,” in ICASSP . IEEE, 2017
2017
Cited alongside, same era.
C.-C. Chiu, T. N. Sainath et al. , “State-of-the-art speech recognition with sequence-to-sequence models,” in Proc. ICASSP , 2018
2018
Cited alongside, same era.
J. Li, G. Ye, A. Das et al. , “Advancing acoustic-to-word CTC model,” in Proc. ICASSP , 2018
2018
Cited alongside, same era.
2019
Later among the works it cites.
T. Sainath, Y. He, B. Li et al. , “A streaming on-device end-to-end model surpassing server-side conventional model quality and latency,” in Proc. ICASSP , 2020, pp. 6059–6063
2020
Later among the works it cites.
J. Li, R. Zhao, Z. Meng et al. , “Developing RNN-T models surpassing high-performance hybrid models with customization capability,” in Interspeech , 2020
2020
Later among the works it cites.
J. Li, Y. Wu et al. , “On the comparison of popular end-to-end models for large scale speech recognition,” in Proc. Interspeech , 2020
2020
Later among the works it cites.
Z. Meng, H. Hu, J. Li et al. , “L-vector: Neural label embedding for domain adaptation,” in Proc. ICASSP . IEEE, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Meng, J. Li, Y. Gong et al. , “Adversarial teacher-student learning for unsupervised domain adaptation,” in Proc. ICASSP , 2018
2018
Cited alongside, same era.
V. Manohar, P. Ghahremani, D. Povey et al. , “A teacher-student learning approach for unsupervised domain adaptation of sequence-trained ASR models,” in Proc. SLT . IEEE, 2018
2018
Cited alongside, same era.
Z. Meng, J. Li, Z. Chen et al. , “Speaker-invariant training via adversarial learning,” in Proc. ICASSP , 2018
2018
Cited alongside, same era.
T. Ochiai, S. Watanabe et al. , “Speaker adaptation for multichannel end-to-end speech recognition,” in Proc. ICASSP , 2018
2018
Cited alongside, same era.
A. Sriram, H. Jun, S. Satheesh et al. , “Cold fusion: Training seq2seq models together with language models,” Proc. Interspeech , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
S. Karita, N. Chen et al. , “A comparative study on transformer vs RNN in speech applications,” in Proc. ASRU , 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
C. Peyser, S. Mavandadi, T. Sainath et al. , “Improving tail performance of a deliberation E2E ASR model using a large text corpus,” in INTERSPEECH , 2020
2020
Later among the works it cites.
Y. Huang, J. Li, L. He et al. , “Rapid RNN-T adaptation using personalized speech synthesis and neural language generator.” in INTERSPEECH , 2020, pp. 1256–1260
2020
Later among the works it cites.
Q. Zhang, H. Lu, H. Sak et al. , “Transformer transducer: A streamable speech recognition model with transformer encoders and RNN-T loss,” in Proc. ICASSP . IEEE, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2021
Closest in time.
Z. Meng, S. Parthasarathy, E. Sun et al. , “Internal language model estimation for domain-adaptive end-to-end speech recognition,” in Proc. SLT . IEEE, 2021
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
Z. Meng, Y. Wu, N. Kanda et al. , “Minimum word error rate training with language model fusion for end-to-end speech recognition,” Proc. Interspeech , 2021
2021
Closest in time.
2021
Closest in time.
Z. Meng, N. Kanda, Y. Gaur et al. , “Internal language model training for domain-adaptive end-to-end speech recognition,” in Proc. ICASSP . IEEE, 2021
2021
Closest in time.
X. Chen, Z. Meng et al. , “Factorized neural transducer for efficient language model adaptation,” in Proc. ICASSP . IEEE, 2022
2022
Closest in time.