Fetching the paper…
Reading the bibliography…
Conventional automatic speech recognition (ASR) systems trained from frame-level alignments can easily leverage posterior fusion to improve ASR accuracy and build a better single model with knowledge distillation.
J. Li, R. Zhao, J.-T. Huang, and Y. Gong, “Learning small-size DNN with output-distribution-based criteria,” in
1914
Earlier work this paper cites.
L. R. Rabiner, “A tutorial on hidden Markov models and selected applications in speech recognition,”
1989
Earlier work this paper cites.
J. G. Fiscus, “A post-processing system to yield reduced word error rates: Recognizer output voting error reduction (ROVER),” in
1997
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in
2006
Earlier work this paper cites.
N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, and P. Ouellet, “Front-end factor analysis for speaker verification,”
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
2012
Earlier work this paper cites.
A. Graves, A. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in
2013
Earlier work this paper cites.
A. Graves, “Generating sequences with recurrent neural networks,”
2013
Earlier work this paper cites.
T. N. Sainath, B. Kingsbury, V. Sindhwani, E. Arisoy, and B. Ramabhadran, “Low-rank matrix factorization for deep neural network training with high-dimensional output targets,” in
2013
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Ba and R. Caruana, “Do deep nets really need to be deep?” in
2014
Earlier work this paper cites.
Y. Miao, M. Gowayyed, and F. Metze, “EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,” in
2015
Earlier work this paper cites.
H. Sak, A. Senior, K. Rao, O. Irsoy, A. Graves, F. Beaufays, and J. Schalkwyk, “Learning acoustic frame labeling for speech recognition with recurrent neural networks,” in
2015
Earlier work this paper cites.
J. K. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” in
2015
Cited alongside, same era.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,”
2015
Cited alongside, same era.
H. Sak, F. de Chaumont Quitry, T. Sainath, K. Rao
2015
Cited alongside, same era.
2015
Cited alongside, same era.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in
2016
Cited alongside, same era.
K. Audhkhasi, B. Ramabhadran, G. Saon, M. Picheny, and D. Nahamoo, “Direct acoustics-to-word models for English conversational speech recognition,” in
2017
Later among the works it cites.
H. Liu, Z. Zhu, X. Li, and S. Satheesh, “Gram-CTC: Automatic unit selection and target decomposition for sequence labelling,” in
2017
Later among the works it cites.
G. Saon, G. Kurata, T. Sercu, K. Audhkhasi, S. Thomas, D. Dimitriadis, X. Cui, B. Ramabhadran, M. Picheny, L.-L. Lim, B. Roomi, and P. Hall, “English conversational telephone speech recognition by humans and machines,” in
2017
Later among the works it cites.
T. Fukuda, M. Suzuki, G. Kurata, S. Thomas, J. Cui, and B. Ramabhadran, “Efficient knowledge distillation from an ensemble of teachers,” in
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Pundak and T. Sainath, “Lower frame rate neural network acoustic models,” in
2016
Cited alongside, same era.
Y. Miao, M. Gowayyed, X. Na, T. Ko, F. Metze, and A. Waibel, “An empirical exploration of CTC acoustic models,” in
2016
Cited alongside, same era.
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in
2016
Cited alongside, same era.
Y. Chebotar and A. Waters, “Distilling knowledge from ensembles of neural networks for speech recognition.” in
2016
Cited alongside, same era.
D. Amodei, S. Ananthanarayanan, R. Anubhai, J. Bai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, Q. Cheng, G. Chen
2016
Cited alongside, same era.
T. Hori, S. Watanabe, and J. Hershey, “Joint CTC/attention decoding for end-to-end speech recognition,” in
2017
Cited alongside, same era.
G. Zweig, C. Yu, J. Droppo, and A. Stolcke, “Advances in all-neural speech recognition,” in
2017
Cited alongside, same era.
2017
Later among the works it cites.
L. Fritz and D. Burshtein, “Simplified end-to-end MMI training and voting for ASR,”
2017
Later among the works it cites.
H. Soltau, H. Liao, and H. Sak, “Reducing the computational complexity for whole word models,” in
2017
Later among the works it cites.
K. Audhkhasi, B. Kingsbury, B. Ramabhadran, G. Saon, and M. Picheny, “Building competitive direct acoustics-to-word models for English conversational speech recognition,” in
2018
Later among the works it cites.
G. Kurata and K. Audhkhasi, “Improved knowledge distillation from bi-directional to uni-directional LSTM CTC for end-to-end speech recognition,” in
2018
Later among the works it cites.
S. Ghorbani, A. E. Bulut, and J. H. Hansen, “Advancing multi-accented LSTM-CTC speech recognition using a domain specific student-teacher learning paradigm,” in
2018
Later among the works it cites.
S. Kim, M. L. Seltzer, J. Li, and R. Zhao, “Improved training for online end-to-end speech recognition systems,” in
2018
Later among the works it cites.
R. Takashima, S. Li, and H. Kaswai, “An investigation of a knowledge distillation method for CTC acoustic models,” in
2018
Later among the works it cites.
R. Sanabria and F. Metze, “Hierarchical multi task learning with CTC,” in
2018
Later among the works it cites.
C. Yu, C. Zhang, C. Weng, J. Cui, and D. Yu, “A multistage training framework for acoustic-to-word model,” in
2018
Later among the works it cites.