Fetching the paper…
Reading the bibliography…
Knowledge Distillation is an effective method of transferring knowledge from a large model to a smaller model.
“The aurora experimental framework for the performance evaluation of speech recognition systems under noisy conditions,”
D. Pearce and H. Hirsch, · 2000
Earlier work this paper cites.
“Model compression,”
C. Buciluǎ, R. Caruana, and A. Niculescu-Mizil, · 2006
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence Transduction with Recurrent Neural Networks,”
A. Graves, · 2012
Earlier work this paper cites.
“Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,”
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, et al., · 2012
Earlier work this paper cites.
“KL-divergence regularized deep neural network adaptation for improved large vocabulary speech recognition,”
D. Yu, K. Yao, H. Su, G. Li, and F. Seide, · 2013
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
G. E. Hinton, O. Vinyals, and J. Dean, · 2015
Earlier work this paper cites.
“Acoustic modelling with cd-ctc-smbr lstm rnns,”
H. Sak, F. de Chaumont Quitry, T. Sainath, K. Rao, et al., · 2015
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Distilling knowledge from ensembles of neural networks for speech recognition.,”
Y. Chebotar and A. Waters, · 2016
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, · 2016
Earlier work this paper cites.
“Student-teacher network learning with enhanced features,”
S. Watanabe, T. Hori, J. Le Roux, and J. R. Hershey, · 2017
Earlier work this paper cites.
“Knowledge distillation for small-footprint highway networks,”
L. Lu, M. Guo, and S. Renals, · 2017
Cited alongside, same era.
“Efficient knowledge distillation from an ensemble of teachers.,”
T. Fukuda, M. Suzuki, G. Kurata, S. Thomas, J. Cui, and B. Ramabhadran, · 2017
Cited alongside, same era.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, · 2017
Cited alongside, same era.
“Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in google home,”
C. Kim, A. Misra, K. Chin, T. Hughes, A. Narayanan, T. Sainath, and M. Bacchiani, · 2017
Cited alongside, same era.
“Block-sparse recurrent neural networks,”
S. Narang, E. Undersander, and G. Diamos, · 2017
Cited alongside, same era.
“Compression of End-to-End Models,”
R. Pang, T. Sainath, R. Prabhavalkar, S. Gupta, Y. Wu, S. Zhang, and C. Chiu, · 2018
“Streaming end-to-end speech recognition for mobile devices,”
Y. He, T. N. Sainath, R. Prabhavalkar, I. McGraw, R. Alvarez, D. Zhao, D. Rybach, A. Kannan, Y. Wu, R. Pang, et al., · 2019
Later among the works it cites.
“Investigation of sequence-level knowledge distillation methods for ctc acoustic models,”
R. Takashima, L. Sheng, and H. Kawai, · 2019
Later among the works it cites.
“Guiding ctc posterior spike timings for improved posterior fusion and knowledge distillation,”
G. Kurata and K. Audhkhasi, · 2019
Later among the works it cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, · 2019
Later among the works it cites.
“Gradient Based Pruning,”
Y. Yang and R. Panigrahy, · 2019
Later among the works it cites.
“SNIP: Single-shot Network Pruning based on Connection Sensitivity,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Toward domain-invariant speech recognition via large scale training,”
A. Narayanan, A. Misra, K. C. Sim, G. Pundak, A. Tripathi, M. Elfeky, P. Haghani, T. Strohman, and M. Bacchiani, · 2018
Cited alongside, same era.
“An investigation of a knowledge distillation method for ctc acoustic models,”
R. Takashima, S. Li, and H. Kawai, · 2018
Cited alongside, same era.
“Improved knowledge distillation from bi-directional to uni-directional lstm ctc for end-to-end speech recognition,”
G. Kurata and K. Audhkhasi, · 2018
Cited alongside, same era.
“Group normalization,”
Y. Wu and K. He, · 2018
Cited alongside, same era.
“To prune, or not to prune: Exploring the efficacy of pruning for model compression,”
M. H. Zhu and S. Gupta, · 2018
Cited alongside, same era.
“Improving Noise Robustness of Automatic Speech Recognition via Parallel Data and Teacher-student Learning,”
L. Mošner, M. Wu, A. Raju, S. H. K. Parthasarathi, K. Kumatani, S. Sundaram, R. Maas, and B. Hoffmeister, · 2019
Cited alongside, same era.
N. Lee, T. Ajanthan, and P. H. S. Torr, · 2019
Later among the works it cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, et al., · 2020
Closest in time.
“Optimizing Speech Recognition for the Edge,”
Y. Shangguan, J. Li, Q. Liang, R. Alvarez, and I. McGraw, · 2020
Closest in time.
“Specaugment on large scale datasets,”
D. S. Park, Y. Zhang, C.-C. Chiu, Y. Chen, B. Li, W. Chan, Q. V. Le, and Y. Wu, · 2020
Closest in time.
“Universal ASR: Unify and Improve Streaming ASR with Full-Context Modeling,”
J. Yu, W. Han, A. Gulati, C.-C. Chiu, B. Li, T. N. Sainath, Y. Wu, and R. Pang, · 2020
Closest in time.
“ContextNet: Improving Convolutional Neural Networks for Automatic Speech Recognition with Global Context,”
W. Han, Z. Zhang, Y. Zhang, J. Yu, C.-C. Chiu, J. Qin, A. Gulati, R. Pang, and Y. Wu, · 2020
Closest in time.
“Dynamic Sparsity Neural Networks for Automatic Speech Recognition,”
Z. Wu, D. Zhao, Q. Liang, J. Yu, A. Gulati, and R. Pang, · 2021
Closest in time.