Fetching the paper…
Reading the bibliography…
Although large foundation models pre-trained by self-supervised learning have achieved state-of-the-art performance in many tasks including automatic speech recognition (ASR), knowledge distillation (KD) is often required in practice to transfer the knowledge learned by large teacher models into much smaller student models with affordable computation and memory costs.
T. G. Dietterich, “Ensemble methods in machine learning,” in proc. MCS , Cagliari, 2000
2000
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in Proc. ICML , Pittsburgh, 2006
2006
Earlier work this paper cites.
L. Rokach, “Ensemble-based classifiers,” Artificial Intelligence Review , vol. 33, p. 1–39, 2010
2010
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” in Proc. ICML Workshop on Representation Learning , Edinburgh, 2012
2012
Earlier work this paper cites.
A. M. A. Graves and G. Hinton, “Speech recognition with deep recurrent neural networks,” in Proc. ICASSP , Vancouver, 2013
2013
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” in Proc. NIPS Deep Learning Workshop , Montreal, 2014
2014
Earlier work this paper cites.
K. Hwang and W. Sung, “Fixed-point feedforward deep neural network design using weights+ 1, 0, and- 1,” in Proc. SiPS , Belfast, 2014
2014
Earlier work this paper cites.
K. C. L. Lu, X. Zhang and S. Renals, “A study of the recurrent neural network encoder-decoder for large vocabulary speech recognition,” in Proc. Interspeech , Dresden, 2015
2015
Earlier work this paper cites.
M. Courbariaux, Y. Bengio, and J.-P. David, “Binaryconnect: Training deep neural networks with binary weights during propagations,” in Proc. NIPS , Montreal, 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in Proc. ICASSP , Brisbane, 2015
2015
Earlier work this paper cites.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in Proc. ICASSP , Shanghai, 2016
2016
Earlier work this paper cites.
W. Chan, N. Jaitly, Q. V. Le, and O. Vinyals, “Listen, attend and spell,” in Proc. ICASSP , Shanghai, 2016
2016
Earlier work this paper cites.
J. H. M. Wong and M. J. F. Gales, “Sequence student-teacher training of deep neural networks,” in Proc. Interspeech , San Francisco, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals et al. , “WaveNet: A generative model for raw audio,” in Proc. SSW , Sunnyvale, 2016
2016
Earlier work this paper cites.
T. H. S. Kim and S. Watanabe, “Joint CTC-attention based end-to-end speech recognition using multi-task learning,” in Proc. ICASSP , New Orleans, 2017
2017
Earlier work this paper cites.
J. Li, M. L. Seltzer, X. Wang, R. Zhao, and Y. Gong, “Large-scale domain adaptation via teacher-student learning,” in Proc. Interspeech , Stockholm, 2017
2017
Earlier work this paper cites.
J. Wong and M. Gales, “Student-teacher training with diverse decision tree ensembles,” in Proc. Interspeech , Stockholm, 2017
2017
Earlier work this paper cites.
T. Fukuda, M. Suzuki, G. Kurata, S. Thomas, J. Cui, and B. Ramabhadran, “Efficient knowledge distillation from an ensemble of teachers,” in Proc. Interspeech , Shanghai, 2017
2017
Earlier work this paper cites.
S. Watanabe, T. Hori, S. Kim, J. R. Hershey, and T. Hayashi, “Hybrid CTC/attention architecture for end-to-end speech recognition,” IEEE Journal of Selected Topics in Signal Processing , vol. 11, pp. 1240–1253, 2017
2017
Cited alongside, same era.
E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with gumbel-softmax,” in Proc. ICLR , Toulon, 2017
2017
Cited alongside, same era.
B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and scalable predictive uncertainty estimation using deep ensembles,” in Proc. NIPS , Long Beach, 2017
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. NIPS , Long Beach, 2017
2017
Cited alongside, same era.
R. Takashima, S. Li, and H. Kawai, “An investigation of a knowledge distillation method for CTC acoustic models,” in Proc. ICASSP , Calgary, 2018
J. Kahn, M. Rivière, W. Zheng, Kharitonov et al. , “Libri-light: A benchmark for ASR with limited or no supervision,” in Proc. ICASSP , Barcelona, 2020
2020
Later among the works it cites.
Q. Xu, T. Likhomanenko, J. Kahn, A. Hannun, G. Synnaeve, and R. Collobert, “Iterative pseudo-labeling for speech recognition,” in Proc. Interspeech , Shanghai, 2020
2020
Later among the works it cites.
R. V. Swaminathan, B. King, G. P. Strimel, J. Droppo, and A. Mouchtaris, “CoDERT: Distilling encoder representations with co-learning for transducer-based speech recognition,” in Proc. Interspeech , Brno, 2021
2021
Later among the works it cites.
Z.-H. Zhou, “Ensemble learning,” in Machine learning . Springer, 2021, pp. 181–210
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
R. Sahraeian and D. Van Compernolle, “Cross-entropy training of DNN ensemble acoustic models for low-resource ASR,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, pp. 1991–2001, 2018
2018
Cited alongside, same era.
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba et al. , “ESPnet: End-to-end speech processing toolkit,” in Proc. Interspeech , Hyderabad, 2018
2018
Cited alongside, same era.
T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in Proc. EMNLP , Brussels, 2018
2018
Cited alongside, same era.
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph et al. , “SpecAugment: A simple data augmentation method for automatic speech recognition,” in Proc. Interspeech , Graz, 2019
2019
Cited alongside, same era.
J. Lindqvist, A. Olmin, Lindsten F., and Svensson L., “A general framework for ensemble distribution distillation,” in Proc. MLSP , Espoo, 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in NeurIPS , Vancouver, 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
A. Jaiswal, A. R. Babu, M. Z. Zadeh, D. Banerjee, and F. Makedon, “A survey on contrastive self-supervised learning,” Technologies , vol. 9, no. 1, 2021
2021
Later among the works it cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai et al. , “HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, Language Processing , pp. 3451–3460, 2021
2021
Later among the works it cites.
C. Wang, Y. Wu, Y. Qian, K. Kumatani, S. Liu, F. Wei, M. Zeng, and X. Huang, “Unispeech: Unified speech representation learning with labeled and unlabeled data,” in Proc. ICML , Baltimore, 2021
2021
Later among the works it cites.
L. Pepino, P. Riera, and L. Ferrer, “Emotion recognition from speech using wav2vec 2.0 embeddings,” in Proc. Interspeech , Brno, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Panchapagesan, D. S. Park, C.-C. Chiu, Y. Shangguan, Q. Liang, and A. Gruenstein, “Efficient knowledge distillation for RNN-transducer models,” in Proc. ICASSP , Toronto, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
X. Zheng, C. Zhang, and P. C. Woodland, “Tandem multitask training of speaker diarisation and speech recognition for meeting transcription,” in Proc. Interspeech , Incheon, 2022
2022
Later among the works it cites.
H.-J. Chang, S.-w. Yang, and H.-y. Lee, “DistilHuBERT: Speech representation learning by layer-wise distillation of hidden-unit bert,” in Proc. ICASSP , Toronto, 2022
2022
Later among the works it cites.
X. Yang, Q. Li, and P. C. Woodland, “Knowledge distillation for neural transducers from large self-supervised pre-trained models,” in Proc. ICASSP , Singapore, 2022
2022
Later among the works it cites.