Fetching the paper…
Reading the bibliography…
Pre-training with self-supervised models, such as Hidden-unit BERT (HuBERT) and wav2vec 2.0, has brought significant improvements in automatic speech recognition (ASR).
A. Graves, S. Fernandez, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proc. ICML , 2006, pp. 369–376
2006
Earlier work this paper cites.
C. Bucila, R. Caruana, and A. Niculescu-Mizil, “Model compression,” in Proc. ACM SIGKDD , 2006, p. 535–541
2006
Earlier work this paper cites.
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” in Proc. NIPS Workshop Deep Learn. , 2014
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in Proc. ICASSP , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
S. Han, H. Mao, and W. J. Dally, “Deep compression: compressing deep neural network with pruning, trained quantization and huffman coding,” in Proc. ICLR , 2016
2016
Earlier work this paper cites.
J. Wu, C. Leng, Y. Wang, Q. Hu, and J. Cheng, “Quantized convolutional neural networks for mobile devices,” in Proc. CVPR , 2016, p. 4820–4828
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. NIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning filters for efficient convnets,” in Proc. ICLR , 2017
2017
Earlier work this paper cites.
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli, “Fairseq: a fast, extensible toolkit for sequence modeling,” in Proc. NAACL , 2019, p. 48–53
2019
Earlier work this paper cites.
S. Kaya, Y. Hong and T. Dumitras, “Shallow-deep networks: understanding and mitigating network overthinking,” in Proc. ICML , 2019
2019
Earlier work this paper cites.
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “Wav2vec 2.0: a framework for self-supervised learning of speech representations,” in Proc. NIPS , 2020
2020
Cited alongside, same era.
J. Xin, R. Tang, J. Lee, Y. Yu, and J. Lin, “Deebert: dynamic early exiting for accelerating bERT inference,” in Proc. ACL , 2020, pp. 2246–2251
2020
Cited alongside, same era.
W. Liu, P. Zhou, Z. Zhao, Z. Wang, H. Deng, and Q. Ju, “Fastbert: a self-distilling bert with adaptive inference time,” in Proc. ACL , 2020, p. 6035–6044
2020
Cited alongside, same era.
R. Schwartz, G. Stanovsky, S. Swayamdipta, J. Dodge, and N. A. Smith, “The right tool for the job: matching model and instance complexities,” in Proc. ACL , 2020, p. 6640–6651
2020
Cited alongside, same era.
W. Zhou, C. Xu, T. Ge, J. McAuley, K. Xu, and F. Wei, “Bert loses patience: fast and robust inference with early exit,” in Proc. NIPS , 2020
J. Lee and S. Watanabe, “Intermediate loss regularization for ctc-based speech recognition,” in Proc. ICASSP , 2021
2021
Later among the works it cites.
J. Nozaki and T. Komatsu, “Relaxing the conditional independence assumption of ctc-based asr by conditioning on intermediate predictions,” in Proc. INTERSPEECH , 2021
2021
Later among the works it cites.
J. Lee, J. Kang, and S. Watanabe, “Layer pruning on demand with intermediate ctc,” in Proc. INTERSPEECH , 2021
2021
Later among the works it cites.
2022
Closest in time.
C. Wang, Y. Wu, S. Chen, S. Liu, J. Li, Y. Qian, and Z. Yang, “Improving self-supervised learning for speech recognition with intermediate layer supervision,” in Proc. ICASSP , 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
A. Fan, E. Grave, and A. Joulin, “Reducing transformer depth on demand with structured dropout,” in Proc. ICLR , 2020
2020
Cited alongside, same era.
W. N. Hsu, B. Bolte, Y. H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
J. Xin, R. Tang, Y. Yu, , and J. Lin, “Berxit: early exiting for bert with better fine-tuning and extension to regression,” in Proc. EACL , 2021, p. 91–104
2021
Cited alongside, same era.
A. Li, C. Zheng, L. Zhang, and X. Li, “Learning to inference with early exit in the progressive speech enhancement,” in Proc. EUSIPCO , 2021, pp. 466–470
2021
Cited alongside, same era.
S. Chen, Y. Wu, Z. Chen, T. Yoshioka, S. Liu, J. Li, and X. Yu, “Don’t shoot butterfly with rifles: multi-channel continuous speech separation with early exit transformer,” in Proc. ICASSP , 2021, pp. 6139–6143
2021
Cited alongside, same era.
2022
Closest in time.
A. Baevski, W. Hsu, Q. Xu, A. Babu, J. Gu, and M. Aulim, “Data2vec: a general framework for self-supervised learning in speech, vision and language,” in Proc. ICML , 2022, pp. 1298–1312
2022
Closest in time.
R. Tang, K. Kumar, J. Xin, P. Vyas, W. Li, G. Yang, Y. Mao, C. Murray, and J. Lin, “Temporal early exiting for streaming speech commands recognition,” in Proc. ICASSP , 2022
2022
Closest in time.
J. W. Yoon, B. J. Woo, S. Ahn, H. Lee, and N. S. Kim, “Inter-kd: intermediate knowledge distillation for ctc-based automatic speech recognition,” in Proc. IEEE SLT , 2022
2022
Closest in time.
H. Chang, S. Yang, and H. Lee, “Distilhubert: speech representation learning by layer-wise distillation of hidden-unit bert,” in Proc. ICASSP , 2022
2022
Closest in time.