Fetching the paper…
Reading the bibliography…
Self-supervised learning (SSL) has achieved notable success in many speech processing tasks, but the large model size and heavy computational cost hinder the deployment.
R. Reed, “Pruning algorithms-a survey,” IEEE Trans. on Neural Networks , vol. 4, no. 5, pp. 740–747, 1993
1993
Earlier work this paper cites.
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in Proc. ICASSP , 2015
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. NeurIPS , 2017
2017
Earlier work this paper cites.
C. Louizos, M. Welling, and D. P. Kingma, “Learning Sparse Neural Networks through L0 Regularization,” in ICLR , 2018
2018
Earlier work this paper cites.
Z. Liu, M. Sun, T. Zhou, G. Huang, and T. Darrell, “Rethinking the value of network pruning,” in Proc. ICLR , 2019
2019
Earlier work this paper cites.
P. Michel, O. Levy, and G. Neubig, “Are sixteen heads really better than one?” in Proc. NeurIPS , 2019
2019
Earlier work this paper cites.
S. Wang, P. Lin, R. Hu et al. , “Acceleration of LSTM With Structured Pruning Method on FPGA,” IEEE Access , 2019
2019
Earlier work this paper cites.
A. Paszke et al. , “Pytorch: An imperative style, high-performance deep learning library,” Proc. NeurIPS , 2019
2019
Earlier work this paper cites.
M. Ott et al. , “fairseq: A fast, extensible toolkit for sequence modeling,” in Proc. NAACL-HLT: Demonstrations , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. NeurIPS , 2020
2020
Earlier work this paper cites.
A. Fan, E. Grave, and A. Joulin, “Reducing transformer depth on demand with structured dropout,” in Proc. ICLR , 2020
2020
Earlier work this paper cites.
Z. Wang, J. Wohlwend, and T. Lei, “Structured Pruning of Large Language Models,” in Proc. EMNLP , 2020
2020
Earlier work this paper cites.
P. Dong, S. Wang, W. Niu et al. , “Rtmobile: Beyond real-time mobile acceleration of rnns for speech recognition,” in ACM/IEEE Design Automation Conference (DAC) , 2020
2020
Cited alongside, same era.
H. Cai, C. Gan, T. Wang, Z. Zhang, and S. Han, “Once-for-all: Train one network and specialize it for efficient deployment,” in Proc. ICLR , 2020
2020
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” IEEE/ACM Trans. Audio, Speech, Lang. Process. , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
A. Baevski, W.-N. Hsu, A. Conneau, and M. Auli, “Unsupervised speech recognition,” in Proc. NeurIPS , 2021
2021
Cited alongside, same era.
Z. Huang, S. Watanabe, S.-w. Yang, P. García, and S. Khudanpur, “Investigating Self-Supervised Learning for Speech Enhancement and Separation,” in Proc. ICASSP , 2022
2022
Later among the works it cites.
Y. Peng, S. Arora, Y. Higuchi, Y. Ueda, S. Kumar, K. Ganesan, S. Dalmia, X. Chang, and S. Watanabe, “A Study on the Integration of Pre-trained SSL, ASR, LM and SLU Models for Spoken Language Understanding,” in Proc. SLT , 2022
2022
Later among the works it cites.
H.-J. Chang, S.-w. Yang, and H.-y. Lee, “DistilHuBERT: Speech representation learning by layer-wise distillation of hidden-unit BERT,” in Proc. ICASSP , 2022
2022
Later among the works it cites.
Y. Lee, K. Jang, J. Goo, Y. Jung, and H. R. Kim, “FitHuBERT: Going Thinner and Deeper for Knowledge Distillation of Speech Self-Supervised Models,” in Proc. Interspeech , 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. wen Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. yi Lee, “SUPERB: Speech Processing Universal PERformance Benchmark,” in Proc. Interspeech , 2021
2021
Cited alongside, same era.
X. Chang, T. Maekaku, P. Guo, J. Shi, Y.-J. Lu, A. S. Subramanian, T. Wang, S.-w. Yang, Y. Tsao, H.-y. Lee et al. , “An exploration of self-supervised pretrained representations for end-to-end speech recognition,” in Proc. ASRU , 2021
2021
Cited alongside, same era.
K. Tan and D. Wang, “Compressing Deep Neural Networks for Efficient Speech Enhancement,” in Proc. ICASSP , 2021
2021
Cited alongside, same era.
C.-I. J. Lai, Y. Zhang, A. H. Liu, S. Chang, Y.-L. Liao, Y.-S. Chuang, K. Qian, S. Khurana, D. Cox, and J. Glass, “PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition,” in Proc. NeurIPS , 2021
2021
Cited alongside, same era.
A. Pasad, J.-C. Chou, and K. Livescu, “Layer-Wise Analysis of a Self-Supervised Speech Representation Model,” in Proc. ASRU , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
S. Zhang, E. Loweimi, P. Bell, and S. Renals, “On the usefulness of self-attention for automatic speech recognition with transformers,” in Proc. SLT , 2021
2021
Cited alongside, same era.
A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y. Saraf, J. Pino, A. Baevski, A. Conneau, and M. Auli, “XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,” in Proc. Interspeech , 2022
2022
Cited alongside, same era.
T. Ashihara, T. Moriya, K. Matsuura, and T. Tanaka, “Deep versus Wide: An Analysis of Student Architectures for Task-Agnostic Knowledge Distillation of Self-Supervised Speech Models,” in Proc. Interspeech , 2022
2022
Later among the works it cites.
M. Xia, Z. Zhong, and D. Chen, “Structured Pruning Learns Compact and Accurate Models,” in Proc. ACL , 2022
2022
Later among the works it cites.
R. Wang, Q. Bai, J. Ao, L. Zhou, Z. Xiong, Z. Wei, Y. Zhang, T. Ko, and H. Li, “LightHuBERT: Lightweight and Configurable Speech Representation Learning with Once-for-All Hidden-Unit BERT,” in Proc. Interspeech , 2022
2022
Later among the works it cites.
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, “Data2vec: A general framework for self-supervised learning in speech, vision and language,” in Proc. ICML , 2022
2022
Later among the works it cites.
K. Shim, J. Choi, and W. Sung, “Understanding the role of self attention for efficient speech recognition,” in Proc. ICLR , 2022
2022
Later among the works it cites.
Y. Peng, S. Dalmia, I. Lane, and S. Watanabe, “Branchformer: Parallel MLP-attention architectures to capture local and global context for speech recognition and understanding,” in Proc. ICML , 2022
2022
Later among the works it cites.
T. Maekaku, Y. Fujita, Y. Peng, and S. Watanabe, “Attention Weight Smoothing Using Prior Distributions for Transformer-Based End-to-End ASR,” in Proc. Interspeech , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
F. Wu, K. Kim, J. Pan, K. J. Han, K. Q. Weinberger, and Y. Artzi, “Performance-efficiency trade-offs in unsupervised pre-training for speech recognition,” in Proc. ICASSP , 2022
2022
Later among the works it cites.
Y. Peng, K. Kim, F. Wu, P. Sridhar, and S. Watanabe, “Structured Pruning of Self-Supervised Pre-trained Models for Speech Recognition and Understanding,” in Proc. ICASSP , 2023
2023
Closest in time.