Fetching the paper…
Reading the bibliography…
Transformer-based self-supervised models have achieved remarkable success in speech processing, but their large size and high inference cost present significant challenges for real-world deployment.
Y. LeCun, J. Denker, and S. Solla, “Optimal Brain Damage,” in Advances in Neural Information Processing Systems , vol. 2, 1989
1989
Earlier work this paper cites.
T. N. Sainath, B. Kingsbury, V. Sindhwani, E. Arisoy, and B. Ramabhadran, “Low-rank matrix factorization for Deep Neural Network training with high-dimensional output targets,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing , 2013, pp. 6655–6659
2013
Earlier work this paper cites.
D. Emily, Z. Wojciech, B. Joan, n. L. Yan, and F. Rob, “Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation,” in Advances in Neural Information Processing systems , 2014
2014
Earlier work this paper cites.
L. J. Ba and R. Caruana, “Do Deep Nets Really Need to be Deep?” in Advances in Neural Information Processing Systems , vol. 27, 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both Weights and Connections for Efficient Neural Network,” in Advances in Neural Information Processing Systems , C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, Eds., vol. 28, 2015
2015
Earlier work this paper cites.
S. Han, H. Mao, and W. J. Dally, “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding,” in International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An Unsupervised Autoregressive Model for Speech Representation Learning,” in Interspeech 2019 , 2019, pp. 146–150
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
P. Michel, O. Levy, and G. Neubig, “Are Sixteen Heads Really Better than One?” in Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
J. Frankle and M. Carbin, “The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
E. Voita, D. Talbot, F. Moiseev, R. Sennrich, and I. Titov, “Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , 2019, pp. 5797–5808
2019
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 12 449–12 460
2020
Cited alongside, same era.
M. A. Gordon, K. Duh, and N. Andrews, “Compressing BERT: Studying the Effects of Weight Pruning on Transfer Learning,” in Proceedings of the 5th Workshop on Representation Learning for NLP , 2020, pp. 143–155
2020
Cited alongside, same era.
A. Fan, E. Grave, and A. Joulin, “Reducing Transformer Depth on Demand with Structured Dropout,” in International Conference on Learning Representations , 2020
2020
Cited alongside, same era.
R. Wang, Q. Bai, J. Ao, L. Zhou, Z. Xiong, Z. Wei, Y. Zhang, T. Ko, and H. Li, “LightHuBERT: Lightweight and Configurable Speech Representation Learning with Once-for-All Hidden-Unit BERT,” in Interspeech 2022 , 2022, pp. 1686–1690
2022
Closest in time.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y. Qian, Y. Qian, J. Wu, M. Zeng, X. Yu, and F. Wei, “WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Closest in time.
K. Jang, S. Kim, S.-Y. Yun, and H. Kim, “Recycle-and-distill: Universal compression strategy for transformer-based speech SSL models with attention map reusing and masking distillation,” in Interspeech 2023 , 2023, pp. 316–320
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Liu, P. Zhou, Z. Zhao, Z. Wang, H. Deng, and Q. Ju, “FastBERT: a Self-distilling BERT with Adaptive Inference Time,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL) , 2020, pp. 6035–6044
2020
Cited alongside, same era.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3451–3460, 2021
2021
Cited alongside, same era.
C.-I. J. Lai, Y. Zhang, A. H. Liu, S. Chang, Y.-L. Liao, Y.-S. Chuang, K. Qian, S. Khurana, D. Cox, and J. Glass, “PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition,” in Advances in Neural Information Processing Systems , vol. 34, 2021, pp. 21 256–21 272
2021
Cited alongside, same era.
Z. Peng, A. Budhkar, I. Tuil, J. Levy, P. Sobhani, R. Cohen, and J. Nassour, “Shrinking bigfoot: Reducing wav2vec 2.0 footprint,” in Workshop on Simple and Efficient Natural Language Processing , 2021
2021
Cited alongside, same era.
S.-w. Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K.-t. Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H.-y. Lee, “SUPERB: Speech Processing Universal PERformance Benchmark,” in Interspeech 2021 , 2021, pp. 1194–1198
2021
Cited alongside, same era.
T. Chen, J. Frankle, S. Chang, S. Liu, Y. Zhang, M. Carbin, and Z. Wang, “The lottery tickets hypothesis for supervised and self-supervised pre-training in computer vision models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 16 306–16 316
2021
Cited alongside, same era.
H.-J. Chang, S.-w. Yang, and H.-y. Lee, “Distilhubert: Speech Representation Learning by Layer-Wise Distillation of Hidden-Unit Bert,” in ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 7087–7091
2022
Cited alongside, same era.
Y. Lee, K. Jang, J. Goo, Y. Jung, and H. R. Kim, “FitHuBERT: Going Thinner and Deeper for Knowledge Distillation of Speech Self-Supervised Models,” in Interspeech 2022 , 2022, pp. 3588–3592
2022
Cited alongside, same era.
2023
Closest in time.
Y. Peng, K. Kim, F. Wu, P. Sridhar, and S. Watanabe, “Structured Pruning of Self-Supervised Pre-Trained Models for Speech Recognition and Understanding,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023, pp. 1–5
2023
Closest in time.
H. Wang, S. Wang, W.-Q. Zhang, H. Suo, and Y. Wan, “Task-Agnostic Structured Pruning of Speech Representation Models,” in Interspeech 2023 , 2023, pp. 231–235
2023
Closest in time.
T.-Q. Lin, H.-Y. Lee, and H. Tang, “Melhubert: A simplified hubert on mel spectrograms,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) , 2023, pp. 1–8
2023
Closest in time.
K. Jang, S. Kim, and H. Kim, “STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 10 721–10 725
2024
Closest in time.
T.-Q. Lin, G.-T. Lin, H.-y. Lee, and H. Tang, “Property Neurons in Self-Supervised Speech Transformers,” in 2024 IEEE Spoken Language Technology Workshop (SLT) , 2024, pp. 401–408
2024
Closest in time.
T.-Q. Lin, H.-y. Lee, and H. Tang, “DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models,” in Interspeech 2024 , 2024, pp. 4513–4517
2024
Closest in time.