Fetching the paper…
Reading the bibliography…
Self-supervised models have had great success in learning speech representations that can generalize to various downstream tasks.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Noise-contrastive estimation: A new estimation principle for unnormalized statistical models,”
Michael Gutmann and Aapo Hyvärinen, · 2010
Earlier work this paper cites.
“Learning the speech front-end with raw waveform CLDNNs,”
Tara Sainath, Ron J. Weiss, Kevin Wilson, Andrew W. Senior, and Oriol Vinyals, · 2015
Earlier work this paper cites.
“Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi.,”
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger, · 2017
Earlier work this paper cites.
“Speaker recognition from raw waveform with SincNet,”
Mirco Ravanelli and Yoshua Bengio, · 2018
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Mixed precision training,”
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al., · 2018
Earlier work this paper cites.
“An unsupervised autoregressive model for speech representation learning,”
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass, · 2019
Earlier work this paper cites.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Earlier work this paper cites.
“Improving transformer-based speech recognition using unsupervised pre-training,”
Dongwei Jiang, Xiaoning Lei, Wubo Li, Ne Luo, Yuxuan Hu, Wei Zou, and Xiangang Li, · 2019
Earlier work this paper cites.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Alexei Baevski, Steffen Schneider, and Michael Auli, · 2019
Earlier work this paper cites.
“Reducing transformer depth on demand with structured dropout,”
Angela Fan, Edouard Grave, and Armand Joulin, · 2019
Earlier work this paper cites.
“Fairseq: A fast, extensible toolkit for sequence modeling,”
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli, · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Vector-quantized autoregressive predictive coding,”
Yu-An Chung, Hao Tang, and James Glass, · 2020
Cited alongside, same era.
“DeCoAR 2.0: Deep contextualized acoustic representations with vector quantization,”
Shaoshi Ling and Yuzong Liu, · 2020
Cited alongside, same era.
“Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,”
Andy T. Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu, and Hung-yi Lee, · 2020
Cited alongside, same era.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“Autoregressive Co-Training for Learning Discrete Speech Representations,”
Sung-Lin Yeh and Hao Tang, · 2022
Closest in time.
“On compressing sequences for self-supervised speech models,”
Yen Meng, Hsuan-Jui Chen, Jiatong Shi, Shinji Watanabe, Paola Garcia, Hung-yi Lee, and Hao Tang, · 2022
Closest in time.
“Performance-efficiency trade-offs in unsupervised pre-training for speech recognition,”
Felix Wu, Kwangyoun Kim, Jing Pan, Kyu J Han, Kilian Q. Weinberger, and Yoav Artzi, · 2022
Closest in time.
“SSAST: Self-supervised audio spectrogram transformer,”
Yuan Gong, Cheng-I Lai, Yu-An Chung, and James Glass, · 2022
Closest in time.
“Self-supervised learning with random-projection quantizer for speech recognition,”
Chung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu, and Yonghui Wu, · 2022
Closest in time.
“Masked spectrogram prediction for self-supervised audio pre-training,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“SUPERB: Speech Processing Universal PERformance Benchmark,”
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al., · 2021
Cited alongside, same era.
“TERA: Self-supervised learning of transformer encoder representation for speech,”
Andy T. Liu, Shang-Wen Li, and Hung-yi Lee, · 2021
Cited alongside, same era.
“Layer-wise analysis of a self-supervised speech representation model,”
Ankita Pasad, Ju-Chieh Chou, and Karen Livescu, · 2021
Cited alongside, same era.
“WavLM: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2022
Cited alongside, same era.
“Masked autoencoders that listen,”
Po-Yao Huang, Hu Xu, Juncheng Li, Alexei Baevski, Michael Auli, Wojciech Galuba, Florian Metze, and Christoph Feichtenhofer, · 2022
Cited alongside, same era.
“BigSSL: Exploring the frontier of large-scale semi-supervised learning for automatic speech recognition,”
Yu Zhang, Daniel S Park, Wei Han, James Qin, Anmol Gulati, Joel Shor, Aren Jansen, Yuanzhong Xu, Yanping Huang, Shibo Wang, et al., · 2022
Cited alongside, same era.
“MAE-AST: Masked autoencoding audio spectrogram transformer,”
Alan Baade, Puyuan Peng, and David Harwath, · 2022
Cited alongside, same era.
Dading Chong, Helin Wang, Peilin Zhou, and Qingcheng Zeng, · 2023
Closest in time.
“Structured pruning of self-supervised pre-trained models for speech recognition and understanding,”
Yifan Peng, Kwangyoun Kim, Felix Wu, Prashant Sridhar, and Shinji Watanabe, · 2023
Closest in time.
“DPHuBERT: Joint distillation and pruning of self-supervised speech models,”
Yifan Peng, Yui Sudo, Shakeel Muhammad, and Shinji Watanabe, · 2023
Closest in time.
“On the (in)efficiency of acoustic feature extractors for self-supervised speech representation learning,”
Titouan Parcollet, Shucong Zhang, Rogier van Dalen, Alberto Gil C. P. Ramos, and Sourav Bhattacharya, · 2023
Closest in time.
“Front-end adapter: Adapting front-end input of speech based self-supervised learning for speech recognition,”
Xie Chen, Ziyang Ma, Changli Tang, Yujin Wang, and Zhisheng Zheng, · 2023
Closest in time.
“Comparative layer-wise analysis of self-supervised speech models,”
Ankita Pasad, Bowen Shi, and Karen Livescu, · 2023
Closest in time.
“On the utility of self-supervised models for prosody-related tasks,”
Guan-Ting Lin, Chi-Luen Feng, Wei-Ping Huang, Yuan Tseng, Tzu-Han Lin, Chen-An Li, Hung-yi Lee, and Nigel G Ward, · 2023
Closest in time.
“Reducing Barriers to Self-Supervised Learning: HuBERT Pre-training with Academic Compute,”
William Chen, Xuankai Chang, Yifan Peng, Zhaoheng Ni, Soumi Maiti, and Shinji Watanabe, · 2023
Closest in time.