Fetching the paper…
Reading the bibliography…
We present the SUPERB challenge at SLT 2022, which aims at learning self-supervised speech representation for better performance, generalization, and efficiency.
“Librispeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“GLUE: A multi-task benchmark and analysis platform for natural language understanding,”
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman, · 2018
Earlier work this paper cites.
“BERT: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Earlier work this paper cites.
“An unsupervised autoregressive model for speech representation learning,”
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass, · 2019
Earlier work this paper cites.
“Improving transformer-based speech recognition using unsupervised pre-training,”
Dongwei Jiang, Xiaoning Lei, Wubo Li, Ne Luo, Yuxuan Hu, Wei Zou, and Xiangang Li, · 2019
Earlier work this paper cites.
“Learning problem-agnostic speech representations from multiple self-supervised tasks,”
Santiago Pascual, Mirco Ravanelli, Joan Serrà, Antonio Bonafonte, and Yoshua Bengio, · 2019
Earlier work this paper cites.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Earlier work this paper cites.
“A large-scale study of representation learning with the visual task adaptation benchmark,”
Xiaohua Zhai, Joan Puigcerver, Alexander Kolesnikov, Pierre Ruyssen, Carlos Riquelme, Mario Lucic, Josip Djolonga, et al., · 2019
Earlier work this paper cites.
“Parameter-efficient transfer learning for NLP,”
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, et al., · 2019
Earlier work this paper cites.
“Language models are few-shot learners,”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, et al., · 2020
Earlier work this paper cites.
“A simple framework for contrastive learning of visual representations,”
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton, · 2020
Earlier work this paper cites.
“Bootstrap your own latent-a new approach to self-supervised learning,”
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, et al., · 2020
Earlier work this paper cites.
“Pushing the limits of semi-supervised learning for automatic speech recognition,”
Yu Zhang, James Qin, Daniel S Park, Wei Han, Chung-Cheng Chiu, Ruoming Pang, Quoc V Le, et al., · 2020
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Earlier work this paper cites.
“Generative pre-training for speech with autoregressive predictive coding,”
Yu-An Chung and James Glass, · 2020
Earlier work this paper cites.
“Improved speech representations with multi-target autoregressive predictive coding,”
Yu-An Chung and James Glass, · 2020
Earlier work this paper cites.
“Vector-Quantized Autoregressive Predictive Coding,”
Yu-An Chung, Hao Tang, and James Glass, · 2020
Earlier work this paper cites.
“Deep contextualized acoustic representations for semi-supervised speech recognition,”
Shaoshi Ling, Yuzong Liu, Julian Salazar, and Katrin Kirchhoff, · 2020
Earlier work this paper cites.
“DeCoAR 2.0: Deep contextualized acoustic representations with vector quantization,”
Shaoshi Ling and Yuzong Liu, · 2020
Earlier work this paper cites.
“Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,”
Andy T Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu, and Hung-yi Lee, · 2020
Cited alongside, same era.
“Audio ALBERT: A lite BERT for self-supervised learning of audio representation,”
Po-Han Chi, Pei-Hung Chung, Tsung-Han Wu, Chun-Cheng Hsieh, Yen-Hao Chen, Shang-Wen Li, and Hung-yi Lee, · 2020
Cited alongside, same era.
“Speech-XLNet: Unsupervised Acoustic Model Pretraining for Self-Attention Networks,”
Xingchen Song, Guangsen Wang, Yiheng Huang, Zhiyong Wu, Dan Su, and Helen Meng, · 2020
Cited alongside, same era.
“Multi-task self-supervised learning for robust speech recognition,”
Mirco Ravanelli, Jianyuan Zhong, Santiago Pascual, Pawel Swietojanski, Joao Monteiro, Jan Trmal, and Yoshua Bengio, · 2020
Cited alongside, same era.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Alexei Baevski, Steffen Schneider, and Michael Auli, · 2020
Cited alongside, same era.
“Self-supervised speech representation learning: A review,”
Abdelrahman Mohamed, Hung-yi Lee, Lasse Borgholt, Jakob D Havtorn, Joakim Edin, Christian Igel, Katrin Kirchhoff, et al., · 2022
Closest in time.
“UniSpeech-SAT: Universal speech representation learning with speaker aware pre-training,”
Sanyuan Chen, Yu Wu, Chengyi Wang, Zhengyang Chen, Zhuo Chen, Shujie Liu, Jian Wu, et al., · 2022
Closest in time.
“SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities,”
Hsiang-Sheng Tsai, Heng-Jui Chang, Wen-Chin Huang, Zili Huang, Kushal Lakhotia, Shu-wen Yang, Shuyan Dong, et al., · 2022
Closest in time.
“Self-supervised representation learning for speech processing,”
Hung-yi Lee, Abdelrahman Mohamed, Shinji Watanabe, Tara Sainath, Karen Livescu, Shang-Wen Li, Shu-wen Yang, et al., · 2022
Closest in time.
“WavLM: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, et al., · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“The zero resource speech challenge 2020: Discovering discrete subword and word units,”
Ewan Dunbar, Julien Karadayi, Mathieu Bernard, Xuan-Nga Cao, Robin Algayres, Lucas Ondel, Laurent Besacier, et al., · 2020
Cited alongside, same era.
“Transformer VQ-VAE for Unsupervised Unit Discovery and Speech Synthesis: ZeroSpeech 2020 Challenge,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2020
Cited alongside, same era.
“Vector-Quantized Neural Networks for Acoustic Unit Discovery in the ZeroSpeech 2020 Challenge,”
Benjamin van Niekerk, Leanne Nortje, and Herman Kamper, · 2020
Cited alongside, same era.
“CIF: Continuous integrate-and-fire for end-to-end speech recognition,”
Linhao Dong and Bo Xu, · 2020
Cited alongside, same era.
Yingzhi Wang, Abdelmoumene Boumadane, and Abdelwahab Heba, · 2021
Cited alongside, same era.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“SUPERB: Speech Processing Universal PERformance Benchmark,”
Shu wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, et al., · 2021
Cited alongside, same era.
Closest in time.
“Ch-marl: A multimodal benchmark for cooperative, heterogeneous multi-agent reinforcement learning,”
Vasu Sharma, Prasoon Goyal, Kaixiang Lin, Govind Thattai, Qiaozi Gao, and Gaurav S Sukhatme, · 2022
Closest in time.
“data2vec: A general framework for self-supervised learning in speech, vision and language,”
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli, · 2022
Closest in time.
“Self-supervised representation learning for speech using visual grounding and masked language modeling,”
Puyuan Peng and David Harwath, · 2022
Closest in time.
“DistilHuBERT: Speech Representation Learning by Layer-wise Distillation of Hidden-unit BERT,”
Heng-Jui Chang, Shu wen Yang, and Hung yi Lee, · 2022
Closest in time.
“LightHuBERT: Lightweight and Configurable Speech Representation Learning with Once-for-All Hidden-Unit BERT,”
Rui Wang, Qibing Bai, Junyi Ao, Long Zhou, Zhixiang Xiong, Zhihua Wei, Yu Zhang, et al., · 2022
Closest in time.
“SpeechCLIP: Integrating speech with pre-trained vision and language model,”
Yi-Jen Shih, Hsuan-Fu Wang, Heng-Jui Chang, Layne Berry, Hung-yi Lee, and David Harwath, · 2022
Closest in time.
“Improving distortion robustness of self-supervised speech processing tasks with domain adaptation,”
Kuan Po Huang, Yu-Kuan Fu, Yu Zhang, and Hung yi Lee, · 2022
Closest in time.
“Improving generalizability of distilled self-supervised speech processing models under distorted settings,”
Kuan-Po Huang, Yu-Kuan Fu, Tsu-Yuan Hsu, Fabian Ritter Gutierrez, Fan-Lin Wang, Liang-Hsuan Tseng, Yu Zhang, et al., · 2022
Closest in time.
“ccc-wav2vec 2.0: Clustering aided cross contrastive self-supervised learning of speech representations,”
Vasista Sai Lodagala, Sreyan Ghosh, and S. Umesh, · 2022
Closest in time.
“Silence is sweeter than speech: Self-supervised model using silence to store speaker information,”
Chi-Luen Feng, Po chun Hsu, and Hung yi Lee, · 2022
Closest in time.
“On compressing sequences for self-supervised speech models,”
Yen Meng, Hsuan-Jui Chen, Jiatong Shi, Shinji Watanabe, Paola Garcia, Hung-yi Lee, and Hao Tang, · 2022
Closest in time.
“Analyzing the robustness of unsupervised speech recognition,”
Guan-Ting Lin, Chan-Jan Hsu, Da-Rong Liu, Hung-yi Lee, and Yu Tsao, · 2022
Closest in time.
“Towards end-to-end unsupervised speech recognition,”
Alexander H Liu, Wei-Ning Hsu, Michael Auli, and Alexei Baevski, · 2022
Closest in time.
“Exploring efficient-tuning methods in self-supervised speech models,”
Zih-Ching Chen, Chin-Lun Fu, Chih Ying Liu, Shang-Wen Li, and Hung-yi Lee, · 2022
Closest in time.
“Extracting speaker and emotion information from self-supervised speech models via channel-wise correlations,”
T. Stafylakis, L. Mošner, S. Kakouros, O. Plchot, L. Burget, and J. Černocký, · 2022
Closest in time.