Fetching the paper…
Reading the bibliography…
Self-supervised learning (SSL) proficiency in speech-related tasks has driven research into utilizing discrete tokens for speech tasks like recognition and translation, which offer lower storage requirements and great potential to employ natural language processing techniques.
“LibriSpeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Earlier work this paper cites.
“AISHELL-1: An open-source mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Earlier work this paper cites.
“Effectiveness of self-supervised pre-training for speech recognition,”
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed, · 2019
Earlier work this paper cites.
“BERT: pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Earlier work this paper cites.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Earlier work this paper cites.
“LibriTTS: A corpus derived from LibriSpeech for text-to-speech,”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, et al., · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Earlier work this paper cites.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Alexei Baevski, Steffen Schneider, and Michael Auli, · 2020
Earlier work this paper cites.
“HiFi-GAN: High-fidelity denoising and dereverberation based on speech deep features in adversarial networks,”
Jiaqi Su, Zeyu Jin, and Adam Finkelstein, · 2020
Earlier work this paper cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang, · 2020
Earlier work this paper cites.
“Rnn-Transducer with stateless prediction network,”
Mohammadreza Ghodsi, Xiaofeng Liu, James Apfel, Rodrigo Cabrera, and Eugene Weinstein, · 2020
Earlier work this paper cites.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“SoundStream: An end-to-end neural audio codec,”
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi, · 2021
Cited alongside, same era.
“SUPERB: Speech processing universal performance benchmark,”
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al., · 2021
Cited alongside, same era.
“FastSpeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2021
Cited alongside, same era.
“GigaSpeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,”
Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du, Wei-Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, et al., · 2021
Cited alongside, same era.
“Pruned RNN-T for fast, memory-efficient asr training,”
Fangjun Kuang, Liyong Guo, Wei Kang, Long Lin, Mingshuang Luo, Zengwei Yao, and Daniel Povey, · 2022
Later among the works it cites.
“MT4SSL: Boosting Self-Supervised Speech Representation Learning by Integrating Multiple Targets,”
Ziyang Ma, Zhisheng Zheng, Changli Tang, Yujin Wang, and Xie Chen, · 2023
Closest in time.
“Pushing the limits of unsupervised unit discovery for SSL speech representation,”
Ziyang Ma, Zhisheng Zheng, Guanrou Yang, Yu Wang, Chao Zhang, and Xie Chen, · 2023
Closest in time.
“AudioLM: a language modeling approach to audio generation,”
Zalan Borsos, Raphael Marinier, Damien Vincent, Eugene Kharitonov, et al., · 2023
Closest in time.
“Exploration of efficient end-to-end asr using discretized input from self-supervised learning,”
Xuankai Chang, Brian Yan, Yuya Fujita, Takashi Maekaku, and Shinji Watanabe, · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Speech recognition with next-generation kaldi (k2, lhotse, icefall),”
Daniel Povey, Piotr Zelasko, and Sanjeev Khudanpur, · 2021
Cited alongside, same era.
“Self-supervised speech representation learning: A review,”
Abdelrahman Mohamed, Hung-yi Lee, Lasse Borgholt, Jakob D Havtorn, Joakim Edin, Christian Igel, Katrin Kirchhoff, Shang-Wen Li, Karen Livescu, Lars Maaløe, et al., · 2022
Cited alongside, same era.
“WavLM: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2022
Cited alongside, same era.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Cited alongside, same era.
“VQTTS: High-fidelity text-to-speech synthesis with self-supervised vq acoustic feature,”
Chenpeng Du, Yiwei Guo, Xie Chen, and Kai Yu, · 2022
Cited alongside, same era.
“Recent advances in end-to-end automatic speech recognition,”
Jinyu Li et al., · 2022
Cited alongside, same era.
Chenpeng Du, Yiwei Guo, Feiyu Shen, Zhijun Liu, Zheng Liang, Xie Chen, Shuai Wang, Hui Zhang, and Kai Yu, · 2023
Closest in time.
“Expresso: A benchmark and analysis of discrete expressive speech resynthesis,”
Tu Anh Nguyen, Wei-Ning Hsu, Antony D’Avirro, Bowen Shi, Itai Gat, Maryam Fazel-Zarani, Tal Remez, Jade Copet, Gabriel Synnaeve, Michael Hassid, et al., · 2023
Closest in time.
“VioLA: Unified codec language models for speech recognition, synthesis, and translation,”
Tianrui Wang, Long Zhou, Ziqiang Zhang, Yu Wu, Shujie Liu, Yashesh Gaur, Zhuo Chen, Jinyu Li, and Furu Wei, · 2023
Closest in time.
“Zipformer: A faster and better encoder for automatic speech recognition,”
Zengwei Yao, Liyong Guo, Xiaoyu Yang, Wei Kang, Fangjun Kuang, Yifan Yang, Zengrui Jin, Long Lin, and Daniel Povey, · 2023
Closest in time.
“LongFNT: Long-form speech recognition with factorized neural Transducer,”
Xun Gong, Yu Wu, Jinyu Li, Shujie Liu, et al., · 2023
Closest in time.
“High-fidelity audio compression with improved RVQGAN,”
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar, · 2023
Closest in time.