Fetching the paper…
Reading the bibliography…
Representing speech as discretized units has numerous benefits in supporting downstream spoken language processing tasks.
The Structure of Tone
Zhiming Bao, · 1999
Earlier work this paper cites.
“Estimating or propagating gradients through stochastic neurons for conditional computation,”
Yoshua Bengio, Nicholas Léonard, and Aaron Courville, · 2013
Earlier work this paper cites.
“A ∗ sampling,”
Chris J Maddison, Daniel Tarlow, and Tom Minka, · 2014
Earlier work this paper cites.
“Librispeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Categorical reparameterization with gumbel-softmax,”
Eric Jang, Shixiang Gu, and Ben Poole, · 2016
Earlier work this paper cites.
“Neural discrete representation learning,”
Aaron van den Oord, Oriol Vinyals, and koray kavukcuoglu, · 2017
Earlier work this paper cites.
“AIShell-1: An open-source Mandarin speech corpus and a speech recognition baseline,”
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng, · 2017
Earlier work this paper cites.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2017
Earlier work this paper cites.
“The challenge of realistic music generation: modelling raw audio at scale,”
Sander Dieleman, Aaron van den Oord, and Karen Simonyan, · 2018
Earlier work this paper cites.
“Effectiveness of self-supervised pre-training for speech recognition,”
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed, · 2019
Earlier work this paper cites.
“vq-wav2vec: Self-supervised learning of discrete speech representations,”
Alexei Baevski, Steffen Schneider, and Michael Auli, · 2019
Earlier work this paper cites.
Ryan Eloff, André Nortje, Benjamin van Niekerk, Avashna Govender, Leanne Nortje, Arnu Pretorius, Elan Van Biljon, Ewald van der Westhuizen, Lisa van Staden, and Herman Kamper, · 2019
Earlier work this paper cites.
“g2pe,” https://github.com/Kyubyong/g2p
Kyubyong Park and Jongseok Kim, · 2019
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Jukebox: A generative model for music,”
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever, · 2020
Cited alongside, same era.
“HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Cited alongside, same era.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2020
Cited alongside, same era.
“Speech resynthesis from discrete disentangled self-supervised representations,”
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux, · 2021
Later among the works it cites.
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
Jaehyeon Kim, Jungil Kong, and Juhee Son, · 2021
Later among the works it cites.
“Self-supervised speech representation learning: A review,”
Abdelrahman Mohamed, Hung-yi Lee, Lasse Borgholt, Jakob D Havtorn, Joakim Edin, Christian Igel, Katrin Kirchhoff, Shang-Wen Li, Karen Livescu, Lars Maaløe, et al., · 2022
Later among the works it cites.
“WavLM: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2022
Later among the works it cites.
“SPIRAL: Self-supervised perturbation-invariant representation learning for speech pre-training,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“A neural grapheme-to-phoneme conversion package for Mandarin Chinese based on a new open benchmark dataset,”
Kyubyong Park and Seanie Lee, · 2020
Cited alongside, same era.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“W2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,”
Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu, · 2021
Cited alongside, same era.
“On generative spoken language modeling from raw audio,”
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al., · 2021
Cited alongside, same era.
“SuperB: Speech processing universal performance benchmark,”
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al., · 2021
Cited alongside, same era.
“Textless speech-to-speech translation on real data,”
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino, Jiatao Gu, et al., · 2021
Cited alongside, same era.
“Direct speech-to-speech translation with discrete units,”
Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Sravya Popuri, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, et al., · 2021
Cited alongside, same era.
Wenyong Huang, Zhenhe Zhang, Yu Ting Yeung, Xin Jiang, and Qun Liu, · 2022
Later among the works it cites.
“SQ-VAE: Variational bayes on discrete representation with self-annealed stochastic quantization,”
Yuhta Takida, Takashi Shibuya, WeiHsiang Liao, Chieh-Hsin Lai, Junki Ohmura, Toshimitsu Uesaka, Naoki Murata, Shusuke Takahashi, Toshiyuki Kumakura, and Yuki Mitsufuji, · 2022
Later among the works it cites.
“Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition,”
Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, et al., · 2022
Later among the works it cites.
“AudioLM: A language modeling approach to audio generation,”
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al., · 2023
Later among the works it cites.
“Finite scalar quantization: VQ-VAE made simple,”
Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen, · 2023
Later among the works it cites.
“Exploration of efficient end-to-end ASR using discretized input from self-supervised learning,”
Xuankai Chang, Brian Yan, Yuya Fujita, Takashi Maekaku, and Shinji Watanabe, · 2023
Later among the works it cites.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever, · 2023
Later among the works it cites.