Fetching the paper…
Reading the bibliography…
This paper proposes a novel semi-supervised TTS framework, QS-TTS, to improve TTS quality with lower supervised data requirements via Vector-Quantized Self-Supervised Speech Representation Learning (VQ-S3RL) utilizing more unlabeled speech audio.
Mel-cepstral distance measure for objective speech quality assessment
Robert Kubichek · 1993
Earlier work this paper cites.
Text-to-speech technology in human-computer interaction
Lehlohonolo Mohasi and Daniel Mashao · 2006
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jegou, Matthijs Douze, and Cordelia Schmid · 2010
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Neural discrete representation learning
Aäron van den Oord, Oriol Vinyals, and Koray Kavukcuoglu · 2017
Earlier work this paper cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Earlier work this paper cites.
Listening while speaking: Speech chain by deep learning
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura · 2017
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Natural TTS synthesis by conditioning wavenet on Mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Earlier work this paper cites.
Understanding disentangling in beta-vae
Christopher P Burgess, Irina Higgins, Arka Pal, Loic Matthey, Nick Watters, Guillaume Desjardins, and Alexander Lerchner · 2018
Earlier work this paper cites.
Fr \ \backslash ’echet audio distance: A metric for evaluating music enhancement algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · 2018
Earlier work this paper cites.
Semi-supervised training for improving data efficiency in end-to-end speech synthesis
Yu-An Chung, Yuxuan Wang, Wei-Ning Hsu, Yu Zhang, and RJ Skerry-Ryan · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Earlier work this paper cites.
Tagged back-translation
Isaac Caswell, Ciprian Chelba, and David Grangier · 2019
Earlier work this paper cites.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli · 2019
Earlier work this paper cites.
Low bit-rate speech coding with VQ-VAE and a WaveNet decoder
Cristina Gârbacea, Aäron van den Oord, Yazhe Li, Felicia SC Lim, Alejandro Luebs, Oriol Vinyals, and Thomas C Walters · 2019
Earlier work this paper cites.
Generating diverse high-fidelity images with vq-vae-2
Ali Razavi, Aaron van den Oord, and Oriol Vinyals · 2019
Earlier work this paper cites.
FastSpeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Earlier work this paper cites.
Simultaneous speech-to-speech translation system with neural incremental ASR, MT, and TTS
Katsuhito Sudoh, Takatomo Kano, Sashi Novitasari, Tomoya Yanagita, Sakriani Sakti, and Satoshi Nakamura · 2020
Earlier work this paper cites.
Semi-supervised learning based on hierarchical generative models for end-to-end speech synthesis
Takato Fujimoto, Shinji Takaki, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda · 2020
Earlier work this paper cites.
Semi-supervised speaker adaptation for end-to-end speech synthesis with pretrained models
Katsuki Inoue, Sunao Hara, Masanobu Abe, Tomoki Hayashi, Ryuichi Yamamoto, and Shinji Watanabe · 2020
Earlier work this paper cites.
Semi-supervised learning for multi-speaker text-to-speech synthesis using discrete speech representation
Tao Tu, Yuan-Jui Chen, Alexander H Liu, and Hung-yi Lee · 2020
Cited alongside, same era.
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli · 2020
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Cited alongside, same era.
A survey on contrastive self-supervised learning
Ashish Jaiswal, Ashwin Ramesh Babu, Mohammad Zaki Zadeh, Debapriya Banerjee, and Fillia Makedon · 2020
Cited alongside, same era.
One-shot voice conversion by vector quantization
Da-Yi Wu and Hung-yi Lee · 2020
Cited alongside, same era.
UnivNet: A neural vocoder with multi-resolution spectrogram discriminators for high-fidelity waveform generation
Won Jang, Dan Lim, Jaesam Yoon, Bongwan Kim, and Juntae Kim · 2021
Later among the works it cites.
End-to-end adversarial text-to-speech
Jeff Donahue, Sander Dieleman, Mikolaj Binkowski, Erich Elsen, and Karen Simonyan · 2021
Later among the works it cites.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Jaehyeon Kim, Jungil Kong, and Juhee Son · 2021
Later among the works it cites.
A multi-stage multi-codebook vq-vae approach to high-performance neural tts
Haohan Guo, Fenglong Xie, Frank K Soong, Xixin Wu, and Helen Meng · 2022
Later among the works it cites.
Minchan Kim, Myeonghun Jeong, Byoung Jin Choi, Sunghwan Ahn, Joun Yeop Lee, and Nam Soo Kim · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tomoki Hayashi and Shinji Watanabe · 2020
Cited alongside, same era.
HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Cited alongside, same era.
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck · 2020
Cited alongside, same era.
Glow-TTS: A Generative flow for text-to-speech via monotonic alignment search
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon · 2020
Cited alongside, same era.
Conversational end-to-end TTS for voice agents
Haohan Guo, Shaofei Zhang, Frank K Soong, Lei He, and Lei Xie · 2021
Cited alongside, same era.
Incremental speech synthesis for speech-to-speech translation
Danni Liu, Changhan Wang, Hongyu Gong, Xutai Ma, Yun Tang, and Juan Pino · 2021
Cited alongside, same era.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux · 2021
Cited alongside, same era.
Self-supervised speech representation learning: A review
Abdelrahman Mohamed, Hung-yi Lee, Lasse Borgholt, Jakob D Havtorn, Joakim Edin, Christian Igel, Katrin Kirchhoff, Shang-Wen Li, Karen Livescu, Lars Maaløe, et al · 2022
Later among the works it cites.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al · 2022
Later among the works it cites.
Self-supervised representation learning for speech processing
Hung-yi Lee, Abdelrahman Mohamed, Shinji Watanabe, Tara Sainath, Karen Livescu, Shang-Wen Li, Shu-wen Yang, and Katrin Kirchhoff · 2022
Later among the works it cites.
Large-scale self-supervised speech representation learning for automatic speaker verification
Zhengyang Chen, Sanyuan Chen, Yu Wu, Yao Qian, Chengyi Wang, Shujie Liu, Yanmin Qian, and Michael Zeng · 2022
Later among the works it cites.
Joint training of speech enhancement and self-supervised model for noise-robust asr
Qiu-Shi Zhu, Jie Zhang, Zi-Qiang Zhang, and Li-Rong Dai · 2022
Later among the works it cites.
Improving automatic speech recognition performance for low-resource languages with self-supervised models
Jing Zhao and Wei-Qiang Zhang · 2022
Later among the works it cites.
Vector quantized diffusion model for text-to-image synthesis
Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo · 2022
Later among the works it cites.
High fidelity neural audio compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi · 2022
Later among the works it cites.
Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition
Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, et al · 2022
Later among the works it cites.
PyCantonese: Cantonese linguistics and NLP in python
Jackson Lee, Litong Chen, Charles Lam, Chaak Ming Lau, and Tsz-Him Tsui · 2022
Later among the works it cites.
A survey on audio diffusion models: Text to speech synthesis and enhancement in generative ai
Chenshuang Zhang, Chaoning Zhang, Sheng Zheng, Mengchun Zhang, Maryam Qamar, Sung-Ho Bae, and In So Kweon · 2023
Closest in time.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
Eugene Kharitonov, Damien Vincent, Zalán Borsos, Raphaël Marinier, Sertan Girgin, Olivier Pietquin, Matt Sharifi, Marco Tagliasacchi, and Neil Zeghidour · 2023
Closest in time.
Superb@ slt 2022: Challenge on generalization and efficiency of self-supervised speech representation learning
Tzu-hsun Feng, Annie Dong, Ching-Feng Yeh, Shu-wen Yang, Tzu-Quan Lin, Jiatong Shi, Kai-Wei Chang, Zili Huang, Haibin Wu, Xuankai Chang, et al · 2023
Closest in time.
A comparative study of self-supervised speech representations in read and spontaneous tts
Siyang Wang, Gustav Eje Henter, Joakim Gustafson, and Éva Székely · 2023
Closest in time.
Neural codec language models are zero-shot text to speech synthesizers
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Closest in time.