Fetching the paper…
Reading the bibliography…
In this paper, we present StyleTTS 2, a text-to-speech (TTS) model that leverages style diffusion and adversarial training with large speech language models (SLMs) to achieve human-level TTS synthesis.
The lj speech dataset
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
Arbitrary style transfer in real-time with adaptive instance normalization
Xun Huang and Serge Belongie · 2017
Earlier work this paper cites.
Least squares generative adversarial networks
Xudong Mao, Qing Li, Haoran Xie, Raymond YK Lau, Zhen Wang, and Stephen Paul Smolley · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Fixing Weight Decay Regularization in Adam, 2018
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
ESPnet: End-to-end speech processing toolkit
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai · 2018
Earlier work this paper cites.
A review of deep learning based speech synthesis
Yishuang Ning, Sheng He, Zhiyong Wu, Chunxiao Xing, and Liang-Jie Zhang · 2019
Earlier work this paper cites.
CSTR VCTK Corpus: English Multi-Speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92)
Junichi Yamagishi, Christophe Veaux, Kirsten MacDonald, et al · 2019
Earlier work this paper cites.
LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
DiffWave: A Versatile Diffusion Model for Audio Synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2020
Earlier work this paper cites.
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Earlier work this paper cites.
End-to-End Adversarial Text-to-Speech
Jeff Donahue, Sander Dieleman, Mikołaj Bińkowski, Erich Elsen, and Karen Simonyan · 2020
Earlier work this paper cites.
Neural networks fail to learn periodic functions and how to fix it
Liu Ziyin, Tilman Hartwig, and Masahito Ueda · 2020
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2020
Earlier work this paper cites.
Jonathan Shen, Ye Jia, Mike Chrzanowski, Yu Zhang, Isaac Elias, Heiga Zen, and Yonghui Wu · 2020
Earlier work this paper cites.
A survey on neural speech synthesis
Xu Tan, Tao Qin, Frank Soong, and Tie-Yan Liu · 2021
Earlier work this paper cites.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Jaehyeon Kim, Jungil Kong, and Juhee Son · 2021
Earlier work this paper cites.
PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS
Ye Jia, Heiga Zen, Jonathan Shen, Yu Zhang, and Yonghui Wu · 2021
Earlier work this paper cites.
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Earlier work this paper cites.
Diff-TTS: A Denoising Diffusion Model for Text-to-Speech
Myeonghun Jeong, Hyeongju Kim, Sung Jun Cheon, Byoung Jin Choi, and Nam Soo Kim · 2021
Earlier work this paper cites.
Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov · 2021
Cited alongside, same era.
Wavegrad 2: Iterative refinement for text-to-speech synthesis
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, Najim Dehak, and William Chan · 2021
Cited alongside, same era.
Parallel Tacotron 2: A Non-Autoregressive Neural TTS Model with Differentiable Duration Modeling
Isaac Elias, Heiga Zen, Jonathan Shen, Yu Zhang, Jia Ye, R. J. Skerry-Ryan, and Yonghui Wu · 2021
Cited alongside, same era.
UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation
Won Jang, Daniel Chung Yong Lim, Jaesam Yoon, Bongwan Kim, and Juntae Kim · 2021
Cited alongside, same era.
A variational perspective on diffusion-based generative models and score matching
Chin-Wei Huang, Jae Hyun Lim, and Aaron C Courville · 2021
Cited alongside, same era.
HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech Synthesis
Sang-Hoon Lee, Seung-Bin Kim, Ji-Hyun Lee, Eunwoo Song, Min-Jae Hwang, and Seong-Whan Lee · 2022
Later among the works it cites.
Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech
Guangyan Zhang, Kaitao Song, Xu Tan, Daxin Tan, Yuzi Yan, Yanqing Liu, G. Wang, Wei Zhou, Tao Qin, Tan Lee, and Sheng Zhao · 2022
Later among the works it cites.
iSTFTNet: Fast and lightweight mel-spectrogram vocoder incorporating inverse short-time Fourier transform
Takuhiro Kaneko, Kou Tanaka, Hirokazu Kameoka, and Shogo Seki · 2022
Later among the works it cites.
Improve gan-based neural vocoder using truncated pointwise relativistic least square gan
Yanli Li and Congyi Wang · 2022
Later among the works it cites.
Elucidating the design space of diffusion-based generative models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Layer-wise analysis of a self-supervised speech representation model
Ankita Pasad, Ju-Chieh Chou, and Karen Livescu · 2021
Cited alongside, same era.
Projected GANs Converge Faster
Axel Sauer, Kashyap Chitta, Jens Müller, and Andreas Geiger · 2021
Cited alongside, same era.
Phonemizer: Text to Phones Transcription for Multiple Languages in Python
Mathieu Bernard and Hadrien Titeux · 2021
Cited alongside, same era.
StarGANv2-VC: A Diverse, Unsupervised, Non-parallel Framework for Natural-Sounding Voice Conversion
Yinghao Aaron Li, Ali Zare, and Nima Mesgarani · 2021
Cited alongside, same era.
NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
Xu Tan, Jiawei Chen, Haohe Liu, Jian Cong, Chen Zhang, Yanqing Liu, Xi Wang, Yichong Leng, Yuanhao Yi, Lei He, Frank K. Soong, Tao Qin, Sheng Zhao, and Tie-Yan Liu · 2022
Cited alongside, same era.
StyleTTS: A Style-Based Generative Model for Natural and Diverse Text-to-Speech Synthesis
Yinghao Aaron Li, Cong Han, and Nima Mesgarani · 2022
Cited alongside, same era.
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al · 2022
Cited alongside, same era.
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine · 2022
Later among the works it cites.
DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu · 2022
Later among the works it cites.
YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone
Edresson Casanova, Julian Weber, Christopher D Shulby, Arnaldo Candido Junior, Eren Gölge, and Moacir A Ponti · 2022
Later among the works it cites.
Phoneme-Level Bert for Enhanced Prosody of Text-To-Speech with Grapheme Predictions
Yinghao Aaron Li, Cong Han, Xilin Jiang, and Nima Mesgarani · 2023
Closest in time.
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Closest in time.
A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI
Chenshuang Zhang, Chaoning Zhang, Sheng Zheng, Mengchun Zhang, Maryam Qamar, Sung-Ho Bae, and In So Kweon · 2023
Closest in time.
Diffsound: Discrete Diffusion Model for Text-to-Sound Generation
Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu · 2023
Closest in time.
WaveFit: An Iterative and Non-autoregressive Neural Vocoder based on Fixed-Point Iteration
Yuma Koizumi, Kohei Yatabe, Heiga Zen, and Michiel Bacchiani · 2023
Closest in time.
NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian · 2023
Closest in time.
GAN You Hear Me? Reclaiming Unconditional Speech Synthesis from Diffusion Models
Matthew Baas and Herman Kamper · 2023
Closest in time.
M2-CTTS: End-to-End Multi-scale Multi-modal Conversational Text-to-Speech Synthesis
Jinlong Xue, Yayue Deng, Fengping Wang, Ya Li, Yingming Gao, Jianhua Tao, Jianqing Sun, and Jiaen Liang · 2023
Closest in time.
FoundationTTS: Text-to-Speech for ASR Custmization with Generative Language Model
Ruiqing Xue, Yanqing Liu, Lei He, Xu Tan, Linquan Liu, Edward Lin, and Sheng Zhao · 2023
Closest in time.
Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model
Kenichi Fujita, Takanori Ashihara, Hiroki Kanagawa, Takafumi Moriya, and Yusuke Ijima · 2023
Closest in time.
Why we should report the details in subjective evaluation of tts more rigorously
Cheng-Han Chiang, Wei-Ping Huang, and Hung-yi Lee · 2023
Closest in time.
Styletts-vc: One-shot voice conversion by knowledge transfer from style-based tts models
Yinghao Aaron Li, Cong Han, and Nima Mesgarani · 2023
Closest in time.
Summary of ChatGPT/GPT-4 Research and Perspective Towards the Future of Large Language Models
Yiheng Liu, Tianle Han, Siyuan Ma, Jiayue Zhang, Yuanyuan Yang, Jiaming Tian, Hao He, Antong Li, Mengshen He, Zhengliang Liu, et al · 2023
Closest in time.