Fetching the paper…
Reading the bibliography…
We propose Guided-TTS 2, a diffusion-based generative model for high-quality adaptive TTS using untranscribed data.
Reverse-time diffusion equation models
Brian D O Anderson · 1982
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
The lj speech dataset
Keith Ito · 2017
Earlier work this paper cites.
Montreal forced aligner: Trainable text-speech alignment using kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger · 2017
Earlier work this paper cites.
Tacotron: Towards End-to-End Speech Synthesis
Yuxuan Wang, R.J. Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous · 2017
Earlier work this paper cites.
Voxceleb2: Deep speaker recognition
J. S. Chung, A. Nagrani, and A. Zisserman · 2018
Earlier work this paper cites.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis
Ye Jia, Yu Zhang, Ron Weiss, Quan Wang, Jonathan Shen, Fei Ren, zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, and Yonghui Wu · 2018
Earlier work this paper cites.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Earlier work this paper cites.
Generalized end-to-end loss for speaker verification
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno · 2018
Earlier work this paper cites.
Sample efficient adaptive text-to-speech
Yutian Chen, Yannis Assael, Brendan Shillingford, David Budden, Scott Reed, Heiga Zen, Quan Wang, Luis C. Cobo, Andrew Trask, Ben Laurie, Caglar Gulcehre, Aäron van den Oord, Oriol Vinyals, and Nando de Freitas · 2019
Earlier work this paper cites.
Adversarial audio synthesis
C. Donahue, J. McAuley, and M. Puckette · 2019
Earlier work this paper cites.
Flowavenet: A generative flow for raw audio
Sungwon Kim, Sang-Gil Lee, Jongyoon Song, Jaehyeon Kim, and Sungroh Yoon · 2019
Earlier work this paper cites.
Nemo: a toolkit for building ai applications using neural modules
Oleksii Kuchaiev, Jason Li, Huyen Nguyen, Oleksii Hrinchuk, Ryan Leary, Boris Ginsburg, Samuel Kriman, Stanislav Beliaev, Vitaly Lavrukhin, Jack Cook, et al · 2019
Earlier work this paper cites.
Resemblyzer
Gilles Louppe · 2019
Cited alongside, same era.
Waveglow: A flow-based generative network for speech synthesis
Ryan Prenger, Rafael Valle, and Bryan Catanzaro · 2019
Cited alongside, same era.
FastSpeech: Fast, Robust and Controllable Text to Speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Cited alongside, same era.
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon · 2019
Cited alongside, same era.
Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92)
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald · 2019
Cited alongside, same era.
LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu · 2019
Cited alongside, same era.
Diffusion models beat GANs on image synthesis
Prafulla Dhariwal and Alexander Quinn Nichol · 2021
Later among the works it cites.
End-to-end Adversarial Text-to-Speech
Jeff Donahue, Sander Dieleman, Mikolaj Binkowski, Erich Elsen, and Karen Simonyan · 2021
Later among the works it cites.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2021
Later among the works it cites.
Diff-TTS: A Denoising Diffusion Model for Text-to-Speech
Myeonghun Jeong, Hyeongju Kim, Sung Jun Cheon, Byoung Jin Choi, and Nam Soo Kim · 2021
Later among the works it cites.
DiffWave: A Versatile Diffusion Model for Audio Synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2021
Later among the works it cites.
Meta-stylespeech : Multi-speaker adaptive text-to-speech generation
Dongchan Min, Dong Bok Lee, Eunho Yang, and Sung Ju Hwang · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings
Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang, Xin Wang, Nanxin Chen, and Junichi Yamagishi · 2020
Cited alongside, same era.
Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Libri-light: A benchmark for asr with limited or no supervision
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux · 2020
Cited alongside, same era.
Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon · 2020
Cited alongside, same era.
HiFi-GAN: Generative Adversarial networks for Efficient and High Fidelity Speech Synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Cited alongside, same era.
Improved techniques for training score-based generative models
Yang Song and Stefano Ermon · 2020
Cited alongside, same era.
GLIDE: towards photorealistic image generation and editing with text-guided diffusion models
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen · 2021
Later among the works it cites.
Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2021
Later among the works it cites.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2021
Later among the works it cites.
Adaspeech 2: Adaptive text to speech with untranscribed data
Yuzi Yan, Xu Tan, Bohan Li, Tao Qin, Sheng Zhao, Yuan Shen, and Tie-Yan Liu · 2021
Later among the works it cites.
Priorgrad: Improving conditional denoising diffusion models with data-dependent adaptive prior
Sang gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan, Chang Liu, Qi Meng, Tao Qin, Wei Chen, Sungroh Yoon, and Tie-Yan Liu · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents, 2022
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Closest in time.
Adaspeech 4: Adaptive text to speech in zero-shot scenarios
Yihan Wu, Xu Tan, Bohan Li, Lei He, Sheng Zhao, Ruihua Song, Tao Qin, and Tie-Yan Liu · 2022
Closest in time.