Fetching the paper…
Reading the bibliography…
This work explores the task of synthesizing speech in nonexistent human-sounding voices.
“Effective use of variational embedding capacity in expressive end-to-end speech synthesis,” 2019,
Eric Battenberg, Soroosh Mariooryad, Daisy Stanton, R.J. Skerry-Ryan, Matt Shannon, David Kao, and Tom Bagby, · 1906
Earlier work this paper cites.
“A variational Bayesian framework for graphical models,”
Hagai Attias, · 1999
Earlier work this paper cites.
“YIN, a fundamental frequency estimator for speech and music,”
Alain de Cheveigné and Hideki Kawahara, · 2002
Earlier work this paper cites.
“Properties of f-divergences and f-GAN training,”
Matt Shannon, · 2009
Earlier work this paper cites.
“Attentron: Few-shot text-to-speech utilizing attention-based variable-length embedding,”
Seungwoo Choi, Seungju Han, Dongyoung Kim, and Sungjoo Ha, · 2011
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
Najim Dehak, Patrick Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet, · 2011
Earlier work this paper cites.
“Stochastic variational inference,”
Matthew D Hoffman, David M Blei, Chong Wang, and John Paisley, · 2013
Earlier work this paper cites.
“Generating sequences with recurrent neural networks,”
Alex Graves, · 2013
Earlier work this paper cites.
“Deep neural networks for small footprint text-dependent speaker verification,”
Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“WaveNet: A generative model for raw audio,” 2016,
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu, · 2016
Earlier work this paper cites.
“f-GAN: Training generative neural samplers using variational divergence minimization,”
Sebastian Nowozin, Botond Cseke, and Ryota Tomioka, · 2016
Earlier work this paper cites.
“Improved generator objectives for GANs,”
Ben Poole, Alexander A. Alemi, Jascha Sohl-Dickstein, and Anelia Angelova, · 2016
Earlier work this paper cites.
“Deep voice 2: Multi-speaker neural text-to-speech,”
Sercan Arık, Gregory Diamos, Andrew Gibiansky, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou, · 2017
Cited alongside, same era.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, R.J. Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous, · 2017
Cited alongside, same era.
“ β \beta -VAE: Learning basic visual concepts with a constrained variational framework,”
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner, · 2017
Cited alongside, same era.
“What does the speaker embedding encode?,”
Shuai Wang, Yanmin Qian, and Kai Yu, · 2017
Cited alongside, same era.
“Neural voice cloning with a few samples,”
Sercan Arık, Jitong Chen, Kainan Peng, Wei Ping, and Yanqi Zhou, · 2018
Cited alongside, same era.
“One-shot voice conversion by separating speaker and content representations with instance normalization,”
Ju-chieh Chou, Cheng-chieh Yeh, and Hung-yi Lee, · 2019
Later among the works it cites.
“Boffin TTS: Few-shot speaker adaptation by Bayesian optimization,”
Henry B Moss, Vatsal Aggarwal, Nishant Prateek, Javier González, and Roberto Barra-Chicote, · 2020
Later among the works it cites.
“Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings,”
Erica Cooper, Jeff Lai, Yusuke Yasuda, Fuming Fang, Xin Wang, Nanxin Chen, and Junichi Yamagishi, · 2020
Later among the works it cites.
“Semi-supervised generative modeling for controllable speech synthesis,”
Raza Habib, Soroosh Mariooryad, Matt Shannon, Eric Battenberg, R. J. Skerry-Ryan, Daisy Stanton, David Kao, and Tom Bagby, · 2020
Later among the works it cites.
“Location-relative attention mechanisms for robust long-form speech synthesis,”
Eric Battenberg, R.J. Skerry-Ryan, Soroosh Mariooryad, Daisy Stanton, David Kao, Matt Shannon, and Tom Bagby, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ye Jia, Yu Zhang, Ron J. Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez-Moreno, and Yonghui Wu, · 2018
Cited alongside, same era.
“X-Vectors: Robust DNN embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, R.J. Skerry-Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Yonghui Wu, · 2018
Cited alongside, same era.
Danilo Jimenez Rezende and Fabio Viola, · 2018
Cited alongside, same era.
“Efficient neural audio synthesis,”
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, et al., · 2018
Cited alongside, same era.
“Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron,”
R.J. Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron J. Weiss, Rob Clark, and Rif A. Saurous, · 2018
Cited alongside, same era.
“Style Tokens: Unsupervised style modeling, control, and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, R.J. Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Fei Ren, Ye Jia, and Rif A. Saurous, · 2018
Cited alongside, same era.
“Parallel WaveGAN: A Fast Waveform Generation Model Based on Generative Adversarial Networks with Multi-Resolution Spectrogram,”
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim, · 2020
Later among the works it cites.
“HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Later among the works it cites.
“Adaspeech: Adaptive text to speech for custom voice,”
Mingjian Chen, Xu Tan, Bohan Li, Yanqing Liu, Tao Qin, Sheng Zhao, and Tie-Yan Liu, · 2021
Closest in time.
“SC-GlowTTS: An efficient zero-shot multi-speaker text-to-speech model,”
Edresson Casanova, Christopher Shulby, Eren Gölge, Nicolas Michael Müller, Frederico Santos de Oliveira, Arnaldo Candido Junior, Anderson da Silva Soares, Sandra Maria Aluisio, and Moacir Antonelli Ponti, · 2021
Closest in time.
“Investigating on incorporating pretrained and learnable speaker representations for multi-speaker multi-style text-to-speech,”
Chung-Ming Chien, Jheng-Hao Lin, Chien-yu Huang, Po-chun Hsu, and Hung-yi Lee, · 2021
Closest in time.
“Deep Gaussian process based multi-speaker speech synthesis with latent speaker representation,”
Kentaro Mitsui, Tomoki Koriyama, and Hiroshi Saruwatari, · 2021
Closest in time.
“Translatotron 2: Robust direct speech-to-speech translation,” 2021
Ye Jia, Michelle Tadmor Ramanovich, Tal Remez, and Roi Pomerantz, · 2021
Closest in time.