Fetching the paper…
Reading the bibliography…
We describe a neural network-based system for text-to-speech (TTS) synthesis that is able to generate speech audio in the voice of many different speakers, including those unseen during training.
Suppression of acoustic noise in speech using spectral subtraction
Steven Boll · 1979
Earlier work this paper cites.
P. 800: Methods for subjective determination of transmission quality
ITUT Rec · 1996
Earlier work this paper cites.
Deep neural networks for small footprint text-dependent speaker verification
Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
LibriSpeech: an ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
End-to-end text-dependent speaker verification
Georg Heigold, Ignacio Moreno, Samy Bengio, and Noam Shazeer · 2016
Earlier work this paper cites.
WaveNet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Speaker adaptation in dnn-based speech synthesis using d-vectors
Rama Doddipatla, Norbert Braunschweiler, and Ranniery Maia · 2017
Earlier work this paper cites.
Deep Voice 2: Multi-speaker neural text-to-speech
Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou · 2017
Cited alongside, same era.
VoxCeleb: A large-scale speaker identification dataset
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman · 2017
Cited alongside, same era.
Char2Wav: End-to-end speech synthesis
Jose Sotelo, Soroush Mehri, Kundan Kumar, João Felipe Santos, Kyle Kastner, Aaron Courville, and Yoshua Bengio · 2017
Cited alongside, same era.
CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit, 2017
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous · 2017
Cited alongside, same era.
VoxCeleb2: Deep speaker recognition
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman · 2018
Closest in time.
Fitting new speakers based on a short untranscribed sample
Eliya Nachmani, Adam Polyak, Yaniv Taigman, and Lior Wolf · 2018
Closest in time.
Deep Voice 3: 2000-speaker neural text-to-speech
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller · 2018
Closest in time.
Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ Skerry-Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Yonghui. Wu · 2018
Closest in time.
Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron J. Weiss, Rob Clark, and Rif A. Saurous · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
https://ai.google/principles/ , 2018
Artificial Intelligence at Google – Our Principles · 2018
Cited alongside, same era.
Neural voice cloning with a few samples
Sercan O Arik, Jitong Chen, Kainan Peng, Wei Ping, and Yanqi Zhou · 2018
Cited alongside, same era.
Sample efficient adaptive text-to-speech
Yutian Chen, Yannis Assael, Brendan Shillingford, David Budden, Scott Reed, Heiga Zen, Quan Wang, Luis C Cobo, Andrew Trask, Ben Laurie, et al · 2018
Cited alongside, same era.
Closest in time.
VoiceLoop: Voice fitting and synthesis via a phonological loop
Yaniv Taigman, Lior Wolf, Adam Polyak, and Eliya Nachmani · 2018
Closest in time.
Generalized end-to-end loss for speaker verification
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno · 2018
Closest in time.
Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Fei Ren, Ye Jia, and Rif A Saurous · 2018
Closest in time.