Fetching the paper…
Reading the bibliography…
This paper proposes a neural sequence-to-sequence text-to-speech (TTS) model which can control latent attributes in the generated speech that are rarely annotated in the training data, such as speaking style, accent, background noise, and recording conditions.
Relationships among fundamental frequency, vocal sound pressure, and rate of speaking
John W Black · 1961
Earlier work this paper cites.
Effects of pitch and speech rate on personal attributions
William Apple, Lynn A Streeter, and Robert M Krauss · 1979
Earlier work this paper cites.
Suppression of acoustic noise in speech using spectral subtraction
Steven Boll · 1979
Earlier work this paper cites.
YIN, a fundamental frequency estimator for speech and music
Alain De Cheveigné and Hideki Kawahara · 2002
Earlier work this paper cites.
https://librivox.org , 2005
LibriVox · 2005
Earlier work this paper cites.
Robust signal-to-noise ratio estimation based on waveform amplitude distribution analysis
Chanwoo Kim and Richard M Stern · 2008
Earlier work this paper cites.
Statistical parametric speech synthesis
Heiga Zen, Keiichi Tokuda, and Alan W Black · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
The Blizzard Challenge 2013
Simon King and Vasilis Karaiskos · 2013
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
LibriSpeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Generating sentences from a continuous space
Samuel R Bowman, Luke Vilnis, Oriol Vinyals, Andrew Dai, Rafal Jozefowicz, and Samy Bengio · 2016
Cited alongside, same era.
Deep unsupervised clustering with Gaussian mixture variational autoencoders
Nat Dilokthanakul, Pedro AM Mediano, Marta Garnelo, Matthew CH Lee, Hugh Salimbeni, Kai Arulkumaran, and Murray Shanahan · 2016
Cited alongside, same era.
Approximate inference for deep latent Gaussian mixtures
Eric Nalisnick, Lars Hertel, and Padhraic Smyth · 2016
Cited alongside, same era.
WaveNet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Deep Voice: Real-time neural text-to-speech
Sercan Arık, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al · 2017
Neural voice cloning with a few samples
Sercan Arik, Jitong Chen, Kainan Peng, Wei Ping, and Yanqi Zhou · 2018
Closest in time.
Back-translation-style data augmentation for end-to-end ASR
Tomoki Hayashi, Shinji Watanabe, Yu Zhang, Tomoki Toda, Takaaki Hori, Ramon Astudillo, and Kazuya Takeda · 2018
Closest in time.
Deep encoder-decoder models for unsupervised learning of controllable speech synthesis
Gustav Eje Henter, Jaime Lorenzo-Trueba, Xin Wang, and Junichi Yamagishi · 2018
Closest in time.
Unsupervised adaptation with interpretable disentangled representations for distant conversational speech recognition
Wei-Ning Hsu, Hao Tang, and James Glass · 2018
Closest in time.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep Voice 2: Multi-speaker neural text-to-speech
Sercan Arik, Gregory Diamos, Andrew Gibiansky, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou · 2017
Cited alongside, same era.
beta-VAE: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2017
Cited alongside, same era.
Variational deep embedding: an unsupervised and generative approach to clustering
Zhuxi Jiang, Yin Zheng, Huachun Tan, Bangsheng Tang, and Hanning Zhou · 2017
Cited alongside, same era.
Generation of large-scale simulated utterances in virtual rooms to train deep-neural networks for far-field speech recognition in Google Home
Chanwoo Kim, Ananya Misra, Kean Chin, Thad Hughes, Arun Narayanan, Tara Sainath, and Michiel Bacchiani · 2017
Cited alongside, same era.
Char2Wav: End-to-End speech synthesis
J. Sotelo, S. Mehri, K. Kumar, J. Santos, K. Kastner, A. Courville, and Y. Bengio · 2017
Cited alongside, same era.
Listening while speaking: Speech chain by deep learning
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Ye Jia, Yu Zhang, Ron J Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, et al · 2018
Closest in time.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Closest in time.
Semi-amortized variational autoencoders
Yoon Kim, Sam Wiseman, Andrew C Miller, David Sontag, and Alexander M Rush · 2018
Closest in time.
Fitting new speakers based on a short untranscribed sample
Eliya Nachmani, Adam Polyak, Yaniv Taigman, and Lior Wolf · 2018
Closest in time.
Deep Voice 3: 2000-speaker neural text-to-speech
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller · 2018
Closest in time.
Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ Skerry-Ryan, et al · 2018
Closest in time.
Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron J Weiss, Rob Clark, and Rif A Saurous · 2018
Closest in time.
VoiceLoop: Voice fitting and synthesis via a phonological loop
Yaniv Taigman, Lior Wolf, Adam Polyak, and Eliya Nachmani · 2018
Closest in time.
Machine speech chain with one-shot speaker adaptation
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura · 2018
Closest in time.
Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Fei Ren, Ye Jia, and Rif A Saurous · 2018
Closest in time.