Fetching the paper…
Reading the bibliography…
Modern text-to-speech synthesis pipelines typically involve multiple processing stages, each of which is designed or learnt independently from the rest.
WaveFlow: A compact flow-based model for raw audio
Wei Ping, Kainan Peng, Kexin Zhao, and Zhao Song · 1912
Earlier work this paper cites.
Dynamic-programming approach to continuous speech recognition
Hiroaki Sakoe · 1971
Earlier work this paper cites.
Minimum prediction residual principle applied to speech recognition
Fumitada Itakura · 1975
Earlier work this paper cites.
Dynamic programming algorithm optimization for spoken word recognition
Hiroaki Sakoe and Seibi Chiba · 1978
Earlier work this paper cites.
Signal estimation from modified short-time Fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Ronald J. Williams and David Zipser · 1989
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Text-to-speech synthesis
Paul Taylor · 2009
Earlier work this paper cites.
Statistical parametric speech synthesis
Heiga Zen, Keiichi Tokuda, and Alan W Black · 2009
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Fusion of magnitude and phase-based features for objective evaluation of TTS voice
Hardik B Sailor and Hemant A Patil · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dandelion Mané, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viégas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2015
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Earlier work this paper cites.
How to compare TTS systems: A new subjective evaluation methodology focused on differences
Jonathan Chevelu, Damien Lolive, Sébastien Le Maguer, and David Guennec · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
WORLD: A vocoder-based high-quality speech synthesis system for real-time applications
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa · 2016
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2016
Earlier work this paper cites.
WaveNet: A generative model for raw audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Deep Voice: Real-time neural text-to-speech
Sercan Ö Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, Shubho Sengupta, and Mohammad Shoeybi · 2017
Earlier work this paper cites.
Soft-DTW: a differentiable loss function for time-series
Marco Cuturi and Mathieu Blondel · 2017
Earlier work this paper cites.
Modulating early visual processing by language
Harm De Vries, Florian Strub, Jérémie Mary, Hugo Larochelle, Olivier Pietquin, and Aaron C Courville · 2017
Earlier work this paper cites.
A learned representation for artistic style
Vincent Dumoulin, Jonathon Shlens, and Manjunath Kudlur · 2017
Cited alongside, same era.
Deep Voice 2: Multi-speaker neural text-to-speech
Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou · 2017
Cited alongside, same era.
GANs trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Cited alongside, same era.
Jae Hyun Lim and Jong Chul Ye · 2017
Cited alongside, same era.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
SampleRNN: An unconditional end-to-end neural audio generation model
A new GAN-based end-to-end TTS training algorithm
Haohan Guo, Frank K Soong, Lei He, and Lei Xie · 2019
Later among the works it cites.
Robust sequence-to-sequence acoustic modeling with stepwise monotonic attention for neural TTS
Mutian He, Yan Deng, and Lei He · 2019
Later among the works it cites.
FloWaveNet: A generative flow for raw audio
Sungwon Kim, Sang-Gil Lee, Jongyoon Song, Jaehyeon Kim, and Sungroh Yoon · 2019
Later among the works it cites.
MelGAN: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brebisson, Yoshua Bengio, and Aaron Courville · 2019
Later among the works it cites.
A large-scale study on regularization and normalization in GANs
Karol Kurach, Mario Lučić, Xiaohua Zhai, Marcin Michalski, and Sylvain Gelly · 2019
Later among the works it cites.
Neural speech synthesis with transformer network
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2017
Cited alongside, same era.
Online and linear-time attention by enforcing monotonic alignments
Colin Raffel, Minh-Thang Luong, Peter J Liu, Ron J Weiss, and Douglas Eck · 2017
Cited alongside, same era.
Open sourcing Sonnet - a new library for constructing neural networks
Malcolm Reynolds, Gabriel Barth-Maron, Frederic Besse, Diego de Las Casas, Andreas Fidjeland, Tim Green, Adrià Puigdomènech, Sébastien Racanière, Jack Rae, and Fabio Viola · 2017
Cited alongside, same era.
Char2Wav: End-to-end speech synthesis
Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aaron Courville, and Yoshua Bengio · 2017
Cited alongside, same era.
VoiceLoop: Voice fitting and synthesis via a phonological loop
Yaniv Taigman, Lior Wolf, Adam Polyak, and Eliya Nachmani · 2017
Cited alongside, same era.
Hierarchical implicit models and likelihood-free variational inference
Dustin Tran, Rajesh Ranganath, and David M. Blei · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu · 2019
Later among the works it cites.
Expediting TTS synthesis with adversarial vocoding
Paarth Neekhara, Chris Donahue, Miller Puckette, Shlomo Dubnov, and Julian McAuley · 2019
Later among the works it cites.
Normalizing flows for probabilistic modeling and inference
George Papamakarios, Eric Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, and Balaji Lakshminarayanan · 2019
Later among the works it cites.
Parallel neural text-to-speech
Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao · 2019
Later among the works it cites.
WaveGlow: A flow-based generative network for speech synthesis
Ryan Prenger, Rafael Valle, and Bryan Catanzaro · 2019
Later among the works it cites.
FastSpeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Later among the works it cites.
LPCNet: Improving neural speech synthesis through linear prediction
Jean-Marc Valin and Jan Skoglund · 2019
Later among the works it cites.
MelNet: A generative model for audio in the frequency domain
Sean Vasquez and Mike Lewis · 2019
Later among the works it cites.
Neural source-filter-based waveform model for statistical parametric speech synthesis
Xin Wang, Shinji Takaki, and Junichi Yamagishi · 2019
Later among the works it cites.
Probability density distillation with generative adversarial networks for high-quality parallel waveform generation
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim · 2019
Later among the works it cites.
Neural models of text normalization for speech applications
Hao Zhang, Richard Sproat, Axel H Ng, Felix Stahlberg, Xiaochang Peng, Kyle Gorman, and Brian Roark · 2019
Later among the works it cites.
Location-relative attention mechanisms for robust long-form speech synthesis
Eric Battenberg, RJ Skerry-Ryan, Soroosh Mariooryad, Daisy Stanton, David Kao, Matt Shannon, and Tom Bagby · 2020
Closest in time.
Phonemizer
Mathieu Bernard · 2020
Closest in time.
High fidelity speech synthesis with adversarial networks
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C. Cobo, and Karen Simonyan · 2020
Closest in time.
DDSP: Differentiable digital signal processing
Jesse Engel, Lamtharn (Hanoi) Hantrakul, Chenjie Gu, and Adam Roberts · 2020
Closest in time.
Glow-TTS: A generative flow for text-to-speech via monotonic alignment search
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon · 2020
Closest in time.
Flow-TTS: A non-autoregressive network for text to speech based on flow
Chenfeng Miao, Shuang Liang, Minchuan Chen, Jun Ma, Shaojun Wang, and Jing Xiao · 2020
Closest in time.
Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis
Rafael Valle, Kevin Shih, Ryan Prenger, and Bryan Catanzaro · 2020
Closest in time.
Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim · 2020
Closest in time.
Multi-band MelGAN: Faster waveform generation for high-quality text-to-speech
Geng Yang, Shan Yang, Kai Liu, Peng Fang, Wei Chen, and Lei Xie · 2020
Closest in time.