Fetching the paper…
Reading the bibliography…
Non-autoregressive text to speech (TTS) models such as FastSpeech can synthesize speech significantly faster than previous autoregressive models with comparable quality.
Decomposition of hardy functions into square integrable wavelets of constant shape
Alexander Grossmann and Jean Morlet · 1984
Earlier work this paper cites.
Wavelet transformations in signal detection
Franz B Tuteur · 1988
Earlier work this paper cites.
Ricker, ormsby; klander, bntterwo-a choice of wavelets, 1994
Harold Ryan · 1994
Earlier work this paper cites.
Objective measure for estimating mean opinion score of synthesized speech, April 4 2006
Min Chu and Hu Peng · 2006
Earlier work this paper cites.
Dynamic time warping
Meinard Müller · 2007
Earlier work this paper cites.
Speech quality assessment
Philipos C Loizou · 2011
Earlier work this paper cites.
One-to-many neural network mapping techniques for face image synthesis
Chrisina Jayne, Andreas Lanitis, and Chris Christodoulou · 2012
Earlier work this paper cites.
Wavelets for intonation modeling in hmm speech synthesis
Antti Santeri Suni, Daniel Aalto, Tuomo Raitio, Paavo Alku, Martti Vainio, et al · 2013
Earlier work this paper cites.
Statistical parametric speech synthesis using deep neural networks
Heiga Ze, Andrew Senior, and Mike Schuster · 2013
Earlier work this paper cites.
Differences of pitch profiles in germanic and slavic languages
Bistra Andreeva, Grażyna Demenko, Bernd Möbius, Frank Zimmerer, Jeanin Jügler, and Magdalena Oleskowicz-Popiel · 2014
Earlier work this paper cites.
Tts synthesis with bidirectional lstm based recurrent neural networks
Yuchen Fan, Yao Qian, Feng-Long Xie, and Frank K Soong · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Speech Prosody in Speech Synthesis: Modeling and generation of prosody for high quality and flexible speech synthesis
Keikichi Hirose and Jianhua Tao · 2015
Earlier work this paper cites.
Deep bidirectional lstm modeling of timbre and prosody for emotional voice conversion
Huaiping Ming, Dongyan Huang, Lei Xie, Jie Wu, Minghui Dong, and Haizhou Li · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aäron Van Den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al · 2016
Earlier work this paper cites.
Deep voice: Real-time neural text-to-speech
Sercan O Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al · 2017
Cited alongside, same era.
Deep voice 2: Multi-speaker neural text-to-speech
Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou · 2017
Cited alongside, same era.
The lj speech dataset
Keith Ito · 2017
Cited alongside, same era.
Montreal forced aligner: Trainable text-speech alignment using kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger · 2017
Cited alongside, same era.
Parallel wavenet: Fast high-fidelity speech synthesis
Aaron van den Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George van den Driessche, Edward Lockhart, Luis C Cobo, Florian Stimberg, et al · 2017
Cited alongside, same era.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2019
Later among the works it cites.
Token-level ensemble distillation for grapheme-to-phoneme conversion
Hao Sun, Xu Tan, Jun-Wei Gan, Hongzhi Liu, Sheng Zhao, Tao Qin, and Tie-Yan Liu · 2019
Later among the works it cites.
Toward multimodal image-to-image translation
Jun-Yan Zhu, Richard Zhang, Deepak Pathak, Trevor Darrell, Alexei A Efros, Oliver Wang, and Eli Shechtman · 2019
Later among the works it cites.
Multispeech: Multi-speaker text to speech with transformer
Mingjian Chen, Xu Tan, Yi Ren, Jin Xu, Hao Sun, Sheng Zhao, and Tao Qin · 2020
Closest in time.
End-to-end adversarial text-to-speech
Jeff Donahue, Sander Dieleman, Mikołaj Bińkowski, Erich Elsen, and Karen Simonyan · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al · 2017
Cited alongside, same era.
Flowavenet: A generative flow for raw audio
Sungwon Kim, Sang-gil Lee, Jongyoon Song, Jaehyeon Kim, and Sungroh Yoon · 2018
Cited alongside, same era.
Deep voice 3: 2000-speaker neural text-to-speech
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller · 2018
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Cited alongside, same era.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C Courville · 2019
Cited alongside, same era.
Neural speech synthesis with transformer network
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu · 2019
Cited alongside, same era.
Jesse Engel, Lamtharn Hantrakul, Chenjie Gu, and Adam Roberts · 2020
Closest in time.
Michael Gadermayr, Maximilian Tschuchnig, Dorit Merhof, Nils Krämer, Daniel Truhn, and Burkhard Gess · 2020
Closest in time.
Glow-tts: A generative flow for text-to-speech via monotonic alignment search
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon · 2020
Closest in time.
Fastpitch: Parallel text-to-speech with pitch prediction
Adrian Łańcucki · 2020
Closest in time.
Jdi-t: Jointly trained duration informed transformer for text-to-speech without explicit alignment
Dan Lim, Won Jang, Hyeyeong Park, Bongwan Kim, Jesam Yoon, et al · 2020
Closest in time.
Flow-tts: A non-autoregressive network for text to speech based on flow
Chenfeng Miao, Shuang Liang, Minchuan Chen, Jun Ma, Shaojun Wang, and Jing Xiao · 2020
Closest in time.
Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim · 2020
Closest in time.
Aligntts: Efficient feed-forward text-to-speech system without explicit alignment
Zhen Zeng, Jianzong Wang, Ning Cheng, Tian Xia, and Jing Xiao · 2020
Closest in time.
Adaspeech: Adaptive text to speech for custom voice
Mingjian Chen, Xu Tan, Bohan Li, Yanqing Liu, Tao Qin, sheng zhao, and Tie-Yan Liu · 2021
Closest in time.
Lightspeech: Lightweight and fast text to speech with neural architecture search
Renqian Luo, Xu Tan, Rui Wang, Tao Qin, Jinzhu Li, Sheng Zhao, Enhong Chen, and Tie-Yan Liu · 2021
Closest in time.
Denoispeech: Denoising text to speech with frame-level noise modeling
Chen Zhang, Yi Ren, Xu Tan, Jinglin Liu, Kejun Zhang, Tao Qin, Sheng Zhao, and Tie-Yan Liu · 2021
Closest in time.