Fetching the paper…
Reading the bibliography…
Non-autoregressive text to speech (NAR-TTS) models have attracted much attention from both academia and industry due to their fast generation speed.
Libritts: A corpus derived from librispeech for text-to-speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. 2019 · 1904
Earlier work this paper cites.
Melnet: A generative model for audio in the frequency domain
Sean Vasquez and Mike Lewis. 2019 · 1906
Earlier work this paper cites.
High fidelity speech synthesis with adversarial networks
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C Cobo, and Karen Simonyan. 2019 · 1909
Earlier work this paper cites.
The theory behind controllable expressive speech synthesis: a cross-disciplinary approach
Noé Tits, Kevin El Haddad, and Thierry Dutoit. 2019 · 1910
Earlier work this paper cites.
The dip test of unimodality
John A Hartigan, Pamela M Hartigan, et al. 1985 · 1985
Earlier work this paper cites.
Density estimation for statistics and data analysis
Khosrow Dehnad. 1987 · 1987
Earlier work this paper cites.
Diatom autofocusing in brightfield microscopy: a comparative study
José Luis Pech-Pacheco, Gabriel Cristóbal, Jesús Chamorro-Martinez, and Joaquín Fernández-Valdivia. 2000 · 2000
Earlier work this paper cites.
Deep voice 3: 2000-speaker neural text-to-speech
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller. 2018 · 2000
Earlier work this paper cites.
Speech probability distribution
Saeed Gazor and Wei Zhang. 2003 · 2003
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004 · 2004
Earlier work this paper cites.
Glow-tts: A generative flow for text-to-speech via monotonic alignment search
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon. 2020 · 2005
Earlier work this paper cites.
Fastpitch: Parallel text-to-speech with pitch prediction
Adrian Łańcucki. 2020 · 2006
Earlier work this paper cites.
Fastspeech 2: Fast and high-quality end-to-end text-to-speech
Yi Ren, Chenxu Hu, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020 · 2006
Earlier work this paper cites.
Adversarially trained multi-singer sequence-to-sequence singing synthesizer
Jie Wu and Jian Luan. 2020 · 2006
Earlier work this paper cites.
Speedyspeech: Efficient neural speech synthesis
Jan Vainer and Ondřej Dušek. 2020 · 2008
Cited alongside, same era.
Speech quality assessment
Philipos C Loizou. 2011 · 2011
Cited alongside, same era.
Sang-Hoon Lee, Hyun-Wook Yoon, Hyeong-Rae Noh, Ji-Hoon Kim, and Seong-Whan Lee. 2020 · 2012
Cited alongside, same era.
Modeling spectral envelopes using restricted boltzmann machines and deep belief networks for statistical parametric speech synthesis
Zhen-Hua Ling, Li Deng, and Dong Yu. 2013 · 2013
Cited alongside, same era.
Root mean square error (rmse) or mean absolute error (mae)? –arguments against avoiding rmse in the literature
Tianfeng Chai and Roland R Draxler. 2014 · 2014
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Tacotron: Towards end-to-end speech synthesis
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al. 2017 · 2017
Later among the works it cites.
Glow: Generative flow with invertible 1x1 convolutions
Durk P Kingma and Prafulla Dhariwal. 2018 · 2018
Later among the works it cites.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al. 2018 · 2018
Later among the works it cites.
Probabilistic modeling of speech in spectral domain using maximum likelihood estimation
Mohammed Usman, Mohammed Zubair, Mohammad Shiblee, Paul Rodrigues, and Syed Jaffar. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun. 2015 · 2015
Cited alongside, same era.
Wavenet: A generative model for raw audio
Aäron Van Den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W Senior, and Koray Kavukcuoglu. 2016 · 2016
Cited alongside, same era.
Deep voice: Real-time neural text-to-speech
Sercan O Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al. 2017 · 2017
Cited alongside, same era.
Non-autoregressive neural machine translation
Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Image-to-image translation with conditional adversarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. 2017 · 2017
Cited alongside, same era.
Neural speech synthesis with transformer network
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. 2019 · 2019
Later among the works it cites.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2019 · 2019
Later among the works it cites.
Reducing over-smoothness in speech synthesis using generative adversarial networks
Leyuan Sheng and Evgeniy N Pavlovskiy. 2019 · 2019
Later among the works it cites.
Non-autoregressive machine translation with auxiliary regularization
Yiren Wang, Fei Tian, Di He, Tao Qin, ChengXiang Zhai, and Tie-Yan Liu. 2019 · 2019
Later among the works it cites.
Deep mixture density networks for acoustic modeling in statistical parametric speech synthesis
Heiga Zen and Andrew Senior. 2014 · 2019
Later among the works it cites.
Flow-tts: A non-autoregressive network for text to speech based on flow
Chenfeng Miao, Shuang Liang, Minchuan Chen, Jun Ma, Shaojun Wang, and Jing Xiao. 2020 · 2020
Later among the works it cites.
Non-autoregressive neural text-to-speech
Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao. 2020 · 2020
Later among the works it cites.
Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim. 2020 · 2020
Later among the works it cites.