Fetching the paper…
Reading the bibliography…
Text to speech (TTS), or speech synthesis, which aims to synthesize intelligible and natural speech given text, is a hot research topic in speech, language, and machine learning communities and has broad applications in the industry.
Efficient bidirectional neural machine translation
Xu Tan, Yingce Xia, Lijun Wu, and Tao Qin · 1908
Earlier work this paper cites.
Maximizing mutual information for tacotron
Peng Liu, Xixin Wu, Shiyin Kang, Guangzhi Li, Dan Su, and Dong Yu · 1909
Earlier work this paper cites.
Towards robust neural vocoding for speech generation: A survey
Po-chun Hsu, Chun-hsuan Wang, Andy T Liu, and Hung-yi Lee · 1912
Earlier work this paper cites.
A study of multilingual neural machine translation
Xu Tan, Yichong Leng, Jiale Chen, Yi Ren, Tao Qin, and Tie-Yan Liu · 1912
Earlier work this paper cites.
The speaking machine of wolfgang von kempelen
Homer Dudley and Thomas H Tarnoczy · 1950
Earlier work this paper cites.
Line spectrum representation of linear predictor coefficients of speech signals
Fumitada Itakura · 1975
Earlier work this paper cites.
A model of articulatory dynamics and control
Cecil H Coker · 1976
Earlier work this paper cites.
Automatic generation of control signals for a parallel formant speech synthesizer
P Seeviour, J Holmes, and M Judd · 1976
Earlier work this paper cites.
Rule synthesis of speech from dyadic units
Joseph Olive · 1977
Earlier work this paper cites.
Mitalk-79: The 1979 mit text-to-speech system
Jonathan Allen, Sharon Hunnicutt, Rolf Carlson, and Bjorn Granstrom · 1979
Earlier work this paper cites.
Software for a cascade/parallel formant synthesizer
Dennis H Klatt · 1980
Earlier work this paper cites.
Cepstral analysis synthesis on the mel frequency scale
Satoshi Imai · 1983
Earlier work this paper cites.
Mel log spectrum approximation (mlsa) filter for speech synthesis
Satoshi Imai, Kazuo Sumita, and Chieko Furuichi · 1983
Earlier work this paper cites.
Signal estimation from modified short-time fourier transform
Daniel Griffin and Jae Lim · 1984
Earlier work this paper cites.
An introduction to hidden markov models
Lawrence Rabiner and Biinghwang Juang · 1986
Earlier work this paper cites.
Review of text-to-speech conversion for english
Dennis H Klatt · 1987
Earlier work this paper cites.
Digital signal processing
William D Stanley, Gary R Dougherty, Ray Dougherty, and H Saunders · 1988
Earlier work this paper cites.
The evolutionary history of the human speech organs
Jan Wind · 1989
Earlier work this paper cites.
Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones
Eric Moulines and Francis Charpentier · 1990
Earlier work this paper cites.
Understanding human communication
Ronald Brian Adler, George R Rodman, and Alexandre Sévigny · 1991
Earlier work this paper cites.
An adaptive algorithm for mel-cepstral analysis of speech
Toshiaki Fukada, Keiichi Tokuda, Takao Kobayashi, and Satoshi Imai · 1992
Earlier work this paper cites.
Atr μ \mu -talk speech synthesis system
Yoshinori Sagisaka, Nobuyoshi Kaiki, Naoto Iwahashi, and Katsuhiko Mimura · 1992
Earlier work this paper cites.
Tobi: A standard for labeling english prosody
Kim Silverman, Mary Beckman, John Pitrelli, Mori Ostendorf, Colin Wightman, Patti Price, Janet Pierrehumbert, and Julia Hirschberg · 1992
Earlier work this paper cites.
Mel-generalized cepstral analysis-a unified approach to speech spectral estimation
Keiichi Tokuda, Takao Kobayashi, Takashi Masuko, and Satoshi Imai · 1994
Earlier work this paper cites.
Unit selection in a concatenative speech synthesis system using a large speech database
Andrew J Hunt and Alan W Black · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The aligner: Text-to-speech alignment using markov models
Colin W Wightman and David T Talkin · 1997
Earlier work this paper cites.
The festival speech synthesis system, 1998
Alan Black, Paul Taylor, Richard Caley, and Rob Clark · 1998
Earlier work this paper cites.
The tilt intonation model
Paul Taylor · 1998
Earlier work this paper cites.
Restructuring speech representations using a pitch-adaptive time–frequency smoothing and an instantaneous-frequency-based f0 extraction: Possible role of a repetitive structure in sounds
Hideki Kawahara, Ikuyo Masuda-Katsuse, and Alain De Cheveigne · 1999
Earlier work this paper cites.
Fundamentals of acoustics
Lawrence E Kinsler, Austin R Frey, Alan B Coppens, and James V Sanders · 1999
Earlier work this paper cites.
Foundations of statistical natural language processing
Christopher Manning and Hinrich Schutze · 1999
Earlier work this paper cites.
Simultaneous modeling of spectrum, pitch and duration in hmm-based speech synthesis
Takayoshi Yoshimura, Keiichi Tokuda, Takashi Masuko, Takao Kobayashi, and Tadashi Kitamura · 1999
Earlier work this paper cites.
Speech & language processing
Dan Jurafsky · 2000
Earlier work this paper cites.
Deep voice 3: 2000-speaker neural text-to-speech
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller · 2000
Earlier work this paper cites.
Speech parameter generation algorithms for hmm-based speech synthesis
Keiichi Tokuda, Takayoshi Yoshimura, Takashi Masuko, Takao Kobayashi, and Tadashi Kitamura · 2000
Earlier work this paper cites.
Locating boundaries for prosodic constituents in unrestricted mandarin texts
Min Chu and Yao Qian · 2001
Earlier work this paper cites.
Automatic analysis of prosody for multilingual speech corpora
Daniel Hirst · 2001
Earlier work this paper cites.
Aperiodicity extraction and control using mixed mode excitation and group delay manipulation for a high quality speech analysis, modification and synthesis system straight
Hideki Kawahara, Jo Estill, and Osamu Fujimura · 2001
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John Lafferty, Andrew McCallum, and Fernando CN Pereira · 2001
Earlier work this paper cites.
Prospects for articulatory synthesis: A position paper
Christine H Shadle and Robert I Damper · 2001
Earlier work this paper cites.
Normalization of non-standard words
Richard Sproat, Alan W Black, Stanley Chen, Shankar Kumar, Mari Ostendorf, and Christopher Richards · 2001
Earlier work this paper cites.
An rnn-based algorithm to detect prosodic phrase for chinese tts
Zhiwei Ying and Xiaohua Shi · 2001
Earlier work this paper cites.
Artificial intelligence: a modern approach
Stuart Russell and Peter Norvig · 2002
Earlier work this paper cites.
Simultaneous modeling of phonetic and prosodic parameters, and characteristic conversion for hmm-based text-to-speech systems
Takayoshi Yoshimura · 2002
Earlier work this paper cites.
An efficient way to learn rules for grapheme-to-phoneme conversion in chinese
Zi-Rong Zhang, Min Chu, and Eric Chang · 2002
Earlier work this paper cites.
Conditional and joint models for grapheme-to-phoneme conversion
Stanley F Chen · 2003
Earlier work this paper cites.
Speech probability distribution
Saeed Gazor and Wei Zhang · 2003
Earlier work this paper cites.
Cmu arctic databases for speech synthesis
John Kominek, Alan W Black, and Ver Ver · 2003
Earlier work this paper cites.
Chinese word segmentation as character tagging
Nianwen Xue · 2003
Earlier work this paper cites.
Application of neural networks for pos tagging and intonation control in speech synthesis for polish
Artur Janicki · 2004
Earlier work this paper cites.
Grapheme-to-phoneme conversion for chinese text-to-speech
Jun Xu, Guohong Fu, and Haizhou Li · 2004
Earlier work this paper cites.
What comprises a good talking-head video generation?: A survey and benchmark
Lele Chen, Guofeng Cui, Ziyi Kou, Haitian Zheng, and Chenliang Xu · 2005
Earlier work this paper cites.
Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis
Rafael Valle, Kevin Shih, Ryan Prenger, and Bryan Catanzaro · 2005
Earlier work this paper cites.
Multi-band melgan: Faster waveform generation for high-quality text-to-speech
Geng Yang, Shan Yang, Kai Liu, Peng Fang, Wei Chen, and Lei Xie · 2005
Earlier work this paper cites.
Adadurian: Few-shot adaptation for neural text-to-speech with durian
Zewang Zhang, Qiao Tian, Heng Lu, Ling-Hui Chen, and Shan Liu · 2005
Earlier work this paper cites.
Pattern recognition and machine learning
Christopher M Bishop · 2006
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Straight, exploitation of the other aspect of vocoder: Perceptually isomorphic decomposition of speech sounds
Hideki Kawahara · 2006
Earlier work this paper cites.
Hkust/mts: A very large scale mandarin telephone speech corpus
Yi Liu, Pascale Fung, Yongsheng Yang, Christopher Cieri, Shudong Huang, and David Graff · 2006
Earlier work this paper cites.
Statistical parametric speech synthesis
Alan W Black, Heiga Zen, and Keiichi Tokuda · 2007
Earlier work this paper cites.
Inequality maximum entropy classifier with character features for polyphone disambiguation in mandarin tts systems
Xinnian Mao, Yuan Dong, Jinyu Han, Dezhi Huang, and Haila Wang · 2007
Earlier work this paper cites.
Exploiting acoustic and syntactic features for prosody labeling in a maximum entropy framework
Vivek Kumar Rangarajan Sridhar, Srinivas Bangalore, and Shrikanth Narayanan · 2007
Earlier work this paper cites.
A speech parameter generation algorithm considering global variance for hmm-based speech synthesis
Tomoki Toda and Keiichi Tokuda · 2007
Earlier work this paper cites.
Joint-sequence models for grapheme-to-phoneme conversion
Maximilian Bisani and Hermann Ney · 2008
Earlier work this paper cites.
Intonational phonology
D Robert Ladd · 2008
Earlier work this paper cites.
Automatic prosodic labeling with conditional random fields and rich acoustic features
Gina-Anne Levow · 2008
Earlier work this paper cites.
Expressive tts training with frame and style reconstruction loss
Rui Liu, Berrak Sisman, Guanglai Gao, and Haizhou Li · 2008
Earlier work this paper cites.
Hifisinger: Towards high-fidelity neural singing voice synthesis
Jiawei Chen, Xu Tan, Jian Luan, Tao Qin, and Tie-Yan Liu · 2009
Earlier work this paper cites.
Automatic prosodic events detection using syllable-based acoustic and syntactic features
Je Hun Jeon and Yang Liu · 2009
Earlier work this paper cites.
Chinese prosody structure prediction based on conditional random fields
Jingwei Sun, Jing Yang, Jianping Zhang, and Yonghong Yan · 2009
Earlier work this paper cites.
Text-to-speech synthesis
Paul Taylor · 2009
Earlier work this paper cites.
Statistical parametric speech synthesis
Heiga Zen, Keiichi Tokuda, and Alan W Black · 2009
Earlier work this paper cites.
Tts-by-tts: Tts-driven data augmentation for fast and high-quality speech synthesis
Min-Jae Hwang, Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim · 2010
Earlier work this paper cites.
Graphspeech: Syntax-aware graph attention network for neural speech synthesis
Rui Liu, Berrak Sisman, and Haizhou Li · 2010
Earlier work this paper cites.
Automatic prosody prediction and detection with conditional random field (crf) models
Yao Qian, Zhizheng Wu, Xuezhe Ma, and Frank Soong · 2010
Earlier work this paper cites.
Autobi-a tool for automatic tobi annotation
Andrew Rosenberg · 2010
Earlier work this paper cites.
The effects of part–of–speech tagging on text–to–speech synthesis for resource–scarce languages
Georg Isaac Schlünz · 2010
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2010
Earlier work this paper cites.
Experimental and theoretical advances in prosody: A review
Michael Wagner and Duane G Watson · 2010
Earlier work this paper cites.
Erica Cooper, Xin Wang, Yi Zhao, Yusuke Yasuda, and Junichi Yamagishi · 2011
Earlier work this paper cites.
Course in general linguistics
Ferdinand De Saussure · 2011
Earlier work this paper cites.
The blizzard challenge 2011
S. King and V. Karaiskos · 2011
Earlier work this paper cites.
Morphological analysis based part-of-speech tagging for uyghur speech synthesis
Guljamal Mamateli, Askar Rozi, Gulnar Ali, and Askar Hamdulla · 2011
Earlier work this paper cites.
Improved pos tagging for text-to-speech synthesis
Ming Sun and Jerome R Bellegarda · 2011
Earlier work this paper cites.
Speech synthesis techniques. a survey
Youcef Tabet and Mohamed Boughazi · 2011
Earlier work this paper cites.
Feathertts: Robust and efficient attention based neural tts
Qiao Tian, Zewang Zhang, Chao Liu, Heng Lu, Linghui Chen, Bin Wei, Pujiang He, and Shan Liu · 2011
Earlier work this paper cites.
Improving prosody modelling with cross-utterance bert embeddings for end-to-end speech synthesis
Guanghui Xu, Wei Song, Zhengchen Zhang, Chao Zhang, Xiaodong He, and Bowen Zhou · 2011
Earlier work this paper cites.
One-to-many neural network mapping techniques for face image synthesis
Chrisina Jayne, Andreas Lanitis, and Chris Christodoulou · 2012
Earlier work this paper cites.
Sang-Hoon Lee, Hyun-Wook Yoon, Hyeong-Rae Noh, Ji-Hoon Kim, and Seong-Whan Lee · 2012
Earlier work this paper cites.
Efficienttts: An efficient and high-quality text-to-speech architecture
Chenfeng Miao, Shuang Liang, Zhencheng Liu, Minchuan Chen, Jun Ma, Shaojun Wang, and Jing Xiao · 2012
Earlier work this paper cites.
Ted-lium: an automatic speech recognition dedicated corpus
Anthony Rousseau, Paul Deléglise, and Yannick Esteve · 2012
Earlier work this paper cites.
Unified mandarin tts front-end based on distilled bert model
Yang Zhang, Liqun Deng, and Yasheng Wang · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
The blizzard challenge 2013
S. King and V. Karaiskos · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Speech synthesis based on hidden markov models
Keiichi Tokuda, Yoshihiko Nankaku, Tomoki Toda, Heiga Zen, Junichi Yamagishi, and Keiichiro Oura · 2013
Earlier work this paper cites.
Statistical parametric speech synthesis using deep neural networks
Heiga Zen, Andrew Senior, and Mike Schuster · 2013
Earlier work this paper cites.
Deep learning for chinese word segmentation and pos tagging
Xiaoqing Zheng, Hanyang Chen, and Tianyu Xu · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Nice: Non-linear independent components estimation
Laurent Dinh, David Krueger, and Yoshua Bengio · 2014
Earlier work this paper cites.
Tts synthesis with bidirectional lstm based recurrent neural networks
Yuchen Fan, Yao Qian, Feng-Long Xie, and Frank K Soong · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
A survey on text to speech translation of multi language
Preranasri Mali · 2014
Earlier work this paper cites.
Slam: Automatic stylization and labelling of speech melody
Nicolas Obin, Julie Beliao, Christophe Veaux, and Anne Lacheret · 2014
Earlier work this paper cites.
Max-margin tensor neural network for chinese word segmentation
Wenzhe Pei, Tao Ge, and Baobao Chang · 2014
Earlier work this paper cites.
Tts tutorial at iscslp 2014
Yao Qian and Frank K Soong · 2014
Earlier work this paper cites.
On the training aspects of deep neural network (dnn) for parametric tts synthesis
Yao Qian, Yuchen Fan, Wenping Hu, and Frank K Soong · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Earlier work this paper cites.
Attention-based models for speech recognition
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Automatic prosody prediction for chinese speech synthesis using blstm-rnn and embedding features
Chuang Ding, Lei Xie, Jie Yan, Weini Zhang, and Yang Liu · 2015
Earlier work this paper cites.
Multi-speaker modeling and speaker adaptation for dnn-based tts synthesis
Yuchen Fan, Yao Qian, Frank K Soong, and Lei He · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Machine learning: Trends, perspectives, and prospects
Michael I Jordan and Tom M Mitchell · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2015
Earlier work this paper cites.
Grapheme-to-phoneme conversion using long short-term memory recurrent neural networks
Kanishka Rao, Fuchun Peng, Haşim Sak, and Françoise Beaufays · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
Danilo Rezende and Shakir Mohamed · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Word embedding for recurrent neural network based tts synthesis
Peilu Wang, Yao Qian, Frank K Soong, Lei He, and Hai Zhao · 2015
Earlier work this paper cites.
A study of speaker adaptation for dnn-based speech synthesis
Zhizheng Wu, Pawel Swietojanski, Christophe Veaux, Steve Renals, and Simon King · 2015
Earlier work this paper cites.
Sequence-to-sequence neural net models for grapheme-to-phoneme conversion
Kaisheng Yao and Geoffrey Zweig · 2015
Earlier work this paper cites.
Acoustic modeling in statistical parametric speech synthesis-from hmm to lstm-rnn
Heiga Zen · 2015
Earlier work this paper cites.
Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis
Heiga Zen and Haşim Sak · 2015
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals · 2016
Earlier work this paper cites.
Density estimation using real nvp
Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio · 2016
Earlier work this paper cites.
Speaker and language factorization in dnn-based tts synthesis
Yuchen Fan, Yao Qian, Frank K Soong, and Lei He · 2016
Earlier work this paper cites.
Deep learning , volume 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Professor forcing: a new algorithm for training recurrent networks
Anirudh Goyal, Alex Lamb, Ying Zhang, Saizheng Zhang, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush · 2016
Earlier work this paper cites.
Improved variational inference with inverse autoregressive flow
Durk P Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling · 2016
Earlier work this paper cites.
Autoencoding beyond pixels using a learned similarity metric
Anders Boesen Lindbo Larsen, Søren Kaae Sønderby, Hugo Larochelle, and Ole Winther · 2016
Earlier work this paper cites.
Deep learning for statistical parametric speech synthesis
Zhen-Hua Ling · 2016
Earlier work this paper cites.
World: a vocoder-based high-quality speech synthesis system for real-time applications
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Fast wavenet generation algorithm
Tom Le Paine, Pooya Khorrami, Shiyu Chang, Yang Zhang, Prajit Ramachandran, Mark A Hasegawa-Johnson, and Thomas S Huang · 2016
Earlier work this paper cites.
A bi-directional lstm approach for polyphone disambiguation in mandarin chinese
Changhao Shan, Lei Xie, and Kaisheng Yao · 2016
Earlier work this paper cites.
Rnn approaches to text normalization: A challenge
Richard Sproat and Navdeep Jaitly · 2016
Earlier work this paper cites.
Postfilters to modify the modulation spectrum for statistical parametric speech synthesis
Shinnosuke Takamichi, Tomoki Toda, Alan W Black, Graham Neubig, Sakriani Sakti, and Satoshi Nakamura · 2016
Earlier work this paper cites.
Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al · 2016
Earlier work this paper cites.
First step towards end-to-end parametric tts synthesis: Generating spectral parameters with neural attention
Wenfu Wang, Shuang Xu, and Bo Xu · 2016
Cited alongside, same era.
Fast, compact, and high quality lstm-rnn based statistical parametric speech synthesizers for mobile devices
Heiga Zen, Yannis Agiomyrgiannakis, Niels Egberts, Fergus Henderson, and Przemysław Szczepaniak · 2016
Cited alongside, same era.
Mandarin prosodic phrase prediction based on syntactic trees
Zhengchen Zhang, Fuxiang Wu, Chenyu Yang, Minghui Dong, and Fugen Zhou · 2016
Cited alongside, same era.
Speaker representations for speaker adaptation in multiple speakers blstm-rnn-based speech synthesis
Yi Zhao, Daisuke Saito, and Nobuaki Minematsu · 2016
Cited alongside, same era.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Cited alongside, same era.
Noise robust tts for low resource speakers using pre-trained model and speech enhancement
Dongyang Dai, Li Chen, Yuping Wang, Mu Wang, Rui Xia, Xuchen Song, Zhiyong Wu, and Yuxuan Wang · 2020
Later among the works it cites.
Efficient neural speech synthesis for low-resource languages through multilingual modeling
Marcel de Korte, Jaebok Kim, and Esther Klabbers · 2020
Later among the works it cites.
Parallel tacotron: Non-autoregressive and controllable tts
Isaac Elias, Heiga Zen, Jonathan Shen, Yu Zhang, Ye Jia, Ron Weiss, and Yonghui Wu · 2020
Later among the works it cites.
High quality streaming speech synthesis with low, sentence-length-independent latency
Nikolaos Ellinas, Georgios Vamvoukakis, Konstantinos Markopoulos, Aimilios Chalamandaris, Georgia Maniati, Panos Kakoulidis, Spyros Raptis, June Sig Sung, Hyoungmin Park, and Pirros Tsiakoulis · 2020
Later among the works it cites.
Interactive text-to-speech via semi-supervised style transfer learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep voice: Real-time neural text-to-speech
Sercan Ö Arık, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al · 2017
Cited alongside, same era.
Chinese standard mandarin speech corpus
Data Baker · 2017
Cited alongside, same era.
Deep generative models for speech and images
Yoshua Bengio · 2017
Cited alongside, same era.
Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng · 2017
Cited alongside, same era.
Speaker adaptation in dnn-based speech synthesis using d-vectors
Rama Doddipatla, Norbert Braunschweiler, and Ranniery Maia · 2017
Cited alongside, same era.
Deep voice 2: Multi-speaker neural text-to-speech
Andrew Gibiansky, Sercan Ömer Arik, Gregory Frederick Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou · 2017
Cited alongside, same era.
Improved training of wasserstein gans
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron Courville · 2017
Cited alongside, same era.
Yang Gao, Weiyi Zheng, Zhaojun Yang, Thilo Kohler, Christian Fuegen, and Qing He · 2020
Later among the works it cites.
A spectral energy distance for parallel speech synthesis
Alexey Gritsenko, Tim Salimans, Rianne van den Berg, Jasper Snoek, and Nal Kalchbrenner · 2020
Later among the works it cites.
Didispeech: A large scale mandarin speech corpus
Tingwei Guo, Cheng Wen, Dongwei Jiang, Ne Luo, Ruixiong Zhang, Shuaijiang Zhao, Wubo Li, Cheng Gong, Wei Zou, Kun Han, et al · 2020
Later among the works it cites.
Espnet-tts: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit
Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue, Takenori Yoshimura, Shinji Watanabe, Tomoki Toda, Kazuya Takeda, Yu Zhang, and Xu Tan · 2020
Later among the works it cites.
Open-source multi-speaker speech corpora for building gujarati, kannada, malayalam, marathi, tamil and telugu speech synthesis systems
Fei He, Shan-Hui Cathy Chu, Oddur Kjartansson, Clara Rivera, Anna Katanova, Alexander Gutkin, Isin Demirsahin, Cibu Johny, Martin Jansche, Supheakmungkol Sarin, et al · 2020
Later among the works it cites.
Hamed Hemati and Damian Borth · 2020
Later among the works it cites.
Speaker adaptation of a multilingual acoustic model for cross-language synthesis
Ivan Himawan, Sandesh Aryal, Iris Ouyang, Sam Kang, Pierre Lanchantin, and Simon King · 2020
Later among the works it cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Later among the works it cites.
Hierarchical multi-grained generative model for expressive speech synthesis
Yukiya Hono, Kazuna Tsuboi, Kei Sawada, Kei Hashimoto, Keiichiro Oura, Yoshihiko Nankaku, and Keiichi Tokuda · 2020
Later among the works it cites.
Wg-wavenet: Real-time high-fidelity speech synthesis without gpu
Po-chun Hsu and Hung-yi Lee · 2020
Later among the works it cites.
Devicetts: A small-footprint, fast, stable network for on-device text-to-speech
Zhiying Huang, Hao Li, and Ming Lei · 2020
Later among the works it cites.
Low-resource expressive text-to-speech using data augmentation
Goeric Huybrechts, Thomas Merritt, Giulia Comini, Bartek Perz, Raahil Shah, and Jaime Lorenzo-Trueba · 2020
Later among the works it cites.
Improving lpcnet-based text-to-speech with linear prediction-structured mixture density network
Min-Jae Hwang, Eunwoo Song, Ryuichi Yamamoto, Frank Soong, and Hong-Goo Kang · 2020
Later among the works it cites.
Semi-supervised speaker adaptation for end-to-end speech synthesis with pretrained models
Katsuki Inoue, Sunao Hara, Masanobu Abe, Tomoki Hayashi, Ryuichi Yamamoto, and Shinji Watanabe · 2020
Later among the works it cites.
Universal melgan: A robust neural vocoder for high-fidelity waveform generation in multiple domains
Won Jang, Dan Lim, and Jaesam Yoon · 2020
Later among the works it cites.
Lightweight lpcnet-based neural vocoder with tensor decomposition
Hiroki Kanagawa and Yusuke Ijima · 2020
Later among the works it cites.
Copycat: Many-to-many fine-grained prosody transfer for neural text-to-speech
Sri Karlapati, Alexis Moinet, Arnaud Joly, Viacheslav Klimkov, Daniel Sáez-Trigueros, and Thomas Drugman · 2020
Later among the works it cites.
Glow-tts: A generative flow for text-to-speech via monotonic alignment search
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon · 2020
Later among the works it cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae · 2020
Later among the works it cites.
Fastpitch: Parallel text-to-speech with pitch prediction
Adrian Łańcucki · 2020
Later among the works it cites.
Moboaligner: A neural alignment model for non-autoregressive tts with monotonic boundary search
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, Ming Liu, and Ming Zhou · 2020
Later among the works it cites.
Jdi-t: Jointly trained duration informed transformer for text-to-speech without explicit alignment
Dan Lim, Won Jang, O Gyeonghwan, Heayoung Park, Bongwan Kim, and Jaesam Yoon · 2020
Later among the works it cites.
Towards unsupervised speech recognition and synthesis with quantized speech representation learning
Alexander H Liu, Tao Tu, Hung-yi Lee, and Lin-shan Lee · 2020
Later among the works it cites.
Teacher-student training for robust tacotron-based tts
Rui Liu, Berrak Sisman, Jingdong Li, Feilong Bao, Guanglai Gao, and Haizhou Li · 2020
Later among the works it cites.
Neural homomorphic vocoder
Zhijun Liu, Kuan Chen, and Kai Yu · 2020
Later among the works it cites.
Xiaoicesing: A high-quality and integrated singing voice synthesis system
Peiling Lu, Jie Wu, Jian Luan, Xu Tan, and Li Zhou · 2020
Later among the works it cites.
Neural architecture search with gbdt
Renqian Luo, Xu Tan, Rui Wang, Tao Qin, Enhong Chen, and Tie-Yan Liu · 2020
Later among the works it cites.
Nautilus: a versatile voice cloning system
Hieu-Thi Luong and Junichi Yamagishi · 2020
Later among the works it cites.
Incremental text-to-speech synthesis with prefix-to-prefix framework
Mingbo Ma, Baigong Zheng, Kaibo Liu, Renjie Zheng, Hairong Liu, Kainan Peng, Kenneth Church, and Liang Huang · 2020
Later among the works it cites.
Generating multilingual voices using speaker space translation based on bilingual speaker data
Soumi Maiti, Erik Marchi, and Alistair Conkie · 2020
Later among the works it cites.
Flow-tts: A non-autoregressive network for text to speech based on flow
Chenfeng Miao, Shuang Liang, Minchuan Chen, Jun Ma, Shaojun Wang, and Jing Xiao · 2020
Later among the works it cites.
Incremental text to speech for neural sequence-to-sequence models using reinforcement learning
Devang S Ram Mohan, Raphael Lenain, Lorenzo Foglianti, Tian Huey Teh, Marlene Staib, Alexandra Torresquintero, and Jiameng Gao · 2020
Later among the works it cites.
Controllable neural prosody synthesis
Max Morrison, Zeyu Jin, Justin Salamon, Nicholas J Bryan, and Gautham J Mysore · 2020
Later among the works it cites.
Boffin tts: Few-shot speaker adaptation by bayesian optimization
Henry B Moss, Vatsal Aggarwal, Nishant Prateek, Javier González, and Roberto Barra-Chicote · 2020
Later among the works it cites.
One model, many languages: Meta-learning for multilingual text-to-speech
Tomáš Nekvinda and Ondřej Dušek · 2020
Later among the works it cites.
A unified sequence-to-sequence front-end model for mandarin text-to-speech synthesis
Junjie Pan, Xiang Yin, Zhiling Zhang, Shichao Liu, Yang Zhang, Zejun Ma, and Yuxuan Wang · 2020
Later among the works it cites.
A survey on speech synthesis techniques in indian languages
Soumya Priyadarsini Panda, Ajit Kumar Nayak, and Satyananda Champati Rai · 2020
Later among the works it cites.
g2pm: A neural grapheme-to-phoneme conversion package for mandarin chinese based on a new open benchmark dataset
Kyubyong Park and Seanie Lee · 2020
Later among the works it cites.
Speaker conditional wavernn: Towards universal neural vocoder for unseen speaker and recording conditions
Dipjyoti Paul, Yannis Pantazis, and Yannis Stylianou · 2020
Later among the works it cites.
Enhancing speech intelligibility in text-to-speech synthesis using speaking style conversion
Dipjyoti Paul, Muhammed PV Shifas, Yannis Pantazis, and Yannis Stylianou · 2020
Later among the works it cites.
Non-autoregressive neural text-to-speech
Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao · 2020
Later among the works it cites.
Waveflow: A compact flow-based model for raw audio
Wei Ping, Kainan Peng, Kexin Zhao, and Zhao Song · 2020
Later among the works it cites.
Fast and lightweight on-device tts with tacotron2 and lpcnet
Vadim Popov, Stanislav Kamenev, Mikhail Kudinov, Sergey Repyevsky, Tasnima Sadekova, Vladimir Kryzhanovskiy Bushaev, and Denis Parkhomenko · 2020
Later among the works it cites.
Gaussian lpcnet for multisample speech synthesis
Vadim Popov, Mikhail Kudinov, and Tasnima Sadekova · 2020
Later among the works it cites.
Mls: A large-scale multilingual dataset for speech research
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert · 2020
Later among the works it cites.
Unsupervised speech decomposition via triple information bottleneck
Kaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson, and David Cox · 2020
Later among the works it cites.
Dual Learning
Tao Qin · 2020
Later among the works it cites.
Jonathan Shen, Ye Jia, Mike Chrzanowski, Yu Zhang, Isaac Elias, Heiga Zen, and Yonghui Wu · 2020
Later among the works it cites.
Aishell-3: A multi-speaker mandarin tts corpus and the baselines
Yao Shi, Hui Bu, Xin Xu, Shaoji Zhang, and Ming Li · 2020
Later among the works it cites.
An overview of voice conversion and its challenges: From statistical modeling to deep learning
Berrak Sisman, Junichi Yamagishi, Simon King, and Haizhou Li · 2020
Later among the works it cites.
Neural text-to-speech with a modeling-by-generation excitation vocoder
Eunwoo Song, Min-Jae Hwang, Ryuichi Yamamoto, Jin-Seob Kim, Ohsung Kwon, and Jae-Min Kim · 2020
Later among the works it cites.
Phonological features for 0-shot multilingual speech synthesis
Marlene Staib, Tian Huey Teh, Alexandra Torresquintero, Devang S Ram Mohan, Lorenzo Foglianti, Raphael Lenain, and Jiameng Gao · 2020
Later among the works it cites.
What the future brings: Investigating the impact of lookahead for incremental neural tts
Brooke Stephenson, Laurent Besacier, Laurent Girin, and Thomas Hueber · 2020
Later among the works it cites.
Generating diverse and natural text-to-speech samples using a quantized fine-grained vae and autoregressive prosody prior
Guangzhi Sun, Yu Zhang, Ron J Weiss, Yuan Cao, Heiga Zen, Andrew Rosenberg, Bhuvana Ramabhadran, and Yonghui Wu · 2020
Later among the works it cites.
Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis
Guangzhi Sun, Yu Zhang, Ron J Weiss, Yuan Cao, Heiga Zen, and Yonghui Wu · 2020
Later among the works it cites.
Featherwave: An efficient high-fidelity neural vocoder with multi-band linear prediction
Qiao Tian, Zewang Zhang, Heng Lu, Ling-Hui Chen, and Shan Liu · 2020
Later among the works it cites.
Semi-supervised learning for multi-speaker text-to-speech synthesis using discrete speech representation
Tao Tu, Yuan-Jui Chen, Alexander H Liu, and Hung-yi Lee · 2020
Later among the works it cites.
Emotional speech synthesis with rich and granularized control
Se-Yun Um, Sangshin Oh, Kyungguen Byun, Inseon Jang, ChungHyun Ahn, and Hong-Goo Kang · 2020
Later among the works it cites.
Speedyspeech: Efficient neural speech synthesis
Jan Vainer and Ondřej Dušek · 2020
Later among the works it cites.
Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens
Rafael Valle, Jason Li, Ryan Prenger, and Bryan Catanzaro · 2020
Later among the works it cites.
Bunched lpcnet: Vocoder for low-cost neural text-to-speech systems
Ravichander Vipperla, Sangjun Park, Kihyun Choo, Samin Ishtiaq, Kyoungbo Min, Sourav Bhattacharya, Abhinav Mehrotra, Alberto Gil CP Ramos, and Nicholas D Lane · 2020
Later among the works it cites.
s-transformer: Segment-transformer for robust neural speech synthesis
Xi Wang, Huaiping Ming, Lei He, and Frank K Soong · 2020
Later among the works it cites.
Multi-reference neural tts stylization with adversarial cycle consistency
Matt Whitehill, Shuang Ma, Daniel McDuff, and Yale Song · 2020
Later among the works it cites.
Quasi-periodic parallel wavegan vocoder: A non-autoregressive pitch-dependent dilated convolution model for parametric speech generation
Yi-Chiao Wu, Tomoki Hayashi, Takuma Okamoto, Hisashi Kawai, and Tomoki Toda · 2020
Later among the works it cites.
Improving prosody with linguistic and bert derived features in multi-speaker based mandarin chinese neural tts
Yujia Xiao, Lei He, Huaiping Ming, and Frank K Soong · 2020
Later among the works it cites.
Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim · 2020
Later among the works it cites.
Towards universal text-to-speech
Jingzhou Yang and Lei He · 2020
Later among the works it cites.
Vocgan: A high-fidelity real-time vocoder with a hierarchically-nested adversarial network
Jinhyeok Yang, Junmo Lee, Youngik Kim, Hoon-Young Cho, and Injung Kim · 2020
Later among the works it cites.
Durian: Duration informed attention network for speech synthesis
Chengzhu Yu, Heng Lu, Na Hu, Meng Yu, Chao Weng, Kun Xu, Peng Liu, Deyi Tuo, Shiyin Kang, Guangzhi Lei, et al · 2020
Later among the works it cites.
Aligntts: Efficient feed-forward text-to-speech system without explicit alignment
Zhen Zeng, Jianzong Wang, Ning Cheng, Tian Xia, and Jing Xiao · 2020
Later among the works it cites.
Prosody learning mechanism for speech synthesis system without text length limit
Zhen Zeng, Jianzong Wang, Ning Cheng, and Jing Xiao · 2020
Later among the works it cites.
Squeezewave: Extremely lightweight vocoders for on-device speech synthesis
Bohan Zhai, Tianren Gao, Flora Xue, Daniel Rothchild, Bichen Wu, Joseph E Gonzalez, and Kurt Keutzer · 2020
Later among the works it cites.
Unsupervised learning for sequence-to-sequence text-to-speech for low-resource languages
Haitong Zhang and Yue Lin · 2020
Later among the works it cites.
A hybrid text normalization system using multi-head self-attention for mandarin
Junhui Zhang, Junjie Pan, Xiang Yin, Chen Li, Shichao Liu, Yang Zhang, Yuxuan Wang, and Zejun Ma · 2020
Later among the works it cites.
Towards natural bilingual and code-switched speech synthesis based on mix of monolingual recordings and cross-lingual voice conversion
Shengkui Zhao, Trung Hieu Nguyen, Hao Wang, and Bin Ma · 2020
Later among the works it cites.
End-to-end code-switching tts with cross-lingual language model
Xuehao Zhou, Xiaohai Tian, Grandee Lee, Rohan Kumar Das, and Haizhou Li · 2020
Later among the works it cites.
Improving performance of seen and unseen speech style transfer in end-to-end neural tts
Xiaochun An, Frank K Soong, and Lei Xie · 2021
Closest in time.
Hi-fi multi-speaker english tts dataset
Evelina Bakhturina, Vitaly Lavrukhin, Boris Ginsburg, and Yang Zhang · 2021
Closest in time.
Stanislav Beliaev and Boris Ginsburg · 2021
Closest in time.
Callhome american english speech
Alexandra Canavan, Graff David, and Zipperlen George · 2021
Closest in time.
Sc-glowtts: an efficient zero-shot multi-speaker text-to-speech model
Edresson Casanova, Christopher Shulby, Eren Gölge, Nicolas Michael Müller, Frederico Santos de Oliveira, Arnaldo Candido Junior, Anderson da Silva Soares, Sandra Maria Aluisio, and Moacir Antonelli Ponti · 2021
Closest in time.
Speech bert embedding for improving prosody in neural tts
Liping Chen, Yan Deng, Xi Wang, Frank K Soong, and Lei He · 2021
Closest in time.
Hierarchical prosody modeling for non-autoregressive speech synthesis
Chung-Ming Chien and Hung-yi Lee · 2021
Closest in time.
Chung-Ming Chien, Jheng-Hao Lin, Chien-yu Huang, Po-chun Hsu, and Hung-yi Lee · 2021
Closest in time.
Jian Cong, Shan Yang, Lei Xie, and Dan Su · 2021
Closest in time.
End-to-end adversarial text-to-speech
Jeff Donahue, Sander Dieleman, Mikołaj Bińkowski, Erich Elsen, and Karen Simonyan · 2021
Closest in time.
Mixture density network for phone-level prosody modelling in speech synthesis
Chenpeng Du and Kai Yu · 2021
Closest in time.
Parallel tacotron 2: A non-autoregressive neural tts model with differentiable duration modeling
Isaac Elias, Heiga Zen, Jonathan Shen, Yu Zhang, Jia Ye, RJ Ryan, and Yonghui Wu · 2021
Closest in time.
An asymmetric cycle-consistency loss for dealing with many-to-one mappings in image translation: a study on thigh mr scans
Michael Gadermayr, Maximilian Tschuchnig, Laxmi Gupta, Nils Krämer, Daniel Truhn, D Merhof, and Burkhard Gess · 2021
Closest in time.
Multilingual byte2speech text-to-speech models are few-shot spoken language learners
Mutian He, Jingzhou Yang, and Lei He · 2021
Closest in time.
Whispered and lombard neural speech synthesis
Qiong Hu, Tobias Bleisch, Petko Petkov, Tuomo Raitio, Erik Marchi, and Varun Lakshminarasimhan · 2021
Closest in time.
Model architectures to extrapolate emotional expressions in dnn-based text-to-speech
Katsuki Inoue, Sunao Hara, Masanobu Abe, Nobukatsu Hojo, and Yusuke Ijima · 2021
Closest in time.
Diff-tts: A denoising diffusion model for text-to-speech
Myeonghun Jeong, Hyeongju Kim, Sung Jun Cheon, Byoung Jin Choi, and Nam Soo Kim · 2021
Closest in time.
Png bert: Augmented bert on phonemes and graphemes for neural tts
Ye Jia, Heiga Zen, Jonathan Shen, Yu Zhang, and Yonghui Wu · 2021
Closest in time.
Universal neural vocoding with parallel wavenet
Yunlong Jiao, Adam Gabrys, Georgi Tinchev, Bartosz Putrycz, Daniel Korzekwa, and Viacheslav Klimkov · 2021
Closest in time.
Fast dctts: Efficient deep convolutional text-to-speech
Minsu Kang, Jihyun Lee, Simin Kim, and Injung Kim · 2021
Closest in time.
On fast sampling of diffusion probabilistic models
Zhifeng Kong and Wei Ping · 2021
Closest in time.
Diffwave: A versatile diffusion model for audio synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2021
Closest in time.
Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis
Yi Lei, Shan Yang, and Lei Xie · 2021
Closest in time.
Controllable emotion transfer for end-to-end speech synthesis
Tao Li, Shan Yang, Liumeng Xue, and Lei Xie · 2021
Closest in time.
Shilun Lin, Fenglong Xie, Li Meng, Xinhui Li, and Li Lu · 2021
Closest in time.
Lightspeech: Lightweight and fast text to speech with neural architecture search
Renqian Luo, Xu Tan, Rui Wang, Tao Qin, Jinzhu Li, Sheng Zhao, Enhong Chen, and Tie-Yan Liu · 2021
Closest in time.
Meta-stylespeech: Multi-speaker adaptive text-to-speech generation
Dongchan Min, Dong Bok Lee, Eunho Yang, and Sung Ju Hwang · 2021
Closest in time.
Review of end-to-end speech synthesis technology based on deep learning
Zhaoxi Mu, Xinyu Yang, and Yizhuo Dong · 2021
Closest in time.
Kazakhtts: An open-source kazakh text-to-speech synthesis dataset
Saida Mussakhojayeva, Aigerim Janaliyeva, Almas Mirzakhmetov, Yerbolat Khassanov, and Huseyin Atakan Varol · 2021
Closest in time.
Expressive neural voice cloning
Paarth Neekhara, Shehzeen Hussain, Shlomo Dubnov, Farinaz Koushanfar, and Julian McAuley · 2021
Closest in time.
Speech resynthesis from discrete disentangled self-supervised representations
Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, and Emmanuel Dupoux · 2021
Closest in time.
Grad-tts: A diffusion probabilistic model for text-to-speech
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov · 2021
Closest in time.
Data-efficient training strategies for neural tts systems
KR Prajwal and CV Jawahar · 2021
Closest in time.
Hui-audio-corpus-german: A high quality tts dataset
Pascal Puchtler, Johannes Wirth, and René Peinl · 2021
Closest in time.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu · 2021
Closest in time.
Supervised and unsupervised approaches for controlling narrow lexical focus in sequence-to-sequence speech synthesis
Slava Shechtman, Raul Fernandez, and David Haws · 2021
Closest in time.
Improved parallel wavegan vocoder with perceptually weighted spectrogram loss
Eunwoo Song, Ryuichi Yamamoto, Min-Jae Hwang, Jin-Seob Kim, Ohsung Kwon, and Jae-Min Kim · 2021
Closest in time.
Alternate endings: Improving prosody for incremental neural tts with predicted future text input
Brooke Stephenson, Thomas Hueber, Laurent Girin, and Laurent Besacier · 2021
Closest in time.
Cuhk-ee voice cloning system for icassp 2021 m2voc challenge
Daxin Tan, Hingpang Huang, Guangyan Zhang, and Tan Lee · 2021
Closest in time.
Tts tutorial at iscslp 2021
Xu Tan · 2021
Closest in time.
Tts tutorial at ijcai 2021
Xu Tan and Tao Qin · 2021
Closest in time.
Analysis and assessment of controllability of an expressive deep learning-based tts system
Noé Tits, Kevin El Haddad, and Thierry Dutoit · 2021
Closest in time.
Improve gan-based neural vocoder using pointwise relativistic leastsquare gan
Congyi Wang, Yu Chen, Bin Wang, and Yi Shi · 2021
Closest in time.
Learning to efficiently sample from diffusion probabilistic models
Daniel Watson, Jonathan Ho, Mohammad Norouzi, and William Chan · 2021
Closest in time.
Wave-tacotron: Spectrogram-free end-to-end text-to-speech synthesis
Ron J Weiss, RJ Skerry-Ryan, Eric Battenberg, Soroosh Mariooryad, and Diederik P Kingma · 2021
Closest in time.
Speech synthesis — Wikipedia, the free encyclopedia
Wikipedia · 2021
Closest in time.
The multi-speaker multi-style voice cloning challenge 2021
Qicong Xie, Xiaohai Tian, Guanghou Liu, Kun Song, Lei Xie, Zhiyong Wu, Hai Li, Song Shi, Haizhou Li, Fen Hong, et al · 2021
Closest in time.
Nas-bert: Task-agnostic and adaptive-size bert compression with neural architecture search
Jin Xu, Xu Tan, Renqian Luo, Kaitao Song, Jian Li, Tao Qin, and Tie-Yan Liu · 2021
Closest in time.
Cycle consistent network for end-to-end style transfer tts training
Liumeng Xue, Shifeng Pan, Lei He, Lei Xie, and Frank K Soong · 2021
Closest in time.
Adaspeech 2: Adaptive text to speech with untranscribed data
Yuzi Yan, Xu Tan, Bohan Li, Tao Qin, Sheng Zhao, Yuan Shen, and Tie-Yan Liu · 2021
Closest in time.
Reo Yoneyama, Yi-Chiao Wu, and Tomoki Toda · 2021
Closest in time.
Gan vocoder: Multi-resolution discriminator is all you need
Jaeseong You, Dalhyun Kim, Gyuhyeon Nam, Geumbyeol Hwang, and Gyeongsu Chae · 2021
Closest in time.
Exploring machine speech chain for domain adaptation and few-shot speaker adaptation
Fengpeng Yue, Yan Deng, Lei He, and Tom Ko · 2021
Closest in time.
Ryanspeech: A corpus for conversational text-to-speech synthesis
Rohola Zandie, Mohammad H. Mahoor, Julia Madse, and Eshrat S. Emamian · 2021
Closest in time.
Lvcnet: Efficient condition-dependent modeling network for waveform generation
Zhen Zeng, Jianzong Wang, Ning Cheng, and Jing Xiao · 2021
Closest in time.
Denoispeech: Denoising text to speech with frame-level noise modeling
Chen Zhang, Yi Ren, Xu Tan, Jinglin Liu, Kejun Zhang, Tao Qin, Sheng Zhao, and Tie-Yan Liu · 2021
Closest in time.
Yixuan Zhou, Changhe Song, Jingbei Li, Zhiyong Wu, and Helen Meng · 2021
Closest in time.
End-to-end text-to-speech for low-resource languages by cross-lingual transfer learning
Yuan-Jui Chen, Tao Tu, Cheng-chieh Yeh, and Hung-Yi Lee · 2079
Closest in time.
Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning
Yu Zhang, Ron J Weiss, Heiga Zen, Yonghui Wu, Zhifeng Chen, RJ Skerry-Ryan, Ye Jia, Andrew Rosenberg, and Bhuvana Ramabhadran · 2084
Closest in time.
Neural autoregressive flows
Chin-Wei Huang, David Krueger, Alexandre Lacoste, and Aaron Courville · 2087
Closest in time.