Fetching the paper…
Reading the bibliography…
While neural end-to-end text-to-speech (TTS) is superior to conventional statistical methods in many ways, the exposure bias problem in the autoregressive models remains an issue to be resolved.
“A voice conversion framework with tandem feature sparse representation and speaker-adapted wavenet vocoder,”
Berrak Sisman, Mingyang Zhang, and Haizhou Li, · 1982
Earlier work this paper cites.
“Signal estimation from modified short-time fourier transform,”
Daniel Griffin and Jae Lim, · 1984
Earlier work this paper cites.
“Mixture autoregressive hidden markov models for speech signals,”
Biing-Hwang Juang and Lawrence Rabiner, · 1985
Earlier work this paper cites.
“Imagenet classification with deep convolutional neural networks,”
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, · 2012
Earlier work this paper cites.
“Speech synthesis based on hidden markov models,”
Keiichi Tokuda, Yoshihiko Nankaku, Tomoki Toda, Heiga Zen, Junichi Yamagishi, and Keiichiro Oura, · 2013
Earlier work this paper cites.
“Statistical parametric speech synthesis using deep neural networks,”
Heiga Zen, Andrew Senior, and Mike Schuster, · 2013
Earlier work this paper cites.
“Scheduled sampling for sequence prediction with recurrent neural networks,”
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer, · 2015
Earlier work this paper cites.
“How (not) to train your generative model: Scheduled sampling, likelihood, adversary?,”
Ferenc Huszár, · 2015
Earlier work this paper cites.
“Sequence level training with recurrent neural networks,”
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba, · 2016
Earlier work this paper cites.
“Mongolian text-to-speech system based on deep neural network,”
Rui Liu, Feilong Bao, Guanglai Gao, and Yonghe Wang, · 2017
Earlier work this paper cites.
“Tacotron: A fully end-to-end text-to-speech synthesis model,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Earlier work this paper cites.
“An investigation of multi-speaker training for wavenet vocoder,”
Tomoki Hayashi, Akira Tamamori, Kazuhiro Kobayashi, Kazuya Takeda, and Tomoki Toda, · 2017
Cited alongside, same era.
“A gift from knowledge distillation: Fast optimization, network minimization and transfer learning,”
Junho Yim, Donggyu Joo, Jihoon Bae, and Junmo Kim, · 2017
Cited alongside, same era.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Cited alongside, same era.
“Adaptive wavenet vocoder for residual compensation in gan-based voice conversion,”
Berrak Sisman, Mingyang Zhang, Sakriani Sakti, Haizhou Li, and Satoshi Nakamura, · 2018
Cited alongside, same era.
“Improving mongolian phrase break prediction by using syllable and morphological embeddings with bilstm model.,”
Rui Liu, Feilong Bao, Guanglai Gao, Hui Zhang, and Yonghe Wang, · 2018
Cited alongside, same era.
“Training Multi-Speaker Neural Text-to-Speech Systems Using Speaker-Imbalanced Speech Corpora,”
Hieu-Thi Luong, Xin Wang, Junichi Yamagishi, and Nobuyuki Nishizawa, · 2019
Closest in time.
“Real-Time Neural Text-to-Speech with Sequence-to-Sequence Acoustic Model and WaveGlow or Single Gaussian WaveRNN Vocoders,”
Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, and Hisashi Kawai, · 2019
Closest in time.
“Group Sparse Representation with WaveNet Vocoder Adaptation for Spectrum and Prosody Conversion,”
Berrak Sisman, Mingyang Zhang, and Haizhou Li, · 2019
Closest in time.
“Generalization in generation: A closer look at exposure bias,”
Florian Schmidt, · 2019
Closest in time.
“Fastspeech: Fast, robust and controllable text to speech,”
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu, · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention,”
Hideyuki Tachibana, Katsuya Uenoyama, and Shunsuke Aihara, · 2018
Cited alongside, same era.
“Learning attribute representations with localization for flexible fashion search,”
Kenan E Ak, Ashraf A Kassim, Joo Hwee Lim, and Jo Yew Tham, · 2018
Cited alongside, same era.
“Error reduction network for dblstm-based voice conversion,”
Mingyang Zhang, Berrak Sisman, Sai Sirisha Rallabandi, Haizhou Li, and Li Zhao, · 2018
Cited alongside, same era.
“Robust and fine-grained prosody control of end-to-end speech synthesis,”
Younggun Lee and Taesu Kim, · 2019
Cited alongside, same era.
“Semi-supervised training for improving data efficiency in end-to-end speech synthesis,”
Yu-An Chung, Yuxuan Wang, Wei-Ning Hsu, Yu Zhang, and RJ Skerry-Ryan, · 2019
Cited alongside, same era.
“Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS,”
Mutian He, Yan Deng, and Lei He, · 2019
Cited alongside, same era.
“Pre-alignment guided attention for improving training efficiency and model stability in end-to-end speech synthesis,”
Xiaolian Zhu, Yuchao Zhang, Shan Yang, Liumeng Xue, and Lei Xie, · 2019
Closest in time.
“A New GAN-Based End-to-End TTS Training Algorithm,”
Haohan Guo, Frank K. Soong, Lei He, and Lei Xie, · 2019
Closest in time.
“Semantically consistent hierarchical text to fashion image synthesis with an enhanced-attentional generative adversarial network,”
Kenan Emir Ak, Joo Hwee Lim, Jo Yew Tham, and Ashraf Kassim, · 2019
Closest in time.
“The imu speech synthesis entry for blizzard challenge 2019,”
Rui Liu, Jingdong Li, Feilong Bao, and Guanglai Gao, · 2019
Closest in time.
“Wavetts: Tacotron-based tts with joint time-frequency domain loss,”
Rui Liu, Berrak Sisman, Feilong Bao, Guanglai Gao, and Haizhou Li, · 2020
Closest in time.