Fetching the paper…
Reading the bibliography…
Attention-based end-to-end text-to-speech synthesis (TTS) is superior to conventional statistical methods in many ways.
“Signal estimation from modified short-time fourier transform,”
Daniel Griffin and Jae Lim, · 1984
Earlier work this paper cites.
“Mel-cepstral distance measure for objective speech quality assessment,”
R Kubichek, · 1993
Earlier work this paper cites.
“Unit selection in a concatenative speech synthesis system using a large speech database,”
Andrew J Hunt and Alan W Black, · 1996
Earlier work this paper cites.
“Maltparser: A language-independent system for data-driven dependency parsing,”
Joakim Nivre, Johan Hall, Jens Nilsson, Atanas Chanev, and Erwin Marsi, · 2007
Earlier work this paper cites.
Text-to-speech synthesis
Paul Taylor, · 2009
Earlier work this paper cites.
“Statistical parametric speech synthesis,”
Heiga Zen, Keiichi Tokuda, and Alan W Black, · 2009
Earlier work this paper cites.
“The graph neural network model,”
Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini, · 2009
Earlier work this paper cites.
“Statistical parametric speech synthesis using deep neural networks,”
Heiga Zen, Andrew Senior, and Mike Schuster, · 2013
Earlier work this paper cites.
“Learning phrase representations using rnn encoder–decoder for statistical machine translation,”
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Tacotron: A fully end-to-end text-to-speech synthesis model,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“The lj speech dataset,”
K. Ito, · 2017
Cited alongside, same era.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Cited alongside, same era.
“Improving mongolian phrase break prediction by using syllable and morphological embeddings with bilstm model.,”
Rui Liu, Feilong Bao, Guanglai Gao, Hui Zhang, and Yonghe Wang, · 2018
Cited alongside, same era.
“A lstm approach with sub-word embeddings for mongolian phrase break prediction,”
Rui Liu, Feilong Bao, Guanglai Gao, Hui Zhang, and Yonghe Wang, · 2018
Cited alongside, same era.
“Graph-to-sequence learning using gated graph neural networks,”
Daniel Beck, Gholamreza Haffari, and Trevor Cohn, · 2018
Cited alongside, same era.
“Wavelet analysis of speaker dependent and independent prosody for voice conversion.,”
“Improving prosody with linguistic and bert derived features in multi-speaker based mandarin chinese neural tts,”
Yujia Xiao, Lei He, Huaiping Ming, and Frank K Soong, · 2020
Closest in time.
“Acoustic model-based subword tokenization and prosodic-context extraction without language knowledge for text-to-speech synthesis,”
Masashi Aso, Shinnosuke Takamichi, Norihiro Takamune, and Hiroshi Saruwatari, · 2020
Closest in time.
“Graphtts: graph-to-sequence modelling in neural text-to-speech,”
Aolan Sun, Jianzong Wang, Ning Cheng, Huayi Peng, Zhen Zeng, and Jing Xiao, · 2020
Closest in time.
“Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,”
Kun Zhou, Berrak Sisman, Rui Liu, and Haizhou Li, · 2020
Closest in time.
“SG-Net: Syntax-guided machine reading comprehension,”
Zhuosheng Zhang, Yuwei Wu, Junru Zhou, Sufeng Duan, Hai Zhao, and Rui Wang, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Berrak Sisman and Haizhou Li, · 2018
Cited alongside, same era.
“Neural speech synthesis with transformer network,”
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu, · 2019
Cited alongside, same era.
“Teacher-student training for robust tacotron-based tts,”
Rui Liu, Berrak Sisman, Jingdong Li, Feilong Bao, Guanglai Gao, and Haizhou Li, · 2020
Cited alongside, same era.
“Modeling prosodic phrasing with multi-task learning in tacotron-based tts,”
Rui Liu, Berrak Sisman, Feilong Bao, Guanglai Gao, and Haizhou Li, · 2020
Cited alongside, same era.
“Expressive tts training with frame and style reconstruction loss,”
Rui Liu, Berrak Sisman, Guanglai Gao, and Haizhou Li, · 2020
Cited alongside, same era.
“On the localness modeling for the self-attention based end-to-end speech synthesis,”
Shan Yang, Heng Lu, Shiyin Kang, Liumeng Xue, Jinba Xiao, Dan Su, Lei Xie, and Dong Yu, · 2020
Cited alongside, same era.
“Graph transformer for graph-to-sequence learning,”
Deng Cai and Wai Lam, · 2020
Closest in time.
“Global greedy dependency parsing,”
Zuchao Li, Hai Zhao, and Kevin Parnow, · 2020
Closest in time.
“Stanza: A Python natural language processing toolkit for many human languages,”
Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning, · 2020
Closest in time.
“An overview of voice conversion and its challenges: From statistical modeling to deep learning,”
Berrak Sisman, Junichi Yamagishi, Simon King, and Haizhou Li, · 2021
Closest in time.
“Exploiting morphological and phonological features to improve prosodic phrasing for mongolian speech synthesis,”
Rui Liu, Berrak Sisman, Feilong Bao, Jichen Yang, Guanglai Gao, and Haizhou Li, · 2021
Closest in time.