Fetching the paper…
Reading the bibliography…
Despite the ability to produce human-level speech for in-domain text, attention-based end-to-end text-to-speech (TTS) systems suffer from text alignment failures that increase in frequency for out-of-domain text.
“Mel-cepstral distance measure for objective speech quality assessment,”
R Kubichek, · 1993
Earlier work this paper cites.
“Generating Sequences With Recurrent Neural Networks,”
Alex Graves, · 2013
Earlier work this paper cites.
“Neural Machine Translation by Jointly Learning to Align and Translate,”
Dzmitry Bahdanau, KyungHyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Attention-based Models for Speech Recognition,”
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Tacotron: Towards End-to-End Speech Synthesis,”
Yuxuan Wang, R.J. Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous, · 2017
Earlier work this paper cites.
“Online and Linear-time Attention by Enforcing Monotonic Alignments,”
Colin Raffel, Minh-Thang Luong, Peter J. Liu, Ron J. Weiss, and Douglas Eck, · 2017
Earlier work this paper cites.
“The LJ Speech Dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
Keith Ito, · 2017
Cited alongside, same era.
“Natural TTS Synthesis by Conditioning Wavenet on Mel Spectrogram Predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu, · 2018
Cited alongside, same era.
“Forward attention in sequence- to-sequence acoustic modeling for speech synthesis,”
J. Zhang, Z. Ling, and L. Dai, · 2018
Cited alongside, same era.
“Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron,”
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron Weiss, Rob Clark, and Rif A. Saurous, · 2018
Cited alongside, same era.
“Efficient Neural Audio Synthesis,”
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron van den Oord, Sander Dieleman, and Koray Kavukcuoglu, · 2018
Cited alongside, same era.
“Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS,”
Mutian He, Yan Deng, and Lei He, · 2019
Closest in time.
“Maximizing Mutual Information for Tacotron,”
Peng Liu, Xixin Wu, Shiyin Kang, Guangzhi Li, Dan Su, and Dong Yu, · 2019
Closest in time.
“Representation Mixing for TTS Synthesis,”
K. Kastner, J. F. Santos, Y. Bengio, and A. Courville, · 2019
Closest in time.
“Effective Use of Variational Embedding Capacity in Expressive End-to-End Speech Synthesis,”
E Battenberg, Soroosh Mariooryad, Daisy Stanton, R J Skerry-Ryan, Matt Shannon, David Kao, and Tom Bagby, · 2019
Closest in time.
“Robust and Fine-grained Prosody Control of End-to-end Speech Synthesis,”
Y. Lee and T. Kim, · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…