“Tacotron: Towards end-to-end speech synthesis,”
Y. Wang, R. Skerry-Ryan, D. Santon, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, · 2017
Cited alongside, same era.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Cited alongside, same era.
“Fastspeech: Fast, robust and controllable text to speech,”
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Y. Liu, · 2017
Cited alongside, same era.
“Mobilenets: Efficient convolutional neural networks for mobile vision applications,”
Original
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam, · 2017
Cited alongside, same era.
“The lj speech dataset,”
K. Ito, · 2017
Cited alongside, same era.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ Skerry-Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Y. Wu, · 2018
Cited alongside, same era.
“Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention,”
H. Tachibana, K. Uenoyama, and S. Aihara, · 2018
Cited alongside, same era.
“Neural speech synthesis with transformer network,”
N. Li, S. Liu, Y. Liu, S. Zhao, M. Liu, and M. Zhou, · 2018
Cited alongside, same era.