Fetching the paper…
Reading the bibliography…
Although neural end-to-end text-to-speech models can synthesize highly natural speech, there is still room for improvements to its efficiency and naturalness.
“A Learning Algorithm for Continually Running Fully Recurrent Neural Networks,”
R. J. Williams and D. Zipser, · 1989
Earlier work this paper cites.
“The Aligner: Text to Speech Alignment using Markov Models and a Pronunciation Dictionary,”
D. Talkin and C. W. Wightman, · 1994
Earlier work this paper cites.
“Statistical Parametric Speech Synthesis,”
H. Zen, K. Tokuda, and A. Black, · 2009
Earlier work this paper cites.
“Statistical Parametric Speech Synthesis Using Deep Neural Networks,”
H. Zen, A. Senior, and M. Schuster, · 2013
Earlier work this paper cites.
“Generating Sequences with Recurrent Neural Networks,”
A. Graves, · 2013
Earlier work this paper cites.
“Neural Machine Translation by Jointly Learning to Align and Translate,”
D. Bahdanau, K. Cho, and Y. Bengio, · 2015
Earlier work this paper cites.
“Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks,”
S. Bengio, O. Vinyals, N. Jaitly, and N. Shazeer, · 2015
Earlier work this paper cites.
“WaveNet: A Generative Model for Raw Audio,”
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, · 2016
Earlier work this paper cites.
“Professor Forcing: A New Algorithm for Training Recurrent Networks,”
A. Goyal, A. Lamb, Y. Zhang, S. Zhang, A. Courville, and Y. Bengio, · 2016
Earlier work this paper cites.
“Layer Normalization,” 2016
J. L. Ba, J. R. Kiros, and G. E. Hinton, · 2016
Earlier work this paper cites.
“Char2Wav: End-to-End Speech Synthesis,”
J. Sotelo, S. Mehri, K. Kumar, J. F. Santos, K. Kastner, A. C. Courville, and Y. Bengio, · 2017
Earlier work this paper cites.
“Tacotron: Towards End-to-End Speech Synthesis,”
Y. Wang, RJ Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, · 2017
Earlier work this paper cites.
“Non-Autoregressive Neural Machine Translation,”
J. Gu, J. Bradbury, C. Xiong, V. O. K. Li, and R. Socher, · 2017
Earlier work this paper cites.
“Attention Is All You Need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“Deep Voice 3: 2000-Speaker Neural Text-to-Speech,”
W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, · 2018
Cited alongside, same era.
“Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, RJ Skerrv-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, · 2018
Cited alongside, same era.
“Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis,”
Y. Wang, D. Stanton, Y. Zhang, RJ Skerry-Ryan, E. Battenberg, J. Shor, Y. Xiao, Y. Jia, F. Ren, and R. A. Saurous, · 2018
Cited alongside, same era.
“Deterministic Non-Autoregressive Neural Sequence Modeling by Iterative Refinement,”
J. Lee, E. Mansimov, and K. Cho, · 2018
Cited alongside, same era.
“An Intriguing Failing of Convolutional Neural Networks and the CoordConv Solution,”
R. Liu, J. Lehman, P. Molino, F. P. Such, E. Frank, A. Sergeev, and J. Yosinski, · 2018
“Pay Less Attention with Lightweight and Dynamic Convolutions,”
F. Wu, A. Fan, A. Baevski, Y. N. Dauphin, and M. Auli, · 2019
Later among the works it cites.
“Location-relative attention mechanisms for robust long-form speech synthesis,”
E. Battenberg, RJ Skerry-Ryan, S. Mariooryad, D. Stanton, D. Kao, M. Shannon, and T. Bagby, · 2020
Closest in time.
“FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,”
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2020
Closest in time.
“TalkNet: Fully-Convolutional Non-Autoregressive Speech Synthesis Model,”
S. Beliaev, Y. Rebryk, and B. Ginsburg, · 2020
Closest in time.
D. Lim, W. Jang, H. Park, B. Kim, and J. Yoon, · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron,”
RJ Skerry-Ryan, E. Battenberg, Y. Xiao, Y. Wang, D. Stanton, J. Shor, R. Weiss, R. Clark, and R. A. Saurous, · 2018
Cited alongside, same era.
“Efficient Neural Audio Synthesis,”
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. van den Oord, S. Dieleman, and K. Kavukcuoglu, · 2018
Cited alongside, same era.
“Neural Speech Synthesis with Transformer Network,”
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, · 2019
Cited alongside, same era.
“Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS,”
M. He, Y. Deng, and L. He, · 2019
Cited alongside, same era.
“Forward–backward decoding sequence for regularizing end-to-end tts,”
Y. Zheng, J. Tao, Z. Wen, and J. Yi, · 2019
Cited alongside, same era.
“A New GAN-Based End-to-End TTS Training Algorithm,”
H. Guo, F. K. Soong, L. He, and L. Xie, · 2019
Cited alongside, same era.
“FastSpeech: Fast, Robust and Controllable Text to Speech,”
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2019
Cited alongside, same era.
Closest in time.
“AlignTTS: Efficient Feed-Forward Text-to-Speech System without Explicit Alignment,”
Z. Zeng, J. Wang, N. Cheng, T. Xia, and J. Xiao, · 2020
Closest in time.
“Flow-TTS: A Non-Autoregressive Network for Text to Speech Based on Flow,”
C. Miao, S. Liang, M. Chen, J. Ma, S. Wang, and J. Xiao, · 2020
Closest in time.
“End-to-End Adversarial Text-to-Speech,”
J. Donahue, S. Dieleman, M. Bińkowski, E. Elsen, and K. Simonyan, · 2020
Closest in time.
“FastPitch: Parallel Text-to-speech with Pitch Prediction,”
A. Lańcucki, · 2020
Closest in time.
J. Shen, Y. Jia, M. Chrzanowski, Y. Zhang, I. Elias, H. Zen, and Y. Wu, · 2020
Closest in time.
“Fully-Hierarchical Fine-Grained Prosody Modeling for Interpretable Speech Synthesis,”
G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, and Y. Wu, · 2020
Closest in time.
“DEJA-VU: Double Feature Presentation and Iterated Loss in Deep Transformer Networks,”
A. Tjandra, C. Liu, F. Zhang, X. Zhang, Y. Wang, G. Synnaeve, S. Nakamura, and G. Zweig, · 2020
Closest in time.
“Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta Posterior,”
R. Shu, J. Lee, H. Nakayama, and K. Cho, · 2020
Closest in time.
“Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning,”
Y. Zhang, R. J. Weiss, H. Zen, Wu Y., Z. Chen, RJ Skerry-Ryan, Y. Jia, A. Rosenberg, and B. Ramabhadran, · 2084
Closest in time.