Fetching the paper…
Reading the bibliography…
In this work, we address the Text-to-Speech (TTS) task by proposing a non-autoregressive architecture called EfficientTTS.
Neural Machine Translation by Jointly Learning to Align and Translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, P. and Ba, J · 2015
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Earlier work this paper cites.
The lj speech dataset
Ito, K · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., N.Gomez, A., Kaiser, Å., and Polosukhin, I · 2017
Earlier work this paper cites.
Tacotron: Towards End-to-End Speech Synthesis
Wang, Y., Skerry-Ryan, R., Stanton, D., Y. Wu, R. J. W., Jaitly, N., and Yang, Z · 2017
Earlier work this paper cites.
Deep Voice 3: 2000-Speaker Neural Text-to-Speech
Ping, W., Peng, K., Gibiansky, A., Arik, S. O., Kannan, A., Narang, S., Raiman, J., and Miller, J · 2018
Earlier work this paper cites.
Natural TTS synthesis by conditioning Wavenet on mel spectrogram predictions
Shen, J., Pang, R., Weiss, R. J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., and SkerryRyan, R · 2018
Cited alongside, same era.
Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention
Tachibana, H., Uenoyama, K., and Aihara, S · 2018
Cited alongside, same era.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kumar, K., Kumar, R., de Boissiere, T., Gestin, L., Teoh, W. Z., Sotelo, J., de Brebisson, A., Bengio, Y., and Courville., A · 2019
Cited alongside, same era.
Close to Human Quality TTS with Transformer
Li, N., Liu, S., Liu, Y., Zhao, S., Liu, M., and Zhou, M · 2019
Cited alongside, same era.
Parallel neural text-to-speech
Peng, K., Ping, W., Song, Z., and Zhao, K · 2019
Cited alongside, same era.
End-to-End Adversarial Text-to-Speech
Donahue, J., Dieleman, S., Binkowski, M., Elsen, E., and Simonyan, K · 2020
Closest in time.
Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search
Kim, J., Kim, S., Kong, J., and Yoon, S · 2020
Closest in time.
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
Kong, J., Kim, J., and Bae, J · 2020
Closest in time.
MoBoAligner: a Neural Alignment Model for Non-autoregressive TTS with Monotonic Boundary Search
Li, N., Liu, S., Liu, Y., Zhao, S., Liu, M., and Zhou, M · 2020
Closest in time.
Flow-TTS: A non-autoregressive network for text to speech based on flow
Miao, C., Liang, S., Chen, M., Ma, J., Wang, S., and Xiao, J · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech
Ping, W., Peng, K., and Chen, J · 2019
Cited alongside, same era.
FastSpeech: Fast, Robust and Controllable Text to Speech
Ren, Y., Ruan, Y., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2019
Cited alongside, same era.
FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Ren, Y., Hu, C., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2020
Closest in time.
Flowtron: an autoregressive flow-based generative network for text- to-speech synthesis
Valle, R., Shih, K., Prenger, R., and Catanzaro, B · 2020
Closest in time.