Fetching the paper…
Reading the bibliography…
Several fully end-to-end text-to-speech (TTS) models have been proposed that have shown better performance compared to cascade models (i.e., training acoustic and vocoder models separately).
“Continuous F0 modeling for HMM based statistical parametric speech synthesis,”
K. Yu and S. Young, · 2010
Earlier work this paper cites.
“Auto-encoding variational Bayes,”
D. P. Kingma and M. Welling, · 2014
Earlier work this paper cites.
“Variational inference with normalizing flows,”
D. J. Rezende and S. Mohamed, · 2015
Earlier work this paper cites.
“WaveNet: A generative model for raw audio,”
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, · 2016
Earlier work this paper cites.
“Autoencoding beyond pixels using a learned similarity metric,”
A. B. L. Larsen, S. K. Sønderby, H. Larochelle, and O. Winther, · 2016
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis,”
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio, Q. Le, Y. Agiomyrgiannakis, R. Clark, and R. A. Saurous, · 2017
Earlier work this paper cites.
“Least squares generative adversarial networks,”
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. Paul Smolley, · 2017
Earlier work this paper cites.
“Effective spectral and excitation modeling techniques for LSTM-RNN-based speech synthesis systems,”
E. Song, F. K. Soong, and H.-G. Kang, · 2017
Earlier work this paper cites.
“Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, et al., · 2018
Earlier work this paper cites.
“FastSpeech: Fast, robust and controllable text to speech,”
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2019
Earlier work this paper cites.
“Neural source-filter-based waveform model for statistical parametric speech synthesis,”
X. Wang, S. Takaki, and J. Yamagishi, · 2019
Cited alongside, same era.
“Investigation of enhanced Tacotron text-to-speech synthesis systems with self-attention for pitch accent language,”
Y. Yasuda, X. Wang, S. Takaki, and J. Yamagishi, · 2019
Cited alongside, same era.
“Neural source-filter waveform models for statistical parametric speech synthesis,”
X. Wang, S. Takaki, and J. Yamagishi, · 2019
Cited alongside, same era.
“Decoupled weight decay regularization,”
I. Loshchilov and F. Hutter, · 2019
Cited alongside, same era.
“Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,”
R. Yamamoto, E. Song, and J.-M. Kim, · 2020
Cited alongside, same era.
“End-to-end adversarial text-to-speech,”
J. Donahue, S. Dieleman, M. Bińkowski, E. Elsen, and K. Simonyan, · 2021
Later among the works it cites.
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
J. Kim, J. Kong, and J. Son, · 2021
Later among the works it cites.
“PeriodNet: A non-autoregressive waveform generation model with a structure separating periodic and aperiodic components,”
Y. Hono, S. Takaki, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, · 2021
Later among the works it cites.
“High-fidelity Parallel WaveGAN with multi-band harmonic-plus-noise model,”
M.-J. Hwang, R. Yamamoto, E. Song, and J.-M. Kim, · 2021
Later among the works it cites.
“Parallel Tacotron: Non-autoregressive and controllable TTS,”
I. Elias, H. Zen, J. Shen, Y. Zhang, Y. Jia, R. J. Weiss, and Y. Wu, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Kong, J. Kim, and J. Bae, · 2020
Cited alongside, same era.
“ESPnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit,”
T. Hayashi, R. Yamamoto, K. Inoue, T. Yoshimura, S. Watanabe, T. Toda, K. Takeda, Y. Zhang, and X. Tan, · 2020
Cited alongside, same era.
“On the variance of the adaptive learning rate and beyond,”
L. Liu, H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and J. Han, · 2020
Cited alongside, same era.
“A survey on neural speech synthesis,”
X. Tan, T. Qin, F. Soong, and T.-Y. Liu, · 2021
Cited alongside, same era.
“FastSpeech 2: Fast and high-quality end-to-end text-to-speech,”
Y. Ren, C. Hu, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, · 2021
Cited alongside, same era.
D. Lim, S. Jung, and E. Kim, · 2022
Closest in time.
“VISinger: Variational inference with adversarial learning for end-to-end singing voice synthesis,”
Y. Zhang, J. Cong, H. Xue, L. Xie, P. Zhu, and M. Bi, · 2022
Closest in time.
“Chunked autoregressive GAN for conditional waveform synthesis,”
M. Morrison, R. Kumar, K. Kumar, P. Seetharaman, A. Courville, and Y. Bengio, · 2022
Closest in time.
“Unified Source-Filter GAN with Harmonic-plus-Noise Source Excitation Generation,”
R. Yoneyama, Y.-C. Wu, and T. Toda, · 2022
Closest in time.
“Period-HiFi-GAN: Fast and fundamental frequency controllable neural vocoder,”
K. Matsubara, T. Okamoto, R. Takashima, T. Takiguchi, T. Toda, and H. Kawai, · 2022
Closest in time.