Fetching the paper…
Reading the bibliography…
The mainstream neural text-to-speech(TTS) pipeline is a cascade system, including an acoustic model(AM) that predicts acoustic feature from the input transcript and a vocoder that generates waveform according to the given acoustic feature.
L. Rabiner, M. Cheng, A. Rosenberg, and C. McGonegal, “A comparative performance study of several pitch detection algorithms,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 24, no. 5, pp. 399–418, 1976
1976
Earlier work this paper cites.
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in Proc. IEEE ICASSP , vol. 2, 2001, pp. 749–752 vol.2
2001
Earlier work this paper cites.
H. Zen, K. Tokuda, and A. W. Black, “Statistical parametric speech synthesis,” speech communication , vol. 51, no. 11, pp. 1039–1064, 2009
2009
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial nets,” Proc. NIPS , 2014
2014
Earlier work this paper cites.
P. Ghahremani, B. BabaAli, D. Povey, K. Riedhammer, J. Trmal, and S. Khudanpur, “A pitch extraction algorithm tuned for automatic speech recognition,” in Proc. IEEE ICASSP , 2014, pp. 2494–2498
2014
Earlier work this paper cites.
2016
Earlier work this paper cites.
Y. Wang, R. J. Skerry-Ryan, D. Stanton et al. , “Tacotron: Towards end-to-end speech synthesis,” in Proc. ISCA Interspeech , 2017, pp. 4006–4010
2017
Earlier work this paper cites.
K. Ito, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, “Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,” in Proc. IEEE ICASSP , 2018, pp. 4779–4783
2018
Earlier work this paper cites.
X. Wang, S. Takaki, and J. Yamagishi, “Neural source-filter waveform models for statistical parametric speech synthesis,” IEEE/ACM Trans. ASLP. , vol. 28, pp. 402–415, 2019
2019
Earlier work this paper cites.
X. Wang, S. Takaki, J. Yamagishi, S. King, and K. Tokuda, “A vector quantized variational autoencoder (vq-vae) autoregressive neural f _ 0 f\_0 model for statistical parametric speech synthesis,” IEEE/ACM Trans. ASLP. , vol. 28, pp. 157–170, 2019
2019
Earlier work this paper cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition,” in Proc. ISCA Interspeech , 2019, pp. 3465–3469
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” Proc. NeurIPS , vol. 33, pp. 17 022–17 033, 2020
2020
Cited alongside, same era.
Z. Liu, K. Chen, and K. Yu, “Neural homomorphic vocoder,” in Proc. ISCA Interspeech , 2020, pp. 240–244
2020
Cited alongside, same era.
G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, and Y. Wu, “Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis,” in Proc. IEEE ICASSP , 2020, pp. 6264–6268
2020
Cited alongside, same era.
Y. Hono, K. Tsuboi, K. Sawada, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, “Hierarchical multi-grained generative model for expressive speech synthesis,” in Proc. ISCA Interspeech , 2020, pp. 3441–3445
2020
Cited alongside, same era.
C. M. Chien and H. Y. Lee, “Hierarchical prosody modeling for non-autoregressive speech synthesis,” in Proc. IEEE SLT , 2021, pp. 446–453
2021
Later among the works it cites.
R. Liu, B. Sisman, and H. Li, “Graphspeech: Syntax-aware graph attention network for neural speech synthesis,” in Proc. IEEE ICASSP , 2021, pp. 6059–6063
2021
Later among the works it cites.
F. Shen, C. Du, and K. Yu, “Acoustic word embeddings for end-to-end speech synthesis,” Applied Sciences , vol. 11, no. 19, 2021
2021
Later among the works it cites.
Y. Jia, H. Zen, J. Shen, Y. Zhang, and Y. Wu, “PnG BERT: Augmented BERT on Phonemes and Graphemes for Neural TTS,” in Proc. ISCA Interspeech , 2021, pp. 151–155
2021
Later among the works it cites.
G. Papamakarios, E. Nalisnick, D. J. Rezende, S. Mohamed, and B. Lakshminarayanan, “Normalizing flows for probabilistic modeling and inference,” Journal of Machine Learning Research , vol. 22, no. 57, pp. 1–64, 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, A. Rosenberg, B. Ramabhadran, and Y. Wu, “Generating diverse and natural text-to-speech samples using a quantized fine-grained VAE and autoregressive prosody prior,” in Proc. IEEE ICASSP , 2020, pp. 6699–6703
2020
Cited alongside, same era.
A. Sun, J. Wang, N. Cheng, H. Peng, Z. Zeng, and J. Xiao, “GraphTTS: graph-to-sequence modelling in neural text-to-speech,” in Proc. IEEE ICASSP , 2020, pp. 6719–6723
2020
Cited alongside, same era.
C. Miao, S. Liang, M. Chen, J. Ma, S. Wang, and J. Xiao, “Flow-tts: A non-autoregressive network for text to speech based on flow,” in Proc. IEEE ICASSP , 2020, pp. 7209–7213
2020
Cited alongside, same era.
J. Kim, S. Kim, J. Kong, and S. Yoon, “Glow-tts: A generative flow for text-to-speech via monotonic alignment search,” Proc. NeurIPS , vol. 33, pp. 8067–8077, 2020
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Proc. NeurIPS , 2020
2020
Cited alongside, same era.
A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in Proc. ISCA Interspeech , 2020, pp. 5036–5040
2020
Cited alongside, same era.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in Proc. ICLR , 2021
2021
Cited alongside, same era.
G. Yang, S. Yang, K. Liu, P. Fang, W. Chen, and L. Xie, “Multi-band MelGAN: Faster waveform generation for high-quality text-to-speech,” in Proc. IEEE SLT , 2021, pp. 492–498
2021
Cited alongside, same era.
2021
Later among the works it cites.
W. Hsu, B. Bolte, Y. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM Trans. ASLP. , vol. 29, pp. 3451–3460, 2021
2021
Later among the works it cites.
H.-S. Choi, J. Lee, W. Kim, J. Lee, H. Heo, and K. Lee, “Neural analysis and synthesis: Reconstructing speech from self-supervised representations,” Proc. NeurIPS , 2021
2021
Later among the works it cites.
A. Polyak, Y. Adi, J. Copet, E. Kharitonov, K. Lakhotia, W. Hsu, A. Mohamed, and E. Dupoux, “Speech resynthesis from discrete disentangled self-supervised representations,” in Proc. ISCA Interspeech , 2021, pp. 3615–3619
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
C. Du and K. Yu, “Phone-level prosody modelling with gmm-based mdn for diverse and controllable speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 190–201, 2022
2022
Closest in time.