Fetching the paper…
Reading the bibliography…
This paper proposes a hierarchical, fine-grained and interpretable latent variable model for prosody based on the Tacotron 2 text-to-speech model.
“YIN, a fundamental frequency estimator for speech and music,”
A. de Cheveigné and H. Kawahara, · 2002
Earlier work this paper cites.
“Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,”
W. Chu and A. Alwan, · 2009
Earlier work this paper cites.
“Experimental and theoretical advances in prosody: A review,”
M. Wagner and D. G. Watson, · 2010
Earlier work this paper cites.
“RNADE: The real-valued neural autoregressive density-estimator,”
B. Uria, I. Murray, and H. Larochelle, · 2013
Earlier work this paper cites.
“The blizzard challenge 2013,”
S. King and V. Karaiskos, · 2013
Earlier work this paper cites.
“Sequence to sequence learning with neural networks,”
I. Sutskever, O. Vinyals, and Q. V. Le, · 2014
Earlier work this paper cites.
“Infogan: Interpretable representation learning by information maximizing generative adversarial nets,”
X. Chen, Y. Duan, R. Houthooft, et al., · 2016
Earlier work this paper cites.
“Disentangling factors of variation in deep representations using adversarial training,”
M. Mathieu, J. Zhao, P. Sprechmann, A. Ramesh, and Y. LeCun, · 2016
Earlier work this paper cites.
“Char2wav: End-to-end speech synthesis.,”
J. Sotelo, S. Mehri, K. Kumar, et al., · 2017
Earlier work this paper cites.
“Tacotron: Towards end-to-end speech synthesis.,”
Y. Wang, R. Skerry-Ryan, D. Stanton, et al., · 2017
Earlier work this paper cites.
“Unsupervised learning of disentangled and interpretable representations from sequential data.,”
W.-N. Hsu, Y. Zhang, and J. Glass, · 2017
Earlier work this paper cites.
“Beta-VAE: Learning basic visual concepts with a constrained variational framework,”
I. Higgins, L. Matthey, A. Pal, et al., · 2017
Earlier work this paper cites.
“Variational inference of disentangled latent concepts from unlabeled observations,”
A. Kumar, P. Sattigeri, and A. Balakrishnan, · 2017
Earlier work this paper cites.
“Learning disentangled representations with semi-supervised deep generative models,”
S. Narayanaswamy, T. B. Paige, J.-W. van de Meent, et al., · 2017
Cited alongside, same era.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, et al., · 2017
Cited alongside, same era.
“Masked autoregressive flow for density estimation,”
G. Papamakarios, T. Pavlakou, and I. Murray, · 2017
Cited alongside, same era.
“Deep Voice 3: 2000-speaker neural text-to-speech.,”
W. Ping, K. Peng, A. Gibiansky, et al., · 2018
Cited alongside, same era.
“Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions.,”
J. Shen, R. Pang, R. J. Weiss, et al., · 2018
Cited alongside, same era.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Y. Wang, D. Stanton, Y. Zhang, et al., · 2018
“A generative adversarial network for style modeling in a text-to-speech system,”
S. Ma, D. Mcduff, and Y. Song., · 2019
Later among the works it cites.
“Disentangling correlated speaker and noise for speech synthesis via data augmentation and adversarial factorization,”
W.-N. Hsu, Y. Zhang, R. J. Weiss, et al., · 2019
Later among the works it cites.
“Effective use of variational embedding capacity in expressive end-to-end speech synthesis,”
E. Battenberg, S. Mariooryad, D. Stanton, et al., · 2019
Later among the works it cites.
“Robust and fine-grained prosody control of end-to-end speech synthesis,”
Y. Lee and T. Kim, · 2019
Later among the works it cites.
“Learning latent representations for style control and transfer in end-to-end speech synthesis,”
Y.-J. Zhang, S. Pan, L. He, and Z.-H. Ling, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron,”
R. Skerry-Ryan, E. Battenberg, Y. Xiao, et al., · 2018
Cited alongside, same era.
“Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, et al., · 2018
Cited alongside, same era.
“Deep encoder-decoder models for unsupervised learning of controllable speech synthesis,”
G. E. Henter, J. Lorenzo-Trueba, X. Wang, and J. Yamagishi, · 2018
Cited alongside, same era.
“Expressive speech synthesis via modeling expressions with variational autoencoder,”
K. Akuzawa, Y. Iwasawa, and Y. Matsuo, · 2018
Cited alongside, same era.
“Unsupervised adaptation with interpretable disentangled representations for distant conversational speech recognition,”
W.-N. Hsu, H. Tang, and J. Glass, · 2018
Cited alongside, same era.
“Disentangling by factorising,”
H. Kim and A. Mnih, · 2018
Cited alongside, same era.
A. Razavi, A. van den Oord, and O. Vinyals, · 2019
Later among the works it cites.
“Semi-supervised training for improving data efficiency in end-to-end speech synthesis,”
Y.-A. Chung, Y. Wang, W.-N. Hsu, Y. Zhang, and R. Skerry-Ryan, · 2019
Later among the works it cites.
“Challenging common assumptions in the unsupervised learning of disentangled representations,”
F. Locatello, S. Bauer, M. Lucic, et al., · 2019
Later among the works it cites.
“Semi-supervised learning by disentangling and self-ensembling over stochastic latent space,”
P. K. Gyawali, Z. Li, S. Ghimire, and L. Wang, · 2019
Later among the works it cites.
“Semi-supervised generative modeling for controllable speech synthesis,”
R. Habib, S. Mariooryad, M. Shannon, et al., · 2019
Later among the works it cites.
“Structured disentangled representations,”
B. Esmaeili, H. Wu, S. Jain, et al., · 2019
Later among the works it cites.
“LibriTTS: A corpus derived from LibriSpeech for text-to-speech,”
H. Zen, V. Dang, R. Clark, et al., · 2019
Later among the works it cites.