Fetching the paper…
Reading the bibliography…
The quest for comprehensive generative models of intonation that link linguistic and paralinguistic functions to prosodic forms has been a longstanding challenge of speech communication research.
A generative model for the prosody of connected speech in japanese
Fujisaki, H., 1971 · 1971
Earlier work this paper cites.
Clichés mélodiques
Fónagy, I., Bérard, E., Fónagy, J., 1983 · 1983
Earlier work this paper cites.
A generative model of intonation, in: Prosody: Models and measurements. Springer, pp. 11–25
Gårding, E., 1983 · 1983
Earlier work this paper cites.
Recognition-by-components: a theory of human image understanding
Biederman, I., 1987 · 1987
Earlier work this paper cites.
Principal components analysis of images via back propagation, in: Visual Communications and Image Processing’88: Third in a Series, International Society for Optics and Photonics. pp. 1070–1078
Cottrell, G.W., Munro, P., 1988 · 1988
Earlier work this paper cites.
Adaptive mixtures of local experts
Jacobs, R.A., Jordan, M.I., Nowlan, S.J., Hinton, G.E., 1991 · 1991
Earlier work this paper cites.
Syllable-based segmental duration
Campbell, W.N., 1992 · 1992
Earlier work this paper cites.
Tobi: A standard for labeling english prosody, in: Second international conference on spoken language processing
Silverman, K., Beckman, M., Pitrelli, J., Ostendorf, M., Wightman, C., Price, P., Pierrehumbert, J., Hirschberg, J., 1992 · 1992
Earlier work this paper cites.
Génération multiparamétrique de la prosodie du français par apprentissage automatique
Morlec, Y., 1997 · 1997
Earlier work this paper cites.
Contextual tonal variations in mandarin
Xu, Y., 1997 · 1997
Earlier work this paper cites.
Evaluating the adequacy of synthetic prosody in signalling syntactic boundaries: methodology and first results, in: Proceedings of the first International Conference on Language Resources and Evaluation. Granada, Spain, pp. 647–650
Morlec, Y., Rilliard, A., Bailly, G., Aubergé, V., 1998 · 1998
Earlier work this paper cites.
Effects of tone and focus on the formation and alignment of f0contours
Xu, Y., 1999 · 1999
Earlier work this paper cites.
Simultaneous modeling of spectrum, pitch and duration in hmm-based speech synthesis, in: Sixth European Conference on Speech Communication and Technology
Yoshimura, T., Tokuda, K., Masuko, T., Kobayashi, T., Kitamura, T., 1999 · 1999
Earlier work this paper cites.
Generating prosodic attitudes in french: data, model and evaluation
Morlec, Y., Bailly, G., Aubergé, V., 2001 · 2001
Earlier work this paper cites.
Learning the hidden structure of speech: from communicative functions to prosody
Bailly, G., Holm, B., 2002 · 2002
Earlier work this paper cites.
Mind reading: The interactive guide to emotions
Baron-Cohen, S., Golan, O., Wheelwright, S., Hill, J., 2004 · 2004
Earlier work this paper cites.
La prosodie de la focalisation en français: faits perceptifs et morphogénétiques
Brichet, C., Aubergé, V., 2004 · 2004
Earlier work this paper cites.
A superposed prosodic model for chinese text-to-speech synthesis, in: Chinese Spoken Language Processing, 2004 International Symposium on, IEEE. pp. 177–180
Chen, G.P., Bailly, G., Liu, Q.F., Wang, R.H., 2004 · 2004
Earlier work this paper cites.
SFC: a trainable prosodic model
Bailly, G., Holm, B., 2005 · 2005
Earlier work this paper cites.
Form and function in the representation of speech prosody
Hirst, D.J., 2005 · 2005
Cited alongside, same era.
Parallel encoding of focus and interrogative meaning in mandarin intonation
Liu, F., Xu, Y., 2005 · 2005
Cited alongside, same era.
Speech melody as articulatorily implemented communicative functions
Xu, Y., 2005 · 2005
Cited alongside, same era.
Binary coding of speech spectrograms using a deep auto-encoder, in: Eleventh Annual Conference of the International Speech Communication Association
Deng, L., Seltzer, M.L., Yu, D., Acero, A., Mohamed, A.r., Hinton, G., 2010 · 2010
Cited alongside, same era.
Stylization and trajectory modelling of short and long term speech prosody variations, in: Interspeech
Obin, N., Lacheret, A., Rodet, X., 2011 · 2011
Cited alongside, same era.
Scikit-learn: Machine learning in Python
Acoustic modeling in statistical parametric speech synthesis–from hmm to lstm-rnn
Zen, H., 2015 · 2015
Later among the works it cites.
Chen, X., Kingma, D.P., Salimans, T., Duan, Y., Dhariwal, P., Schulman, J., Sutskever, I., Abbeel, P., 2016 · 2016
Later among the works it cites.
Deep learning. volume 1
Goodfellow, I., Bengio, Y., Courville, A., Bengio, Y., 2016 · 2016
Later among the works it cites.
Toward a bestiary of english intonational contours
Goodhue, D., Harrison, L., Su, Y.C., Wagner, M., 2016 · 2016
Later among the works it cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., Lerchner, A., 2016 · 2016
Later among the works it cites.
Merlin: An open source neural network speech synthesis system
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., Duchesnay, E., 2011 · 2011
Cited alongside, same era.
Semi-supervised recursive autoencoders for predicting sentiment distributions, in: Proceedings of the conference on empirical methods in natural language processing, Association for Computational Linguistics. pp. 151–161
Socher, R., Pennington, J., Huang, E.H., Ng, A.Y., Manning, C.D., 2011 · 2011
Cited alongside, same era.
Introduction to robust estimation and hypothesis testing
Wilcox, R.R., 2011 · 2011
Cited alongside, same era.
Auto-encoding variational bayes
Kingma, D.P., Welling, M., 2013 · 2013
Cited alongside, same era.
Speech enhancement based on deep denoising autoencoder., in: Interspeech, pp. 436–440
Lu, X., Tsao, Y., Matsuda, S., Hori, C., 2013 · 2013
Cited alongside, same era.
Statistical parametric speech synthesis using deep neural networks, in: Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, IEEE. pp. 7962–7966
Zen, H., Senior, A., Schuster, M., 2013 · 2013
Cited alongside, same era.
An autoencoder approach to learning bilingual word representations, in: Advances in Neural Information Processing Systems, pp. 1853–1861
Ap, S.C., Lauly, S., Larochelle, H., Khapra, M., Ravindran, B., Raykar, V.C., Saha, A., 2014 · 2014
Cited alongside, same era.
Wu, Z., Watts, O., King, S., 2016 · 2016
Later among the works it cites.
Modeling f0 trajectories in hierarchically structured deep neural networks
Yin, X., Lei, M., Qian, Y., Soong, F.K., He, L., Ling, Z.H., Dai, L.R., 2016 · 2016
Later among the works it cites.
Deep voice: Real-time neural text-to-speech
Arik, S.O., Chrzanowski, M., Coates, A., Diamos, G., Gibiansky, A., Kang, Y., Li, X., Miller, J., Ng, A., Raiman, J., et al., 2017 · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., Lerer, A., 2017 · 2017
Later among the works it cites.
Char2wav: End-to-end speech synthesis
Sotelo, J., Mehri, S., Kumar, K., Santos, J.F., Kastner, K., Courville, A., Bengio, Y., 2017 · 2017
Later among the works it cites.
Infovae: Information maximizing variational autoencoders
Zhao, S., Song, J., Ermon, S., 2017 · 2017
Later among the works it cites.
Hierarchical generative modeling for controllable speech synthesis
Hsu, W.N., Zhang, Y., Weiss, R.J., Zen, H., Wu, Y., Wang, Y., Cao, Y., Jia, Y., Chen, Z., Shen, J., et al., 2018 · 2018
Closest in time.
Sparse coding of pitch contours with deep auto-encoders, in: Speech Prosody
Obin, N., Beliao, J., 2018 · 2018
Closest in time.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions, in: Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, IEEE. pp. 4779–4783
Shen, J., Pang, R., Weiss, R.J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Skerrv-Ryan, R., et al., 2018 · 2018
Closest in time.
Towards end-to-end prosody transfer for expressive speech synthesis with tacotron, in: Dy, J., Krause, A. (Eds.), Proceedings of the 35th International Conference on Machine Learning, PMLR, Stockholmsmässan, Stockholm Sweden. pp. 4693–4702
Skerry-Ryan, R., Battenberg, E., Xiao, Y., Wang, Y., Stanton, D., Shor, J., Weiss, R., Clark, R., Saurous, R.A., 2018 · 2018
Closest in time.
Voiceloop: Voice fitting and synthesis via a phonological loop
Taigman, Y., Wolf, L., Polyak, A., Nachmani, E., 2018 · 2018
Closest in time.
Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
Wang, Y., Stanton, D., Zhang, Y., Skerry-Ryan, R., Battenberg, E., Shor, J., Xiao, Y., Ren, F., Jia, Y., Saurous, R.A., 2018 · 2018
Closest in time.
A generative adversarial network for style modeling in a text-to-speech system, in: International Conference on Learning Representations
Ma, S., Mcduff, D., Song, Y., 2019 · 2019
Closest in time.