Fetching the paper…
Reading the bibliography…
We propose a novel training strategy for Tacotron-based text-to-speech (TTS) system to improve the expressiveness of speech.
B. Sisman, M. Zhang, and H. Li, “A voice conversion framework with tandem feature sparse representation and speaker-adapted wavenet vocoder,” in
1982
Earlier work this paper cites.
D. Griffin and J. Lim, “Signal estimation from modified short-time fourier transform,”
1984
Earlier work this paper cites.
I. R. Murray and J. L. Arnott, “Toward the simulation of emotion in synthetic speech: A review of the literature on human vocal emotion,”
1993
Earlier work this paper cites.
R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” in
1993
Earlier work this paper cites.
A. J. Hunt and A. W. Black, “Unit selection in a concatenative speech synthesis system using a large speech database,” in
1996
Earlier work this paper cites.
K. Chen, B. Chen, J. Lai, and K. Yu, “High-quality voice conversion using spectrogram-based wavenet vocoder,” in
1997
Earlier work this paper cites.
P. Taylor and A. W. Black, “Assigning phrase breaks from part-of-speech sequences,”
1998
Earlier work this paper cites.
T. Thiede, W. C. Treurniet, R. Bitto, C. Schmidmer, T. Sporer, J. G. Beerends, and C. Colomes, “Peaq-the itu standard for objective measurement of perceived audio quality,”
2000
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in
2001
Earlier work this paper cites.
A. Wennerstrom, “the music of everyday speech prosody and discourse analysis,”
2001
Earlier work this paper cites.
K. Tokuda, H. Zen, and A. W. Black, “An hmm-based speech synthesis system applied to english,” in
2002
Earlier work this paper cites.
J. Yamagishi, K. Onishi, T. Masuko, and T. Kobayashi, “Modeling of various speaking styles and emotions for hmm-based speech synthesis,” in
2003
Earlier work this paper cites.
O. Pierre-Yves, “The production and recognition of emotions in speech: features and algorithms,”
2003
Earlier work this paper cites.
J. Hirschberg, “Pragmatics and intonation,”
2004
Earlier work this paper cites.
M. Tachibana, J. Yamagishi, K. Onishi, T. Masuko, and T. Kobayashi, “Hmm-based speech synthesis with various speaking styles using model interpolation,” in
2004
Earlier work this paper cites.
D. R. Ladd, “Intonational phonology,”
2008
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,”
2008
Earlier work this paper cites.
L. v. d. Maaten and G. Hinton, “Visualizing data using t-sne,”
2008
Earlier work this paper cites.
V. Emiya, E. Vincent, N. Harlander, and V. Hohmann, “The peass toolkit-perceptual evaluation methods for audio source separation,” 2010
2010
Earlier work this paper cites.
Z. Wu, T. Kinnunen, E. S. Chng, and H. Li, “Text-independent f0 transformation with non-parallel data for voice conversion,” in
2010
Earlier work this paper cites.
Y. XU, “Speech prosody: A methodological review,”
2011
Earlier work this paper cites.
D. Yu and M. L. Seltzer, “Improved bottleneck features using pretrained deep neural networks,” in
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in
2012
Earlier work this paper cites.
F. Eyben, S. Buchholz, N. Braunschweiler, J. Latorre, V. Wan, M. J. Gales, and K. Knill, “Unsupervised clustering of emotion and voice styles for expressive tts,” in
2012
Earlier work this paper cites.
K. Tokuda, Y. Nankaku, T. Toda, H. Zen, J. Yamagishi, and K. Oura, “Speech synthesis based on hidden markov models,”
2013
Earlier work this paper cites.
H. Zen, A. Senior, and M. Schuster, “Statistical parametric speech synthesis using deep neural networks,” in
2013
Earlier work this paper cites.
M. Vainio, A. Suni, D. Aalto
2013
Earlier work this paper cites.
A. Suni, D. Aalto, T. Raitio, P. Alku, and M. Vainio, “Wavelets for intonation modeling in HMM speech synthesis,”
2013
Earlier work this paper cites.
J. Latorre, “Multilevel parametric-base F0 model for speech synthesis,” 2014
2014
Earlier work this paper cites.
G. Sanchez, H. Silen, J. Nurminen, and M. Gabbouj, “Hierarchical modeling of F0 contours for voice conversion,” in
2014
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,”
2014
Earlier work this paper cites.
M. Müller, “Dynamic time warping,”
2014
Earlier work this paper cites.
M. S. Ribeiro and R. A. J. Clark, “A multi-level representation of f0 using the continuous wavelet transform and the Discrete Cosine Transform,” in
2015
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in
2015
Cited alongside, same era.
A. Z. Jusoh, R. Togneri, S. Nordholm, N. Sulaiman, and M. H. Khairolanuar, “The investigation of frame disturbance (fd) in perceptual evaluation speech quality (pesq) as a perceptual metric,”
2015
Cited alongside, same era.
A. Dosovitskiy and T. Brox, “Generating images with perceptual similarity metrics based on deep networks,” in
2016
Cited alongside, same era.
J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual losses for real-time style transfer and super-resolution,” in
2016
Cited alongside, same era.
S. Zhang, S. Zhang, T. Huang, and W. Gao, “Speech emotion recognition using deep convolutional neural network and discriminant temporal pyramid matching,”
2018
Later among the works it cites.
M. Chen, X. He, J. Yang, and H. Zhang, “3-d convolutional recurrent neural networks with attention model for speech emotion recognition,”
2018
Later among the works it cites.
Y. Lee and T. Kim, “Robust and fine-grained prosody control of end-to-end speech synthesis,” in
2019
Later among the works it cites.
Y.-A. Chung, Y. Wang, W.-N. Hsu, Y. Zhang, and R. Skerry-Ryan, “Semi-supervised training for improving data efficiency in end-to-end speech synthesis,” in
2019
Later among the works it cites.
M. He, Y. Deng, and L. He, “Robust Sequence-to-Sequence Acoustic Modeling with Stepwise Monotonic Attention for Neural TTS,” in
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
H. Ming, D. Huang, L. Xie, S. Zhang, M. Dong, and H. Li, “Exemplar-based sparse representation of timbre and prosody for voice conversion,” in
2016
Cited alongside, same era.
H. Ming, D. Huang, L. Xie, J. Wu, M. Dong, and H. Li, “Deep bidirectional LSTM modeling of timbre and prosody for emotional voice conversion,” in
2016
Cited alongside, same era.
2016
Cited alongside, same era.
I. Goodfellow, Y. Bengio, and A. Courville,
2016
Cited alongside, same era.
R. Liu, F. Bao, G. Gao, and Y. Wang, “Mongolian text-to-speech system based on deep neural network,” in
2017
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio
2017
Cited alongside, same era.
H.-T. Luong, X. Wang, J. Yamagishi, and N. Nishizawa, “Training Multi-Speaker Neural Text-to-Speech Systems Using Speaker-Imbalanced Speech Corpora,” in
2019
Later among the works it cites.
T. Okamoto, T. Toda, Y. Shiga, and H. Kawai, “Real-Time Neural Text-to-Speech with Sequence-to-Sequence Acoustic Model and WaveGlow or Single Gaussian WaveRNN Vocoders,” in
2019
Later among the works it cites.
——, “Group Sparse Representation with WaveNet Vocoder Adaptation for Spectrum and Prosody Conversion,”
2019
Later among the works it cites.
W.-C. Lin, Y. Tsao, F. Chen, and H.-M. Wang, “Investigation of neural network approaches for unified spectral and prosodic feature enhancement,” in
2019
Later among the works it cites.
K. Emir Ak, J. Hwee Lim, J. Yew Tham, and A. Kassim, “Semantically consistent hierarchical text to fashion image synthesis with an enhanced-attentional generative adversarial network,” in
2019
Later among the works it cites.
Y. Yasuda, X. Wang, S. Takaki, and J. Yamagishi, “Investigation of enhanced tacotron text-to-speech synthesis systems with self-attention for pitch accent language,” in
2019
Later among the works it cites.
T. Kenter, V. Wan, C.-A. Chan, R. Clark, and J. Vit, “Chive: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network,” in
2019
Later among the works it cites.
F. G. Germain, Q. Chen, and V. Koltun, “Speech Denoising with Deep Feature Losses,” in
2019
Later among the works it cites.
C.-C. Lo, S.-W. Fu, W.-C. Huang, X. Wang, J. Yamagishi, Y. Tsao, and H.-M. Wang, “Mosnet: Deep learning-based objective assessment for voice conversion,” in
2019
Later among the works it cites.
E. Kim and J. W. Shin, “Dnn-based emotion recognition based on bottleneck acoustic features and lexical features,” in
2019
Later among the works it cites.
R. Lotfian and C. Busso, “Curriculum learning for speech emotion recognition from crowdsourced labels,”
2019
Later among the works it cites.
P. Wu, Z. Ling, L. Liu, Y. Jiang, H. Wu, and L. Dai, “End-to-end emotional speech synthesis using style tokens and semi-supervised training,” in
2019
Later among the works it cites.
R. Liu, B. Sisman, F. Bao, G. Gao, and H. Li, “Wavetts: Tacotron-based tts with joint time-frequency domain loss,” in
2020
Closest in time.
R. Liu, B. Sisman, J. Li, F. Bao, G. Gao, and H. Li, “Teacher-student training for robust tacotron-based tts,” in
2020
Closest in time.
2020
Closest in time.
Z. Hodari, C. Lai, and S. King, “Perception of prosodic variation for speech synthesis using an unsupervised discrete representation of f0,” in
2020
Closest in time.
G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, and Y. Wu, “Fully-hierarchical fine-grained prosody modeling for interpretable speech synthesis,” in
2020
Closest in time.
G. Sun, Y. Zhang, R. J. Weiss, Y. Cao, H. Zen, A. Rosenberg, B. Ramabhadran, and Y. Wu, “Generating diverse and natural text-to-speech samples using a quantized fine-grained vae and autoregressive prosody prior,” in
2020
Closest in time.
A. Wright and V. Välimäki, “Perceptual loss function for neural modeling of audio systems,” in
2020
Closest in time.
B. Sisman, J. Yamagishi, S. King, and H. Li, “An overview of voice conversion and its challenges: From statistical modeling to deep learning,”
2020
Closest in time.
S. Kataria, P. S. Nidadavolu, J. Villalba, N. Chen, P. García-Perera, and N. Dehak, “Feature enhancement with deep feature losses for speaker verification,” in
2020
Closest in time.
M. Kawanaka, Y. Koizumi, R. Miyazaki, and K. Yatabe, “Stable training of dnn for speech enhancement based on perceptually-motivated black-box cost function,” in
2020
Closest in time.
K.-S. Lee, “Voice conversion using a perceptual criterion,”
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
S.-Y. Um, S. Oh, K. Byun, I. Jang, C. Ahn, and H.-G. Kang, “Emotional speech synthesis with rich and granularized control,” in
2020
Closest in time.