Fetching the paper…
Reading the bibliography…
As part of the Human-Computer Interaction field, Expressive speech synthesis is a very rich domain as it requires knowledge in areas such as machine learning, signal processing, sociology, psychology.
A circumplex model of affect
Russell JA · 1980
Earlier work this paper cites.
An argument for basic emotions
Ekman P · 1992
Earlier work this paper cites.
Verification of acoustical correlates of emotional speech using formant-synthesis
Burkhardt F, Sendlmeier WF · 2000
Earlier work this paper cites.
Emotional speech synthesis: A review
Schröder M · 2001
Earlier work this paper cites.
Information theory, inference and learning algorithms
MacKay DJ, Mac Kay DJ · 2003
Earlier work this paper cites.
Speech probability distribution
Gazor S, Zhang W · 2003
Earlier work this paper cites.
Traitement du signal
Dutoit T · 2005
Earlier work this paper cites.
Statistical parametric speech synthesis
Zen H, Tokuda K, Black AW · 2009
Earlier work this paper cites.
Learning deep physiological models of affect
Martinez HP, Bengio Y, Yannakakis GN · 2013
Cited alongside, same era.
Statistical parametric speech synthesis using deep neural networks
Zen H, Senior A, Schuster M · 2013
Cited alongside, same era.
Emotional speech synthesis
Burkhardt F, Campbell N · 2014
Cited alongside, same era.
Adieu features? end-to-end speech emotion recognition using a deep convolutional recurrent network
Trigeorgis G, Ringeval F, Brueckner R, Marchi E, Nicolaou MA, Schuller B, et al · 2016
Cited alongside, same era.
WaveNet: A Generative Model for Raw Audio
van den Oord A, Dieleman S, Zen H, Simonyan K, Vinyals O, Graves A, et al · 2016
Cited alongside, same era.
From HMMs to DNNs: where do the improvements come from?
Watts O, Henter GE, Merritt T, Wu Z, King S · 2016
Cited alongside, same era.
Deep learning
Goodfellow I, Bengio Y, Courville A · 2016
Later among the works it cites.
The ordinal nature of emotions
Yannakakis GN, Cowie R, Busso C · 2017
Later among the works it cites.
The theory of constructed emotion: an active inference account of interoception and categorization
Barrett LF · 2017
Later among the works it cites.
Tacotron: Towards End-to-End Speech Synthesis
Wang Y, Skerry-Ryan RJ, Stanton D, Wu Y, Weiss RJ, Jaitly N, et al · 2017
Later among the works it cites.
A Comparison of Sequence-to-Sequence Models for Speech Recognition
Prabhavalkar R, Rao K, Sainath TN, Li B, Johnson L, Jaitly N · 2017
Later among the works it cites.
Probabilistic Modeling of Speech in Spectral Domain using Maximum Likelihood Estimation
Usman M, Zubair M, Shiblee M, Rodrigues P, Jaffar S · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Merlin: An open source neural network speech synthesis system
Wu Z, Watts O, King S · 2016
Cited alongside, same era.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
Chan W, Jaitly N, Le Q, Vinyals O · 2016
Cited alongside, same era.
Tits N, Wang F, Haddad KE, Pagel V, Dutoit T · 2019
Closest in time.
Exploring Transfer Learning for Low Resource Emotional TTS
Tits N, El Haddad K, Dutoit T · 2020
Closest in time.