Fetching the paper…
Reading the bibliography…
Emotional voice conversion (EVC) aims to change the emotional state of an utterance while preserving the linguistic content and speaker identity.
D. Griffin and J. Lim, “Signal estimation from modified short-time fourier transform,”
1984
Earlier work this paper cites.
J. Tao, Y. Kang, and A. Li, “Prosody conversion from neutral speech to emotional speech,”
2006
Earlier work this paper cites.
T. Toda, A. W. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,”
2007
Earlier work this paper cites.
L. v. d. Maaten and G. Hinton, “Visualizing data using t-sne,”
2008
Earlier work this paper cites.
Y. Xu, “Speech prosody: A methodological review,”
2011
Earlier work this paper cites.
L.-H. Chen, Z.-H. Ling, L.-J. Liu, and L.-R. Dai, “Voice conversion using deep neural networks with layer-wise generative training,”
2014
Earlier work this paper cites.
R. Aihara, R. Ueda, T. Takiguchi, and Y. Ariki, “Exemplar-based emotional voice conversion using non-negative matrix factorization,” in
2014
Earlier work this paper cites.
D. Bahdanau, K. H. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in
2015
Earlier work this paper cites.
H. Ming, D. Huang, L. Xie, J. Wu, M. Dong, and H. Li, “Deep bidirectional lstm modeling of timbre and prosody for emotional voice conversion,”
2016
Earlier work this paper cites.
C. Veaux, J. Yamagishi, K. MacDonald
2016
Earlier work this paper cites.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. K. J. F. S. Kyle, K. A. C. Y. B. Jose, and S. M. Sotelo, “Char2wav: End-to-end speech synthesis,” in
2017
Earlier work this paper cites.
B. Sisman, H. Li, and K. C. Tan, “Sparse representation of phonetic features for voice conversion with and without parallel data,” in
2017
Cited alongside, same era.
D. Schuller and B. W. Schuller, “The age of artificial emotional intelligence,”
2018
Cited alongside, same era.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in
2018
Cited alongside, same era.
B. Sisman, M. Zhang, and H. Li, “Group sparse representation with wavenet vocoder adaptation for spectrum and prosody conversion,”
2019
Cited alongside, same era.
B. Sisman, M. Zhang, M. Dong, and H. Li, “On the study of generative adversarial networks for cross-lingual voice conversion,”
2019
Cited alongside, same era.
2020
Later among the works it cites.
B. Sisman, J. Yamagishi, S. King, and H. Li, “An overview of voice conversion and its challenges: From statistical modeling to deep learning,”
2020
Later among the works it cites.
K. Zhou, B. Sisman, and H. Li, “Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data,” in
2020
Later among the works it cites.
G. Rizos, A. Baird, M. Elliott, and B. Schuller, “Stargan for emotional speech conversion: Validated by data augmentation of end-to-end emotion recognition,” in
2020
Later among the works it cites.
R. Shankar, J. Sager, and A. Venkataraman, “Non-parallel emotion conversion using a deep-generative hybrid network and an adversarial pair discriminator,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Luo, J. Chen, T. Takiguchi, and Y. Ariki, “Emotional voice conversion using dual supervised adversarial networks with continuous wavelet transform f0 features,”
2019
Cited alongside, same era.
J. Gao, D. Chakraborty, H. Tembine, and O. Olaleye, “Nonparallel emotional speech conversion,”
2019
Cited alongside, same era.
J.-X. Zhang, Z.-H. Ling, L.-J. Liu, Y. Jiang, and L.-R. Dai, “Sequence-to-sequence acoustic modeling for voice conversion,”
2019
Cited alongside, same era.
K. Tanaka, H. Kameoka, T. Kaneko, and N. Hojo, “Atts2s-vc: Sequence-to-sequence voice conversion with attention and context preservation mechanisms,” in
2019
Cited alongside, same era.
J.-X. Zhang, Z.-H. Ling, Y. Jiang, L.-J. Liu, C. Liang, and L.-R. Dai, “Improving sequence-to-sequence voice conversion by adding text-supervision,” in
2019
Cited alongside, same era.
M. Zhang, X. Wang, F. Fang, H. Li, and J. Yamagishi, “Joint training framework for text-to-speech and voice conversion using multi-source tacotron and wavenet,”
2019
Cited alongside, same era.
J.-X. Zhang, Z.-H. Ling, and L.-R. Dai, “Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations,”
2019
Cited alongside, same era.
2020
Later among the works it cites.
K. Zhou, B. Sisman, and H. Li, “Vaw-gan for disentanglement and recomposition of emotional elements in speech,”
2020
Later among the works it cites.
K. Zhou, B. Sisman, M. Zhang, and H. Li, “Converting Anyone’s Emotion: Towards Speaker-Independent Emotional Voice Conversion,”
2020
Later among the works it cites.
H. Kameoka, K. Tanaka, D. Kwaśny, T. Kaneko, and N. Hojo, “Convs2s-vc: Fully convolutional sequence-to-sequence voice conversion,”
2020
Later among the works it cites.
T.-H. Kim, S. Cho, S. Choi, S. Park, and S.-Y. Lee, “Emotional voice conversion using multitask learning with text-to-speech,” in
2020
Later among the works it cites.
D. M. Schuller and B. W. Schuller, “A review on five recent and near-future developments in computational processing of emotion in the human voice,”
2020
Later among the works it cites.
K. Zhou, B. Sisman, R. Liu, and H. Li, “Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,”
2021
Closest in time.