Fetching the paper…
Reading the bibliography…
Prosody plays an important role in characterizing the style of a speaker or an emotion, but most non-parallel voice or emotion style transfer algorithms do not convert any prosody information.
Transforming spectrum and prosody for emotional voice conversion with non-parallel training data
Zhou, K., Sisman, B., and Li, H · 2002
Earlier work this paper cites.
Prosody conversion from neutral speech to emotional speech
Tao, J., Kang, Y., and Li, A · 2006
Earlier work this paper cites.
Data-driven emotion conversion in spoken english
Inanoglu, Z. and Young, S · 2009
Earlier work this paper cites.
Hierarchical prosody conversion using regression-based clustering for emotional speech synthesis
Wu, C.-H., Hsia, C.-C., Lee, C.-H., and Lin, M.-C · 2009
Earlier work this paper cites.
Vaw-gan for disentanglement and recomposition of emotional elements in speech
Zhou, K., Sisman, B., and Li, H · 2011
Earlier work this paper cites.
Gmm-based emotional voice conversion using spectrum and prosody features
Aihara, R., Takashima, R., Takiguchi, T., and Ariki, Y · 2012
Earlier work this paper cites.
Voice conversion from non-parallel corpora using variational auto-encoder
Hsu, C.-C., Hwang, H.-T., Wu, Y.-C., Tsao, Y., and Wang, H.-M · 2016
Earlier work this paper cites.
Emotional voice conversion using deep neural networks with mcc and f0 features
Luo, Z., Takiguchi, T., and Ariki, Y · 2016
Earlier work this paper cites.
Deep bidirectional lstm modeling of timbre and prosody for emotional voice conversion
Ming, H., Huang, D., Xie, L., Wu, J., Dong, M., and Li, H · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Superseded-CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit, 2016
Veaux, C., Yamagishi, J., MacDonald, K., et al · 2016
Earlier work this paper cites.
Voice conversion from unaligned corpora using variational autoencoding Wasserstein generative adversarial networks
Hsu, C.-C., Hwang, H.-T., Wu, Y.-C., Tsao, Y., and Wang, H.-M · 2017
Earlier work this paper cites.
Parallel-data-free voice conversion using cycle-consistent adversarial networks
Kaneko, T. and Kameoka, H · 2017
Earlier work this paper cites.
Joint ctc-attention based end-to-end speech recognition using multi-task learning
Kim, S., Hori, T., and Watanabe, S · 2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Hybrid ctc/attention architecture for end-to-end speech recognition
Watanabe, S., Hori, T., Kim, S., Hershey, J. R., and Hayashi, T · 2017
Cited alongside, same era.
The emotional voices database: Towards controlling the emotion dimension in voice generation systems
Adigwe, A., Tits, N., Haddad, K. E., Ostadabbas, S., and Dutoit, T · 2018
Cited alongside, same era.
StarGAN: Unified generative adversarial networks for multi-domain image-to-image translation
Choi, Y., Choi, M., Kim, M., Ha, J.-W., Kim, S., and Choo, J · 2018
Cited alongside, same era.
Multi-target voice conversion without parallel data by adversarially learning disentangled audio representations
Chou, J.-C., Yeh, C.-C., Lee, H.-Y., and Lee, L.-S · 2018
ACVAE-VC: Non-parallel voice conversion with auxiliary classifier variational autoencoder
Kameoka, H., Kaneko, T., Tanaka, K., and Hojo, N · 2019
Later among the works it cites.
StarGAN-VC2: Rethinking conditional methods for StarGAN-based voice conversion
Kaneko, T., Kameoka, H., Tanaka, K., and Hojo, N · 2019
Later among the works it cites.
CHiVE: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network
Kenter, T., Wan, V., Chan, C.-A., Clark, R., and Vit, J · 2019
Later among the works it cites.
Unsupervised singing voice conversion
Nachmani, E. and Wolf, L · 2019
Later among the works it cites.
Attention-based wavenet autoencoder for universal voice conversion
Polyak, A. and Wolf, L · 2019
Later among the works it cites.
Autovc: Zero-shot voice style transfer with only autoencoder loss
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Voice impersonation using generative adversarial networks
Gao, Y., Singh, R., and Raj, B · 2018
Cited alongside, same era.
StarGAN-VC: Non-parallel many-to-many voice conversion using star generative adversarial networks
Kameoka, H., Kaneko, T., Tanaka, K., and Hojo, N · 2018
Cited alongside, same era.
Statistical voice conversion based on WaveNet
Niwa, J., Yoshimura, T., Hashimoto, K., Oura, K., Nankaku, Y., and Tokuda, K · 2018
Cited alongside, same era.
Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron
Skerry-Ryan, R., Battenberg, E., Xiao, Y., Wang, Y., Stanton, D., Shor, J., Weiss, R., Clark, R., and Saurous, R. A · 2018
Cited alongside, same era.
Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
Wang, Y., Stanton, D., Zhang, Y., Ryan, R.-S., Battenberg, E., Shor, J., Xiao, Y., Jia, Y., Ren, F., and Saurous, R. A · 2018
Cited alongside, same era.
Group normalization
Wu, Y. and He, K · 2018
Cited alongside, same era.
Parrotron: An end-to-end speech-to-speech conversion model and its applications to hearing-impaired speech and speech separation
Biadsy, F., Weiss, R. J., Moreno, P. J., Kanvesky, D., and Jia, Y · 2019
Cited alongside, same era.
Qian, K., Zhang, Y., Chang, S., Yang, X., and Hasegawa-Johnson, M · 2019
Later among the works it cites.
Blow: A single-scale hyperconditioned flow for non-parallel raw-audio voice conversion
Serrà, J., Pascual, S., and Perales, C. S · 2019
Later among the works it cites.
Self-Expressing Autoencoders for Unsupervised Spoken Term Discovery
Bhati, S., Villalba, J., Żelasko, P., and Dehak, N · 2020
Later among the works it cites.
Unsupervised representation disentanglement using cross domain features and adversarial learning in variational autoencoder based voice conversion
Huang, W.-C., Luo, H., Hwang, H.-T., Lo, C.-C., Peng, Y.-H., Tsao, Y., and Wang, H.-M · 2020
Later among the works it cites.
Towards unsupervised speech recognition and synthesis with quantized speech representation learning
Liu, A. H., Tu, T., Lee, H.-y., and Lee, L.-s · 2020
Later among the works it cites.
Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens
Valle, R., Li, J., Prenger, R., and Catanzaro, B · 2020
Later among the works it cites.
On layer normalization in the transformer architecture
Xiong, R., Yang, Y., He, D., Zheng, K., Zheng, S., Xing, C., Zhang, H., Lan, Y., Wang, L., and Liu, T · 2020
Later among the works it cites.
Unsupervised audiovisual synthesis via exemplar autoencoders
Deng, K., Bansal, A., and Ramanan, D · 2021
Closest in time.