Fetching the paper…
Reading the bibliography…
Speech information can be roughly decomposed into four components: language content, timbre, pitch, and rhythm.
Distance measures for speech recognition, psychological and instrumental
Mermelstein, P · 1976
Earlier work this paper cites.
Predictive coding of speech signals and subjective error criteria
Atal, B. and Schroeder, M · 1979
Earlier work this paper cites.
Code-excited linear prediction (CELP): High-quality speech at very low bit rates
Schroeder, M. and Atal, B · 1985
Earlier work this paper cites.
Pitch targets and their realization: Evidence from mandarin chinese
Xu, Y. and Wang, Q. E · 2001
Earlier work this paper cites.
Discrete-time speech signal processing: principles and practice
Quatieri, T. F · 2006
Earlier work this paper cites.
TANDEM-STRAIGHT: A temporally stable power spectral representation for periodic signals and applications to interference-free spectrum, F0, and aperiodicity estimation
Kawahara, H., Morise, M., Takahashi, T., Nisimura, R., Irino, T., and Banno, H · 2008
Earlier work this paper cites.
A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments
Nakatani, T., Amano, S., Irino, T., Ishizuka, K., and Kondo, T · 2008
Earlier work this paper cites.
Reducing F0 frame error of F0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend
Chu, W. and Alwan, A · 2009
Earlier work this paper cites.
Emotional speech processing: Disentangling the effects of prosody and semantic cues
Pell, M. D., Jaywant, A., Monetta, L., and Kotz, S. A · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Improvement of probabilistic acoustic tube model for speech decomposition
Zhang, Y., Ou, Z., and Hasegawa-Johnson, M · 2014
Earlier work this paper cites.
Voice conversion from non-parallel corpora using variational auto-encoder
Hsu, C.-C., Hwang, H.-T., Wu, Y.-C., Tsao, Y., and Wang, H.-M · 2016
Earlier work this paper cites.
Superseded-CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit
Veaux, C., Yamagishi, J., MacDonald, K., et al · 2016
Earlier work this paper cites.
Voice conversion from unaligned corpora using variational autoencoding Wasserstein generative adversarial networks
Hsu, C.-C., Hwang, H.-T., Wu, Y.-C., Tsao, Y., and Wang, H.-M · 2017
Cited alongside, same era.
Parallel-data-free voice conversion using cycle-consistent adversarial networks
Kaneko, T. and Kameoka, H · 2017
Cited alongside, same era.
Fader networks: Manipulating images by sliding attributes
Lample, G., Zeghidour, N., Usunier, N., Bordes, A., Denoyer, L., and Ranzato, M · 2017
Cited alongside, same era.
StarGAN: Unified generative adversarial networks for multi-domain image-to-image translation
Choi, Y., Choi, M., Kim, M., Ha, J.-W., Kim, S., and Choo, J · 2018
Cited alongside, same era.
Multi-target voice conversion without parallel data by adversarially learning disentangled audio representations
Chou, J.-c., Yeh, C.-c., Lee, H.-y., and Lee, L.-s · 2018
Cited alongside, same era.
One-shot voice conversion by separating speaker and content representations with instance normalization
Chou, J.-c. and Lee, H.-Y · 2019
Later among the works it cites.
StarGAN-VC2: Rethinking conditional methods for StarGAN-based voice conversion
Kaneko, T., Kameoka, H., Tanaka, K., and Hojo, N · 2019
Later among the works it cites.
CHiVE: Varying prosody in speech synthesis with a linguistically driven dynamic hierarchical conditional variational network
Kenter, T., Wan, V., Chan, C.-A., Clark, R., and Vit, J · 2019
Later among the works it cites.
Unsupervised singing voice conversion
Nachmani, E. and Wolf, L · 2019
Later among the works it cites.
Attention-based WaveNet autoencoder for universal voice conversion
Polyak, A. and Wolf, L · 2019
Later among the works it cites.
AutoVC: Zero-shot voice style transfer with only autoencoder loss
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Voice impersonation using generative adversarial networks
Gao, Y., Singh, R., and Raj, B · 2018
Cited alongside, same era.
Voice conversion based on cross-domain features using variational auto encoders
Huang, W.-C., Hwang, H.-T., Peng, Y.-H., Tsao, Y., and Wang, H.-M · 2018
Cited alongside, same era.
StarGAN-VC: Non-parallel many-to-many voice conversion using star generative adversarial networks
Kameoka, H., Kaneko, T., Tanaka, K., and Hojo, N · 2018
Cited alongside, same era.
Statistical voice conversion based on WaveNet
Niwa, J., Yoshimura, T., Hashimoto, K., Oura, K., Nankaku, Y., and Tokuda, K · 2018
Cited alongside, same era.
Towards end-to-end prosody transfer for expressive speech synthesis with Tacotron
Skerry-Ryan, R., Battenberg, E., Xiao, Y., Wang, Y., Stanton, D., Shor, J., Weiss, R., Clark, R., and Saurous, R. A · 2018
Cited alongside, same era.
Group normalization
Wu, Y. and He, K · 2018
Cited alongside, same era.
Parrotron: An end-to-end speech-to-speech conversion model and its applications to hearing-impaired speech and speech separation
Biadsy, F., Weiss, R. J., Moreno, P. J., Kanvesky, D., and Jia, Y · 2019
Cited alongside, same era.
Qian, K., Zhang, Y., Chang, S., Yang, X., and Hasegawa-Johnson, M · 2019
Later among the works it cites.
Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion
Serrà, J., Pascual, S., and Perales, C. S · 2019
Later among the works it cites.
Sequence to sequence neural speech synthesis with prosody modification capabilities
Shechtman, S. and Sorin, A · 2019
Later among the works it cites.
r9y9/pysptk: 0.1. 14, 2019
Yamamoto, R., Felipe, J., and Blaauw, M · 2019
Later among the works it cites.
Unsupervised representation disentanglement using cross domain features and adversarial learning in variational autoencoder based voice conversion
Huang, W.-C., Luo, H., Hwang, H.-T., Lo, C.-C., Peng, Y.-H., Tsao, Y., and Wang, H.-M · 2020
Closest in time.
F0-consistent many-to-many non-parallel voice conversion via conditional autoencoder
Qian, K., Jin, Z., Hasegawa-Johnson, M., and Mysore, G. J · 2020
Closest in time.
Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens
Valle, R., Li, J., Prenger, R., and Catanzaro, B · 2020
Closest in time.