Fetching the paper…
Reading the bibliography…
Voice conversion models have developed for decades, and current mainstream research focuses on non-streaming voice conversion.
M. Abe, S. Nakamura, K. Shikano, and H. Kuwabara, “Voice conversion through vector quantization,” Journal of the Acoustical Society of Japan (E) , vol. 11, no. 2, pp. 71–76, 1990
1990
Earlier work this paper cites.
K. Shikano, S. Nakamura, and M. Abe, “Speaker adaptation and voice conversion by codebook mapping,” in 1991 IEEE International Symposium on Circuits and Systems (ISCAS) . IEEE, 1991, pp. 594–597
1991
Earlier work this paper cites.
T. Toda, A. W. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 15, no. 8, pp. 2222–2235, 2007
2007
Earlier work this paper cites.
H. Zen, Y. Nankaku, and K. Tokuda, “Probabilistic feature mapping based on trajectory hmms,” in Ninth Annual Conference of the International Speech Communication Association , 2008
2008
Earlier work this paper cites.
E. Helander, J. Schwarz, J. Nurminen, H. Silen, and M. Gabbouj, “On the impact of alignment on voice conversion performance,” in Ninth Annual Conference of the International Speech Communication Association , 2008
2008
Earlier work this paper cites.
E. Helander, T. Virtanen, J. Nurminen, and M. Gabbouj, “Voice conversion using partial least squares regression,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 18, no. 5, pp. 912–921, 2010
2010
Earlier work this paper cites.
E. Helander, H. Silén, T. Virtanen, and M. Gabbouj, “Voice conversion using dynamic kernel partial least squares regression,” IEEE transactions on audio, speech, and language processing , vol. 20, no. 3, pp. 806–817, 2011
2011
Earlier work this paper cites.
L. Sun, K. Li, H. Wang, S. Kang, and H. Meng, “Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,” in 2016 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 2016, pp. 1–6
2016
Earlier work this paper cites.
T. Kaneko and H. Kameoka, “Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks,” in 2018 26th European Signal Processing Conference (EUSIPCO) . IEEE, 2018, pp. 2100–2104
2018
Earlier work this paper cites.
H. Kameoka, T. Kaneko, K. Tanaka, and N. Hojo, “Stargan-vc: Non-parallel many-to-many voice conversion using star generative adversarial networks,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 266–273
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Kaneko, H. Kameoka, K. Tanaka, and N. Hojo, “Cyclegan-vc2: Improved cyclegan-based non-parallel voice conversion,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6820–6824
2019
Cited alongside, same era.
2019
Cited alongside, same era.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “Autovc: Zero-shot voice style transfer with only autoencoder loss,” in International Conference on Machine Learning . PMLR, 2019, pp. 5210–5219
2019
Cited alongside, same era.
Y. Zhou, X. Tian, H. Xu, R. K. Das, and H. Li, “Cross-lingual voice conversion with bilingual phonetic posteriorgram and average modeling,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6790–6794
T. Saeki, Y. Saito, S. Takamichi, and H. Saruwatari, “Real-time, full-band, online dnn-based voice conversion system using a single cpu.” in INTERSPEECH , 2020, pp. 1021–1022
2020
Later among the works it cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in Neural Information Processing Systems , vol. 33, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Cited alongside, same era.
B. Sisman, J. Yamagishi, S. King, and H. Li, “An overview of voice conversion and its challenges: From statistical modeling to deep learning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2020
2020
Cited alongside, same era.
D.-Y. Wu and H.-y. Lee, “One-shot voice conversion by vector quantization,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7734–7738
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Later among the works it cites.
Z. Chen and P. Zhang, “TVQVC: Transformer Based Vector Quantized Variational Autoencoder with CTC Loss for Voice Conversion,” in Proc. Interspeech 2021 , 2021, pp. 826–830
2021
Later among the works it cites.
2021
Later among the works it cites.
G. Yang, S. Yang, K. Liu, P. Fang, W. Chen, and L. Xie, “Multi-band melgan: Faster waveform generation for high-quality text-to-speech,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 492–498
2021
Later among the works it cites.
J. Pons, S. Pascual, G. Cengarle, and J. Serrà, “Upsampling artifacts in neural audio synthesis,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 3005–3009
2021
Later among the works it cites.
H. Miao, G. Cheng, and P. Zhang, “Low-latency transformer model for streaming automatic speech recognition,” Electronics Letters , vol. 58, no. 1, pp. 44–46, 2022
2022
Closest in time.