Fetching the paper…
Reading the bibliography…
Any-to-any voice conversion (VC) aims to convert the timbre of utterances from and to any speakers seen or unseen during training.
J. Garofolo, L. Lamel, W. Fisher, J. Fiscus, D. Pallett, N. Dahlgren, and V. Zue, “TIMIT Acoustic-Phonetic Continuous Speech Corpus,” 1993. [Online]. Available: https://hdl.handle.net/11272.1/AB2/SWVENO
1993
Earlier work this paper cites.
C.-C. Lo, S.-W. Fu, W.-C. Huang, X. Wang, J. Yamagishi, Y. Tsao, and H.-M. Wang, “MOSNet: Deep Learning-Based Objective Assessment for Voice Conversion,” in Proc. Interspeech 2019 , 2019, pp. 1541–1545. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2019-2003
2003
Earlier work this paper cites.
J. Kominek and A. W. Black, “The cmu arctic speech databases,” in Fifth ISCA workshop on speech synthesis , 2004
2004
Earlier work this paper cites.
L. Sun, K. Li, H. Wang, S. Kang, and H. Meng, “Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,” in 2016 IEEE International Conference on Multimedia and Expo (ICME) , 2016, pp. 1–6
2016
Earlier work this paper cites.
C. Veaux, J. Yamagishi, K. MacDonald et al. , “Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,” 2017
2017
Earlier work this paper cites.
K. Ito and L. Johnson, “The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/, 2017
2017
Earlier work this paper cites.
S. Gidaris, P. Singh, and N. Komodakis, “Unsupervised representation learning by predicting image rotations,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings . OpenReview.net, 2018. [Online]. Available: https://openreview.net/forum?id=S1v4N2l0-
2018
Earlier work this paper cites.
A. van den Oord, Y. Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” 2018
2018
Earlier work this paper cites.
L. Wan, Q. Wang, A. Papir, and I. L. Moreno, “Generalized end-to-end loss for speaker verification,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 4879–4883
2018
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5329–5333
2018
Earlier work this paper cites.
S. Liu, J. Zhong, L. Sun, X. Wu, X. Liu, and H. Meng, “Voice conversion across arbitrary speakers based on a single target-speaker utterance,” in Proc. Interspeech 2018 , 2018, pp. 496–500. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2018-1504
2018
Earlier work this paper cites.
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, z. Chen, P. Nguyen, R. Pang, I. Lopez Moreno, and Y. Wu, “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in Advances in Neural Information Processing Systems 31 . Curran Associates, Inc., 2018, pp. 4480–4490. [Online]. Available: http://papers.nips.cc/paper/7700-transfer-learning-from-speaker-verification-to-multispeaker-text-to-speech-synthesis.pdf
2018
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 4171–4186. [Online]. Available: https://www.aclweb.org/anthology/N19-1423
2019
Cited alongside, same era.
J. chieh Chou and H.-Y. Lee, “One-Shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization,” in Proc. Interspeech 2019 , 2019, pp. 664–668. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2019-2663
Z. Yi, W.-C. Huang, X. Tian, J. Yamagishi, R. K. Das, T. Kinnunen, Z.-H. Ling, and T. Toda, “Voice Conversion Challenge 2020 –- Intra-lingual semi-parallel and cross-lingual voice conversion –-,” in Proc. Joint Workshop for the Blizzard Challenge and Voice Conversion Challenge 2020 , 2020, pp. 80–98. [Online]. Available: http://dx.doi.org/10.21437/VCC_BC.2020-14
2020
Later among the works it cites.
D.-Y. Wu, Y.-H. Chen, and H. yi Lee, “VQVC+: One-Shot Voice Conversion by Vector Quantization and U-Net Architecture,” in Proc. Interspeech 2020 , 2020, pp. 4691–4695. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-1443
2020
Later among the works it cites.
M. Rivière, A. Joulin, P. E. Mazaré, and E. Dupoux, “Unsupervised pretraining transfers well across languages,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 7414–7418
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “Autovc: Zero-shot voice style transfer with only autoencoder loss,” in Proceedings of the 36th International Conference on Machine Learning , 2019, pp. 5210–5219. [Online]. Available: http://proceedings.mlr.press/v97/qian19c.html
2019
Cited alongside, same era.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An Unsupervised Autoregressive Model for Speech Representation Learning,” in Proc. Interspeech 2019 , 2019, pp. 146–150. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2019-1473
2019
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=Bkg6RiCqY7
2019
Cited alongside, same era.
J. Lorenzo-Trueba, T. Drugman, J. Latorre, T. Merritt, B. Putrycz, R. Barra-Chicote, A. Moinet, and V. Aggarwal, “Towards Achieving Robust Universal Neural Vocoding,” in Proc. Interspeech 2019 , 2019, pp. 181–185. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2019-1424
2019
Cited alongside, same era.
Y. A. Chung and J. Glass, “Generative pre-training for speech with autoregressive predictive coding,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 3497–3501
2020
Cited alongside, same era.
A. T. Liu, S. w. Yang, P. H. Chi, P. c. Hsu, and H. y. Lee, “Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 6419–6423
2020
Cited alongside, same era.
S. won Park, D. young Kim, and M. chul Joe, “Cotatron: Transcription-Guided Speech Encoder for Any-to-Many Voice Conversion Without Parallel Data,” in Proc. Interspeech 2020 , 2020, pp. 4696–4700. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-1542
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual , H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., 2020
2020
Later among the works it cites.
P. Safari, M. India, and J. Hernando, “Self-Attention Encoding and Pooling for Speaker Recognition,” in Proc. Interspeech 2020 , 2020, pp. 941–945. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-1446
2020
Later among the works it cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Proc. Interspeech 2020 , 2020, pp. 5036–5040. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-3015
2020
Later among the works it cites.
A. T. Liu and Y. Shu-wen, “S3prl: The self-supervised speech pre-training and representation learning toolkit,” 2020. [Online]. Available: https://github.com/s3prl/s3prl
2020
Later among the works it cites.
W.-C. Huang, Y.-C. Wu, T. Hayashi, and T. Toda, “Any-to-one sequence-to-sequence voice conversion using self-supervised discrete speech representations,” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021
2021
Closest in time.
Y. Y. Lin, C.-M. Chien, J.-H. Lin, H. yi Lee, and L. shan Lee, “Fragmentvc: Any-to-any voice conversion by end-to-end extracting and fusing fine-grained voice fragments with attention,” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021
2021
Closest in time.
T. h. Huang, J. h. Lin, and H. y. Lee, “How far are we from robust voice conversion: A survey,” in 2021 IEEE Spoken Language Technology Workshop (SLT) , 2021, pp. 514–521
2021
Closest in time.