Fetching the paper…
Reading the bibliography…
This paper proposes Scyclone, a high-quality voice conversion (VC) technique without parallel data training.
H. Kawai, T. Toda, J. Ni, M. Tsuzaki, and K. Tokuda, “XIMERA: A new TTS from ATR based on corpus-based technologies,” in Proc. ISCA workshop on speech synthesis (SSW) , 2004, pp. 179–184
2004
Earlier work this paper cites.
2005
Earlier work this paper cites.
T. J. Hazen, W. Shen, and C. White, “Query-by-example spoken term detection using phonetic posteriorgram templates,” in Proc. ASRU , 2009, pp. 421–426
2009
Earlier work this paper cites.
A. L. Maas, A. Y. Hannun, and A. Y. Ng, “Rectifier nonlinearities improve neural network acoustic models,” in Proc. ICML , vol. 30, no. 1, 2013, p. 3
2013
Earlier work this paper cites.
M. Lin, Q. Chen, and S. Yan, “Network in network,” arXiv preprint arXiv:1312.4400 , 2013
2013
Earlier work this paper cites.
P. K. Diederik, M. Welling et al. , “Auto-encoding variational Bayes,” in Proc. ICLR , 2014, pp. 1–14
2014
Earlier work this paper cites.
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder-decoder for statistical machine translation,” in Proc. EMNLP , 2014, pp. 1724–1734
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
C.-C. Hsu, H.-T. Hwang, Y.-C. Wu, Y. Tsao, and H.-M. Wang, “Voice conversion from non-parallel corpora using variational auto-encoder,” in Proc. APSIPA , 2016, pp. 1–6
2016
Earlier work this paper cites.
L. Sun, K. Li, H. Wang, S. Kang, and H. Meng, “Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,” in Proc. ICME , 2016, pp. 1–6
2016
Cited alongside, same era.
M. Morise, F. Yokomori, and K. Ozawa, “World: a vocoder-based high-quality speech synthesis system for real-time applications,” IEICE Trans. Inf. Syst. , vol. 99, no. 7, pp. 1877–1884, 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
S. H. Mohammadi and A. Kain, “An overview of voice conversion systems,” Speech Communication , vol. 88, pp. 65–82, 2017
2017
Cited alongside, same era.
J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in Proc. ICCV , 2017, pp. 2223–2232
Y. Saito, Y. Ijima, K. Nishida, and S. Takamichi, “Non-parallel voice conversion using variational autoencoders conditioned by phonetic posteriorgrams and d-vectors,” in Proc. ICASSP , 2018, pp. 5274–5278
2018
Later among the works it cites.
T. Kaneko and H. Kameoka, “CycleGAN-VC: Non-parallel voice conversion using cycle-consistent adversarial networks,” in Proc. EUSIPCO , 2018, pp. 2100–2104
2018
Later among the works it cites.
2018
Later among the works it cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,” in Proc. ICASSP , 2018, pp. 4779–4783
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
J. Zhao, M. Mathieu, and Y. LeCun, “Energy-based generative adversarial network,” in Proc. ICLR , 2017, pp. 1–17
2017
Cited alongside, same era.
A. Martin and L. Bottou, “Towards principled methods for training generative adversarial networks,” in Proc. ICLR , 2017, pp. 1–17
2017
Cited alongside, same era.
A. Tamamori, T. Hayashi, K. Kobayashi, K. Takeda, and T. Toda, “Speaker-dependent WaveNet vocoder,” in Proc. INTERSPEECH , 2017, pp. 1118–1122
2017
Cited alongside, same era.
F. Fang, J. Yamagishi, I. Echizen, and J. Lorenzo-Trueba, “High-quality nonparallel voice conversion based on cycle-consistent adversarial network,” in Proc. ICASSP , 2018, pp. 5279–5283
2018
Cited alongside, same era.
2018
Later among the works it cites.
T. Kaneko, H. Kameoka, K. Tanaka, and N. Hojo, “CycleGAN-VC2: Improved CycleGAN-based non-parallel voice conversion,” in Proc. ICASSP , 2019, pp. 6820–6824
2019
Later among the works it cites.
W. Ping, K. Peng, and J. Chen, “ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech,” in Proc. ICLR , 2019, pp. 1–15
2019
Later among the works it cites.