Fetching the paper…
Reading the bibliography…
In this paper, a neural network named Sequence-to-sequence ConvErsion NeTwork (SCENT) is presented for acoustic modeling in voice conversion.
D. G. Childers, B. Yegnanarayana, and K. Wu, “Voice conversion: Factors responsible for quality,” in IEEE International Conference on Acoustics, Speech and Signal Processing , 1985, pp. 748–751
1985
Earlier work this paper cites.
D. M. Titterington, A. F. M. Smith, and U. E. Makov, Statistical analysis of finite mixture distributions . Wiley, 1985
1985
Earlier work this paper cites.
L.-J. Liu, Z.-H. Ling, and L.-R. Dai, “WaveNet vocoder with limited training data for voice conversion,” in Annual Conference of the International Speech Communication Association, INTERSPEECH , 2018, pp. 1983–1987
1987
Earlier work this paper cites.
D. G. Childers, K. Wu, D. M. Hicks, and B. Yegnanarayana, “Voice conversion,” Speech Communication , vol. 8, no. 2, pp. 147–158, 1989
1989
Earlier work this paper cites.
K. Hornik, “Multilayer feedforward neural networks are universal approximators,” Neural Networks , vol. 2, 1989
1989
Earlier work this paper cites.
C. M. Bishop, “Mixture density networks,” Technical Report NCRG/4228, Aston University, Birmingham, UK , 1994
1994
Earlier work this paper cites.
S. Hchreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
A. Kain, “Spectral voice conversion for text-to-speech synthesis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , vol. 1, 1998, pp. 285–288
1998
Earlier work this paper cites.
D. T. Chappell and J. H. L. Hansen, “Speaker-specific pitch contour modeling and modification,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , vol. 2, 1998, pp. 885–888
1998
Earlier work this paper cites.
L. M. Arslan, “Speaker transformation algorithm using segmental codebooks (STASC),” Speech Communication , vol. 28, no. 3, pp. 211–226, 1999
1999
Earlier work this paper cites.
H. Kawahara, I. Masuda-Katsuse, and A. D. Cheveigné, “Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency based F0 extraction: Possible role of a repetitive structure in sounds,” Speech Communication , vol. 27, no. 3–4, pp. 187–207, 1999
1999
Earlier work this paper cites.
J. Kominek and A. W. Black, “CMU ARCTIC databases for speech synthesis,” http://festvox.org/cmu_arctic/index.html , 2003, Lang. Technol. Inst., Carnegie Mellon Univ., Pittsburgh, PA
2003
Earlier work this paper cites.
Y. Ohtani, T. Toda, H. Saruwatari, and K. Shikano, “Maximum likelihood voice conversion based on GMM with STRAIGHT mixed excitation,” in Proc. ICSLP , 2006, pp. 2266–2269
2006
Earlier work this paper cites.
M. Müller, “Dynamic time warping,” Information retrieval for music and motion , pp. 69–84, 2007
2007
Earlier work this paper cites.
T. Toda, A. W. Black, and K. Tokuda, “Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,” IEEE Transactions on Audio Speech and Language Processing , vol. 15, no. 8, pp. 2222–2235, 2007
2007
Earlier work this paper cites.
S. Desai, E. V. Raghavendra, B. Yegnanarayana, A. W. Black, and K. Prahallad, “Voice conversion using artificial neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , April 2009, pp. 3893–3896
2009
Earlier work this paper cites.
S. Desai, A. W. Black, B. Yegnanarayana, and K. Prahallad, “Spectral mapping using artificial neural networks for voice conversion,” IEEE Transactions on Audio Speech and Language Processing , vol. 18, no. 5, pp. 954–964, 2010
2010
Earlier work this paper cites.
R. H. Laskar, D. Chakrabarty, F. A. Talukdar, K. S. Rao, and K. Banerjee, “Comparing ANN and GMM in a voice conversion framework,” Applied Soft Computing Journal , vol. 12, no. 11, pp. 3332–3342, 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
L.-H. Chen, Z.-H. Ling, L.-J. Liu, and L.-R. Dai, “Voice conversion using deep neural networks with layer-wise generative training,” IEEE/ACM Transactions on Audio Speech and Language Processing , vol. 22, no. 12, pp. 1859–1872, 2014
2014
Cited alongside, same era.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” Neural Information Processing Systems , pp. 3104–3112, 2014
2014
Cited alongside, same era.
K. Cho, B. Van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder–decoder for statistical machine translation,” Empirical Methods in Natural Language Processing , pp. 1724–1734, 2014
2014
Cited alongside, same era.
H. Zen and A. Senior, “Deep mixture density networks for acoustic modeling in statistical parametric speech synthesis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , May 2014, pp. 3844–3848
2014
Y. Wang, R. J. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio et al. , “Tacotron: Towards end-to-end speech synthesis,” in Annual Conference of the International Speech Communication Association, INTERSPEECH , 2017, pp. 4006–4010
2017
Later among the works it cites.
T. Kaneko, H. Kameoka, K. Hiramatsu, and K. Kashino, “Sequence-to-sequence voice conversion with similarity metric learned using generative adversarial networks,” in Annual Conference of the International Speech Communication Association, INTERSPEECH , 2017, pp. 1283–1287
2017
Later among the works it cites.
H. Miyoshi, Y. Saito, S. Takamichi, and H. Saruwatari, “Voice conversion using sequence-to-sequence learning of context posterior probabilities,” in Annual Conference of the International Speech Communication Association, INTERSPEECH , 2017, pp. 1268–1272
2017
Later among the works it cites.
K. Kobayashi, T. Hayashi, A. Tamamori, and T. Toda, “Statistical voice conversion with WaveNet-based waveform generation,” in Annual Conference of the International Speech Communication Association, INTERSPEECH , 2017, pp. 1138–1142
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
D. Kingma and J. Ba, “Adam: A method for stochastic optimization,” Computer Science , 2014
2014
Cited alongside, same era.
L. Sun, S. Kang, K. Li, and H. Meng, “Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2015, pp. 4869–4873
2015
Cited alongside, same era.
T. Nakashika, T. Takiguchi, and Y. Ariki, “Voice conversion using RNN pre-trained by recurrent temporal restricted Boltzmann machines,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 23, no. 3, pp. 580–587, 2015
2015
Cited alongside, same era.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” International Conference on Learning Representations , 2015
2015
Cited alongside, same era.
T. Luong, H. Pham, and C. D. Manning, “Effective approaches to attention-based neural machine translation,” Empirical Methods in Natural Language Processing , pp. 1412–1421, 2015
2015
Cited alongside, same era.
B. Xu, N. Wang, T. Chen, and M. Li, “Empirical evaluation of recitified acitvations in convolutional network,” ICML Deep Learning Workshop, Lille, France, 06-11 July , 2015
2015
Cited alongside, same era.
J. Lai, B. Chen, T. Tan, S. Tong, and K. Yu, “Phone-aware LSTM-RNN for voice conversion,” in IEEE International Conference on Signal Processing (ICSP) , 2016, pp. 177–182
2016
Cited alongside, same era.
M. V. Ramos, “Voice conversion with deep learning,” Master’s Thesis, Instituto Superior Técnico, 10 2016
2016
Cited alongside, same era.
2017
Later among the works it cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , 2017, pp. 6000–6010
2017
Later among the works it cites.
D. Krueger, T. Maharaj, J. Kramar, M. Pezeshki, N. Ballas, N. R. Ke, A. Goyal, Y. Bengio, A. C. Courville, and C. Pal, “Zoneout: Regularizing RNNs by randomly preserving hidden activations,” International Conference on Learning Representations , 2017
2017
Later among the works it cites.
J.-C. Chou, C.-C. Yeh, H.-Y. Lee, and L.-S. Lee, “Multi-target voice conversion without parallel data by adversarially learning disentangled audio representations,” in Annual Conference of the International Speech Communication Association, INTERSPEECH , 2018, pp. 501–505
2018
Closest in time.
S. O. Arik, J. Chen, K. Peng, W. Ping, and Y. Zhou, “Neural voice cloning with a few samples,” in Advances in Neural Information Processing Systems , 2018, pp. 10 040–10 050
2018
Closest in time.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. J. Skerry-Ryan et al. , “Natural TTS synthesis by conditioning WaveNet on mel spectrogram predictions,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 4779–4783
2018
Closest in time.
W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, J. Raiman, and J. P. Miller, “Deep Voice 3: 2000-speaker neural text-to-speech,” International Conference on Learning Representations , 2018
2018
Closest in time.
H. Tachibana, K. Uenoyama, and S. Aihara, “Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 4784–4788
2018
Closest in time.
Y. Jia, Y. Zhang, R. J. Weiss, Q. Wang, J. Shen, F. Ren, Z. Chen, P. Nguyen, R. Pang, I. L. Moreno et al. , “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,” in Advances in Neural Information Processing Systems , 2018, pp. 4485–4495
2018
Closest in time.
E. Nachmani, A. Polyak, Y. Taigman, and L. Wolf, “Fitting new speakers based on a short untranscribed sample,” in International Conference on Machine Learning , 2018, pp. 3683–3691
2018
Closest in time.
Y. Taigman, L. Wolf, A. Polyak, and E. Nachmani, “VoiceLoop: Voice fitting and synthesis via a phonological loop,” International Conference on Learning Representations , 2018
2018
Closest in time.
J. Niwa, T. Yoshimura, K. Hashimoto, K. Oura, Y. Nankaku, and K. Tokuda, “Statistical voice conversion based on WaveNet,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5289–5293
2018
Closest in time.
X. Wang, J. Lorenzo-Trueba, S. Takaki, L. Juvela, and J. Yamagishi, “A comparison of recent waveform generation and acoustic modeling methods for neural-network-based speech synthesis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 4804–4808
2018
Closest in time.
Y. Ai, H.-C. Wu, and Z.-H. Ling, “SampleRNN-based neural vocoder for statistical parametric speech synthesis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 5659–5663
2018
Closest in time.
J.-X. Zhang, Z.-H. Ling, and L.-R. Dai, “Forward attention in sequence-to-sequence acoustic modeling for speech synthesis,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2018, pp. 4789–4793
2018
Closest in time.