Fetching the paper…
Reading the bibliography…
This paper describes a method based on a sequence-to-sequence learning (Seq2Seq) with attention and context preservation mechanism for voice conversion (VC) tasks.
“Continuous probabilistic transform for voice conversion,”
Yannis Stylianou, Olivier Cappé, and Eric Moulines, · 1998
Earlier work this paper cites.
“The CMU Arctic speech databases,”
John Kominek and Alan W Black, · 2004
Earlier work this paper cites.
“Nonparallel training for voice conversion based on a parameter adaptation approach,”
Athanasios Mouchtaris, Jan Van der Spiegel, and Paul Mueller, · 2006
Earlier work this paper cites.
“Map-based adaptation for speech conversion using adaptation data selection and non-parallel training,”
Chung-Han Lee and Chung-Hsien Wu, · 2006
Earlier work this paper cites.
“Eigenvoice conversion based on gaussian mixture model,”
Toda Tomoki, Ohtani Yamato, and Shikano Kiyohiro, · 2006
Earlier work this paper cites.
“Improving the intelligibility of dysarthric speech,”
Alexander B Kain, John-Paul Hosom, Xiaochuan Niu, Jan PH van Santen, Melanie Fried-Oken, and Janice Staehely, · 2007
Earlier work this paper cites.
“Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,”
Tomoki Toda, Alan W Black, and Keiichi Tokuda, · 2007
Earlier work this paper cites.
“Data-driven emotion conversion in spoken english,”
Zeynep Inanoglu and Steve Young, · 2009
Earlier work this paper cites.
“Foreign accent conversion in computer assisted pronunciation training,”
Daniel Felps, Heather Bortfeld, and Ricardo Gutierrez-Osuna, · 2009
Earlier work this paper cites.
“Spectral mapping using artificial neural networks for voice conversion,”
Srinivas Desai, Alan W Black, B Yegnanarayana, and Kishore Prahallad, · 2010
Earlier work this paper cites.
“Voice conversion using partial least squares regression,”
Elina Helander, Tuomas Virtanen, Jani Nurminen, and Moncef Gabbouj, · 2010
Earlier work this paper cites.
“Rectified linear units improve restricted boltzmann machines,”
Vinod Nair and Geoffrey E Hinton, · 2010
Earlier work this paper cites.
“One-to-many voice conversion based on tensor representation of speaker space,”
Daisuke Saito, Keisuke Yamamoto, Nobuaki Minematsu, and Keikichi Hirose, · 2011
Earlier work this paper cites.
“Speaking-aid systems using GMM-based voice conversion for electrolaryngeal speech,”
Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, and Kiyohiro Shikano, · 2012
Earlier work this paper cites.
“Statistical voice conversion techniques for body-conducted unvoiced speech enhancement,”
Tomoki Toda, Mikihiro Nakagiri, and Kiyohiro Shikano, · 2012
Earlier work this paper cites.
“Exemplar-based voice conversion using sparse representation in noisy environments,”
Ryoichi Takashima, Tetsuya Takiguchi, and Yasuo Ariki, · 2013
Earlier work this paper cites.
“Exemplar-based sparse representation with residual compensation for voice conversion,”
Zhizheng Wu, Tuomas Virtanen, Eng Siong Chng, and Haizhou Li, · 2014
Earlier work this paper cites.
“Voice conversion using deep neural networks with layer-wise generative training,”
Ling-Hui Chen, Zhen-Hua Ling, Li-Juan Liu, and Li-Rong Dai, · 2014
Cited alongside, same era.
“Voice conversion based on speaker-dependent restricted boltzmann machines,”
Toru Nakashika, Tetsuya Takiguchi, and Yasuo Ariki, · 2014
Cited alongside, same era.
“High-order sequence modeling using speaker-dependent recurrent temporal restricted boltzmann machines for voice conversion,”
Toru Nakashika, Tetsuya Takiguchi, and Yasuo Ariki, · 2014
Cited alongside, same era.
“Sequence to sequence learning with neural networks,”
Ilya Sutskever, Oriol Vinyals, and Quoc V Le, · 2014
Cited alongside, same era.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Cited alongside, same era.
“Voice conversion using input-to-output highway networks,”
Yuki Saito, Shinnosuke Takamichi, and Hiroshi Saruwatari, · 2017
Later among the works it cites.
“Tacotron: Towards end-to-end speech synthesis,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Later among the works it cites.
“Online and linear-time attention by enforcing monotonic alignments,”
Colin Raffel, Minh-Thang Luong, Peter J Liu, Ron J Weiss, and Douglas Eck, · 2017
Later among the works it cites.
“Listening while speaking: Speech chain by deep learning,”
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura, · 2017
Later among the works it cites.
“Unpaired image-to-image translation using cycle-consistent adversarial networks,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik Kingma and Jimmy Ba, · 2014
Cited alongside, same era.
“Voice conversion using deep bidirectional long short-term memory based recurrent neural networks,”
Lifa Sun, Shiyin Kang, Kun Li, and Helen Meng, · 2015
Cited alongside, same era.
“Voice conversion from non-parallel corpora using variational auto-encoder,”
Chin-Cheng Hsu, Hsin-Te Hwang, Yi-Chiao Wu, Yu Tsao, and Hsin-Min Wang, · 2016
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Cited alongside, same era.
“Dual learning for machine translation,”
Di He, Yingce Xia, Tao Qin, Liwei Wang, Nenghai Yu, Tieyan Liu, and Wei-Ying Ma, · 2016
Cited alongside, same era.
“Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,”
Sun Lifa, Li Kun, Wang Hao, Kang Shiyin, and Meng Helen, · 2016
Cited alongside, same era.
“Merlin: An open source neural network speech synthesis system,”
Zhizheng Wu, Oliver Watts, and Simon King, · 2016
Cited alongside, same era.
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros, · 2017
Later among the works it cites.
“Voice conversion using sequence-to-sequence learning of context posterior probabilities,”
Hiroyuki Miyoshi, Yuki Saito, Shinnosuke Takamichi, and Hiroshi Saruwatari, · 2017
Later among the works it cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Later among the works it cites.
“Voice impersonation using generative adversarial networks,”
Yang Gao, Rita Singh, and Bhiksha Raj, · 2018
Closest in time.
“StarGAN-VC: Non-parallel many-to-many voice conversion with star generative adversarial networks,”
Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, and Nobukatsu Hojo, · 2018
Closest in time.
“Non-parallel voice conversion using variational autoencoders conditioned by phonetic posteriorgrams and d-vectors,”
Yuki Saito, Yusuke Ijima, Kyosuke Nishida, and Shinnosuke Takamichi, · 2018
Closest in time.
“Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al., · 2018
Closest in time.
“Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention,”
Hideyuki Tachibana, Katsuya Uenoyama, and Shunsuke Aihara, · 2018
Closest in time.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Ekaterina Gonina, et al., · 2018
Closest in time.
“Sequence-to-sequence acoustic modeling for voice conversion,”
Jing-Xuan Zhang, Zhen-Hua Ling, Li-Rong Dai, Li-Juan Liu, and Yuan Jiang, · 2018
Closest in time.
“Wavenet vocoder with limited training data for voice conversion,”
Li-Juan Liu, Zhen-Hua Ling, Yuan Jiang, Ming Zhou, and Li-Rong Dai, · 2018
Closest in time.
“sprocket: Open-source voice conversion software,”
Kazuhiro Kobayashi and Tomoki Toda, · 2018
Closest in time.
“ConvS2S-VC: Fully convolutional sequence-to-sequence voice conversion,”
Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko, and Nobukatsu Hojo, · 2019
Closest in time.