Fetching the paper…
Reading the bibliography…
Speaker identity is one of the important characteristics of human speech.
“Frame alignment method for cross-lingual voice conversion,”
Daniel Erro and Asuncion Moreno, · 1972
Earlier work this paper cites.
“A voice conversion framework with tandem feature sparse representation and speaker-adapted wavenet vocoder.,”
Berrak Sisman, Mingyang Zhang, and Haizhou Li, · 1982
Earlier work this paper cites.
Electronics and Communications in Japan (Part I: Communications)
“Mel log spectrum approximation (mlsa) filter for speech synthesis,” · 1983
Earlier work this paper cites.
“Signal estimation from modified short-time fourier transform,”
Daniel Griffin and Jae Lim, · 1984
Earlier work this paper cites.
“Wavenet vocoder with limited training data for voice conversion,”
Li-Juan Liu, Zhen-Hua Ling, Yuan Jiang, Ming Zhou, and Li-Rong Dai, · 1987
Earlier work this paper cites.
“Multilayer feedforward networks are universal approximators,”
Kurt Hornik, Maxwell Stinchcombe, and Halbert White, · 1989
Earlier work this paper cites.
“Voice conversion through vector quantization,”
Masanobu Abe, Satoshi Nakamura, Kiyohiro Shikano, and Hisao Kuwabara, · 1990
Earlier work this paper cites.
“Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones,”
Eric Moulines and Francis Charpentier, · 1990
Earlier work this paper cites.
“Speaker Adaptation and Voice Conversion by Codebook Mapping,”
Kiyohiro Shikano, Satoshi Nakamura, and Masanobu Abe, · 1991
Earlier work this paper cites.
“Voice transformation using psola technique,”
Hélene Valbret, Eric Moulines, and Jean-Pierre Tubach, · 1992
Earlier work this paper cites.
“Atr μ \mu -talk speech synthesis system,”
Yoshinori Sagisaka, Nobuyoshi Kaiki, Naoto Iwahashi, and Katsuhiko Mimura, · 1992
Earlier work this paper cites.
“Unsupervised speaker adaptation from short utterances based on a minimized fuzzy objective function,”
Hiroshi Matsumoto and Yasuki Yamashita, · 1993
Earlier work this paper cites.
“Mel-cepstral distance measure for objective speech quality assessment,”
R. Kubichek, · 1993
Earlier work this paper cites.
“Transformation of formants for voice conversion using artificial neural networks,”
M Narendranath, Hema A Murthy, S Rajendran, and B Yegnanarayana, · 1995
Earlier work this paper cites.
“Acoustic characteristics of speaker individuality: Control and conversion,”
Hisao Kuwabara and Yoshinori Sagisak, · 1995
Earlier work this paper cites.
“Optimising selection of units from speech databases for concatenative synthesis.,”
Alan W Black and Nick Campbell, · 1995
Earlier work this paper cites.
“Voice conversion algorithm based on piecewise linear conversion rules of formant frequency and spectrum tilt,”
Hideyuki Mizuno and Masanobu Abe, · 1995
Earlier work this paper cites.
“The expectation-maximization algorithm,”
Todd K Moon, · 1996
Earlier work this paper cites.
“High-quality voice conversion using spectrogram-based wavenet vocoder.,”
Kuan Chen, Bo Chen, Jiahao Lai, and Kai Yu, · 1997
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Spectral voice conversion for text-to-speech synthesis,”
Alexander Kain and Michael W Macon, · 1998
Earlier work this paper cites.
“A system for voice conversion based on probabilistic classification and a harmonic plus noise model,”
Yannis Stylianou and Olivier Cappe, · 1998
Earlier work this paper cites.
“Continuous probabilistic transform for voice conversion,”
Yannis Stylianou, Olivier Cappé, and Eric Moulines, · 1998
Earlier work this paper cites.
“Continuous probabilistic transform for voice conversion,”
Yannis Stylianou, Olivier Cappé, and Eric Moulines, · 1998
Earlier work this paper cites.
“Speaker adaptation for hmm-based speech synthesis system using mllr,”
Masatsune Tamura, Takashi Masuko, Keiichi Tokuda, and Takao Kobayashi, · 1998
Earlier work this paper cites.
“Speaker transformation algorithm using segmental codebooks (stasc),”
Levent M Arslan, · 1999
Earlier work this paper cites.
“Restructuring speech representations using a pitch-adaptive time–frequency smoothing and an instantaneous-frequency-based f0 extraction: Possible role of a repetitive structure in sounds,”
Hideki Kawahara, Ikuyo Masuda-Katsuse, and Alain De Cheveigne, · 1999
Earlier work this paper cites.
“Learning to forget: Continual prediction with lstm,”
Felix A Gers, Jürgen Schmidhuber, and Fred Cummins, · 1999
Earlier work this paper cites.
“Digital speech processing, synthesis, and recognition(revised and expanded),”
Sadaoki Furui, · 2000
Earlier work this paper cites.
“Voice conversion algorithm based on gaussian mixture model applied to straight,”
Tomoki Toda, Jinlin Lu, Satoshi Nakamura, and Kiyohiro Shikano, · 2000
Earlier work this paper cites.
“Speech parameter generation algorithms for hmm-based speech synthesis,”
Keiichi Tokuda, Takayoshi Yoshimura, Takashi Masuko, Takao Kobayashi, and Tadashi Kitamura, · 2000
Earlier work this paper cites.
“Mel frequency cepstral coefficients for music modeling.,”
Beth Logan et al., · 2000
Earlier work this paper cites.
“Applying the harmonic plus noise model in concatenative speech synthesis,”
Yannis Stylianou, · 2001
Earlier work this paper cites.
“Voice conversion algorithm based on gaussian mixture model with dynamic frequency warping of straight spectrum,”
Tomoki Toda, Hiroshi Saruwatari, and Kiyohiro Shikano, · 2001
Earlier work this paper cites.
“Em algorithms of gaussian mixture model and hidden markov model,”
Guorong Xuan, Wei Zhang, and Peiqi Chai, · 2001
Earlier work this paper cites.
“Algorithms for non-negative matrix factorization,”
Dd Lee and Hs Seung, · 2001
Earlier work this paper cites.
“Design and evaluation of a voice conversion algorithm based on spectral envelope mapping and residual prediction,”
Alexander Kain and Michael W Macon, · 2001
Earlier work this paper cites.
“1534-1, Method for the subjective assessment of intermediate sound quality (MUSHRA),”
ITUR Recommendation, · 2001
Earlier work this paper cites.
“Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,”
Antony W Rix, John G Beerends, Michael P Hollier, and Andries P Hekstra, · 2001
Earlier work this paper cites.
“Cross-language voice conversion evaluation using bilingual databases,”
Mikiko Mashimo, Tomoki Toda, Hiromichi Kawanami, Kiyohiro Shikano, and Nick Campbell, · 2002
Earlier work this paper cites.
“Transformation of spectral envelope for voice conversion based on radial basis function networks,”
Tomomi Watanabe, Takahiro Murakami, Munehiro Namba, Tetsuya Hoya, and Yoshihisa Ishida, · 2002
Earlier work this paper cites.
“Vtln-based crosslanguage voice conversion,”
David Sundermann, Hermann Ney, and H Hoge, · 2003
Earlier work this paper cites.
“Gmm-based voice conversion applied to emotional speech synthesis,”
Hiromichi Kawanami, Yohei Iwami, Tomoki Toda, Hiroshi Saruwatari, and Kiyohiro Shikano, · 2003
Earlier work this paper cites.
“Vtln-based voice conversion,”
David Sundermann and Hermann Ney, · 2003
Earlier work this paper cites.
“Voice characteristics conversion for tts using reverse vtln,”
Matthias Eichner, Matthias Wolff, and Rüdiger Hoffmann, · 2004
Earlier work this paper cites.
“A first step towards text-independent voice conversion,”
Hermann Ney, David Suendermann, Antonio Bonafonte, and Harald Höge, · 2004
Earlier work this paper cites.
“Voice conversion for unknown speakers,”
Hui Ye and Steve J. Young, · 2004
Earlier work this paper cites.
“The cmu arctic speech databases,”
John Kominek and Alan W Black, · 2004
Earlier work this paper cites.
“Spectral conversion based on maximum likelihood estimation considering global variance of converted parameter,”
Tomoki Toda, Alan W Black, and Keiichi Tokuda, · 2005
Earlier work this paper cites.
“Advantages of the mean absolute error (mae) over the root mean square error (rmse) in assessing average model performance,”
Cort J Willmott and Kenji Matsuura, · 2005
Earlier work this paper cites.
“Measuring speech quality for text-to-speech systems: development and assessment of a modified mean opinion score (mos) scale,”
Mahesh Viswanathan and Madhubalan Viswanathan, · 2005
Earlier work this paper cites.
“An evaluation of synthetic speech using the pesq measure,”
Milos Cernak and Milan Rusko, · 2005
Earlier work this paper cites.
“Text-independent voice conversion based on unit selection,”
David Sundermann, Harald Hoge, Antonio Bonafonte, Hermann Ney, Alan Black, and Shri Narayanan, · 2006
Earlier work this paper cites.
“Maximum likelihood voice conversion based on gmm with straight mixed excitation,”
Yamato Ohtani, Tomoki Toda, Hiroshi Saruwatari, and Kiyohiro Shikano, · 2006
Earlier work this paper cites.
“Non-linear frequency scale mapping for voice conversion in text-to-speech system with cepstral description,”
Anna Přibilová and Jiří Přibil, · 2006
Earlier work this paper cites.
“Quality-enhanced voice morphing using maximum likelihood transformations,”
Hui Ye and Steve Young, · 2006
Earlier work this paper cites.
“Voice conversion of non-aligned data using unit selection,”
Daniel Erro, Ferran Diego, and Antonio Bonafonte, · 2006
Earlier work this paper cites.
“Nonparallel training for voice conversion based on a parameter adaptation approach,”
A. Mouchtaris, J. Van der Spiegel, and P. Mueller, · 2006
Earlier work this paper cites.
“Eigenvoice conversion based on gaussian mixture model,”
Tomoki Toda, Yamato Ohtani, and Kiyohiro Shikano, · 2006
Earlier work this paper cites.
“Robust processing techniques for voice conversion,”
Oytun Turk and Levent M Arslan, · 2006
Earlier work this paper cites.
“Low-complexity, nonintrusive speech quality assessment,”
Volodya Grancharov, David Yuheng Zhao, Jonas Lindblom, and W Bastiaan Kleijn, · 2006
Earlier work this paper cites.
“Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory,”
Tomoki Toda, Alan W. Black, and Keiichi Tokuda, · 2007
Earlier work this paper cites.
“Weighted frequency warping for voice conversion,”
Daniel Erro and Asunción Moreno, · 2007
Earlier work this paper cites.
“High individuality voice conversion based on concatenative speech synthesis,”
Kei Fujii, Jun Okawa, and Kaori Suigetsu, · 2007
Earlier work this paper cites.
“Potential biases in mushra listening tests,”
Slawomir Zielinski, Philip Hardisty, Christopher Hummersone, and Francis Rumsey, · 2007
Earlier work this paper cites.
“On the impact of alignment on voice conversion performance,”
Elina Helander, Jan Schwarz, Jani Nurminen, Hanna Silen, and Moncef Gabbouj, · 2008
Earlier work this paper cites.
“Probabilistic feature mapping based on trajectory hmms,”
Heiga Zen, Yoshihiko Nankaku, and Keiichi Tokuda, · 2008
Earlier work this paper cites.
“What is the expectation maximization algorithm?,”
Chuong B Do and Serafim Batzoglou, · 2008
Earlier work this paper cites.
“Multi-layer f0 modeling for hmm-based speech synthesis,”
Cheng-Cheng Wang, Zhen-Hua Ling, Bu-Fan Zhang, and Li-Rong Dai, · 2008
Earlier work this paper cites.
“Speech denoising using nonnegative matrix factorization with priors,”
Kevin W Wilson, Bhiksha Raj, Paris Smaragdis, and Ajay Divakaran, · 2008
Earlier work this paper cites.
“Speech quality assessment,”
Volodya Grancharov and W Bastiaan Kleijn, · 2008
Earlier work this paper cites.
“Optimization of an objective measure for estimating mean opinion score of synthesized speech,” June 10 2008,
Min Chu, Hu Peng, and Yong Zhao, · 2008
Earlier work this paper cites.
“A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments,”
Tomohiro Nakatani, Shigeaki Amano, Toshio Irino, Kentaro Ishizuka, and Tadahisa Kondo, · 2008
Earlier work this paper cites.
“Text-independent voice conversion based on state mapped codebook,”
Meng Zhang, Jianhua Tao, Jilei Tian, and Xia Wang, · 2008
Earlier work this paper cites.
“Query-by-example spoken term detection using phonetic posteriorgram templates,”
Timothy J Hazen, Wade Shen, and Christopher White, · 2009
Earlier work this paper cites.
“Voice conversion using artificial neural networks,”
Srinivas Desai, E Veera Raghavendra, B Yegnanarayana, Alan W Black, and Kishore Prahallad, · 2009
Earlier work this paper cites.
“Voice conversion based on weighted frequency warping,”
Daniel Erro, Asunción Moreno, and Antonio Bonafonte, · 2009
Earlier work this paper cites.
“Pearson correlation coefficient,”
Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen, · 2009
Earlier work this paper cites.
“Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,”
Wei Chu and Abeer Alwan, · 2009
Earlier work this paper cites.
“Voice conversion using partial least squares regression,”
Elina Helander, Tuomas Virtanen, Jani Nurminen, and Moncef Gabbouj, · 2010
Earlier work this paper cites.
“Inca algorithm for training voice conversion systems from nonparallel corpora,”
D. Erro, A. Moreno, and A. Bonafonte, · 2010
Earlier work this paper cites.
“Supervisory data alignment for text-independent voice conversion,”
Jianhua Tao, Meng Zhang, Jani Nurminen, Jilei Tian, and Xia Wang, · 2010
Earlier work this paper cites.
“Front-end factor analysis for speaker verification,”
Najim Dehak, Patrick J Kenny, Réda Dehak, Pierre Dumouchel, and Pierre Ouellet, · 2010
Earlier work this paper cites.
“Spectral mapping using artificial neural networks for voice conversion,”
Srinivas Desai, Alan W Black, B Yegnanarayana, and Kishore Prahallad, · 2010
Earlier work this paper cites.
“Voice conversion using dynamic kernel partial least squares regression,”
Elina Helander, Hanna Silén, Tuomas Virtanen, and Moncef Gabbouj, · 2011
Earlier work this paper cites.
“Event selection from phone posteriorgrams using matched filters,”
Keith Kintzley, Aren Jansen, and Hynek Hermansky, · 2011
Earlier work this paper cites.
“Theory and use of the em algorithm,”
Maya R Gupta, Yihua Chen, et al., · 2011
Earlier work this paper cites.
“Voice conversion using dynamic frequency warping with amplitude scaling, for parallel or nonparallel corpora,”
Elizabeth Godoy, Olivier Rosec, and Thierry Chonavel, · 2011
Earlier work this paper cites.
“Voice conversion using gmm with enhanced global variance,”
Hadas Benisty and David Malah, · 2011
Earlier work this paper cites.
“Prediction of perceived sound quality of synthetic speech,”
Dong-Yan Huang, · 2011
Earlier work this paper cites.
“Exemplar-based voice conversion in noisy environment,”
Ryoichi Takashima, Tetsuya Takiguchi, and Yasuo Ariki, · 2012
Earlier work this paper cites.
“Voice conversion for non-parallel datasets using dynamic kernel partial least squares regression,”
Hanna Silen, Jani Nurminen, Elina Helander, and Moncef Gabbouj, · 2012
Earlier work this paper cites.
“Comparing ann and gmm in a voice conversion framework,”
Rabul Hussain Laskar, D Chakrabarty, Fazal Ahmed Talukdar, K Sreenivasa Rao, and Kalyan Banerjee, · 2012
Earlier work this paper cites.
“Gmm-based emotional voice conversion using spectrum and prosody features,”
Ryo Aihara, Ryoichi Takashima, Tetsuya Takiguchi, and Yasuo Ariki, · 2012
Earlier work this paper cites.
“Improving the quality of standard gmm-based voice conversion systems by considering physically motivated linear transformations,”
Tudor-Cătălin Zorilă, Daniel Erro, and Inma Hernáez, · 2012
Earlier work this paper cites.
“Pitch synchronous transform warping in voice conversion,”
Robert Vích and Martin Vondra, · 2012
Earlier work this paper cites.
“Mixture of factor analyzers using priors from non-parallel speech for voice conversion,”
Z. Wu, T. Kinnunen, E. S. Chng, and H. Li, · 2012
Earlier work this paper cites.
“Articulatory features for expressive speech synthesis,”
Alan W Black, H Timothy Bunnell, Ying Dou, Prasanna Kumar Muthukumar, Florian Metze, Daniel Perry, Tim Polzehl, Kishore Prahallad, Stefan Steidl, and Callie Vaughn, · 2012
Earlier work this paper cites.
“Towards personalised synthesised voices for individuals with vocal disabilities: Voice banking and reconstruction,”
Christophe Veaux, Junichi Yamagishi, and Simon King, · 2013
Earlier work this paper cites.
“Incorporating global variance in the training phase of gmm-based voice conversion,”
Hsin-Te Hwang, Yu Tsao, Hsin-Min Wang, Yih-Ru Wang, and Sin-Horng Chen, · 2013
Earlier work this paper cites.
“Supervised and Unsupervised Speech Enhancement Using Nonnegative Matrix Factorization,”
Nasser Mohammadiha, Paris Smaragdis, and Arne Leijon, · 2013
Earlier work this paper cites.
“Examplar-Based Voice Conversion Using Non-Negative Spectrogram Deconvolution,”
Zhizheng Wu, Tuomas Virtanen, Tomi Kinnunen, Eng Siong Chng, and Haizhou Li, · 2013
Earlier work this paper cites.
“Voice conversion in high-order eigen space using deep belief nets.,”
Toru Nakashika, Ryoichi Takashima, Tetsuya Takiguchi, and Yasuo Ariki, · 2013
Earlier work this paper cites.
“Objective evaluation measures for speaker-adaptive hmm-tts systems,”
Ulpu Remes, Reima Karhila, and Mikko Kurimo, · 2013
Cited alongside, same era.
“Voice conversion versus speaker verification: an overview,”
Zhizheng Wu and Haizhou Li, · 2014
Cited alongside, same era.
“Semi-supervised noise dictionary adaptation for exemplar-based noise robust speech recognition,”
Yi Luan, Daisuke Saito, Yosuke Kashiwagi, Nobuaki Minematsu, and Keikichi Hirose, · 2014
Cited alongside, same era.
“Exemplar-based sparse representation with residual compensation for voice conversion,”
Zhizheng Wu, Tuomas Virtanen, Eng Siong Chng, and Haizhou Li, · 2014
Cited alongside, same era.
“Voice conversion based on non-negative matrix factorization using phoneme-categorized dictionary,”
Ryo Aihara, Toru Nakashika, Tetsuya Takiguchi, and Yasuo Ariki, · 2014
Cited alongside, same era.
“On the use of wavenet as a statistical vocoder,”
Nagaraj Adiga, Vassilis Tsiaras, and Yannis Stylianou, · 2018
Later among the works it cites.
“Wasserstein gan and waveform loss-based acoustic model training for multi-speaker text-to-speech synthesis systems using a wavenet vocoder,”
Yi Zhao, Shinji Takaki, Hieu-Thi Luong, Junichi Yamagishi, Daisuke Saito, and Nobuaki Minematsu, · 2018
Later among the works it cites.
“Deep voice 3: 2000-speaker neural text-to-speech,”
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller, · 2018
Later among the works it cites.
“Efficiently trainable text-to-speech system based on deep convolutional networks with guided attention,”
Hideyuki Tachibana, Katsuya Uenoyama, and Shunsuke Aihara, · 2018
Later among the works it cites.
“Convs2s-vc: Fully convolutional sequence-to-sequence voice conversion,”
Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko, and Nobukatsu Hojo, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peng Song, Yun Jin, Wenming Zheng, and Li Zhao, · 2014
Cited alongside, same era.
“Modulation spectrum-based post-filter for gmm-based voice conversion,”
Shinnosuke Takamichi, Tomoki Toda, Alan W Black, and Satoshi Nakamura, · 2014
Cited alongside, same era.
“Hierarchical modeling of F0 contours for voice conversion,”
Gerard Sanchez, Hanna Silen, Jani Nurminen, and Moncef Gabbouj, · 2014
Cited alongside, same era.
“Pitch transformation in neural network based voice conversion,”
Feng-Long Xie, Yao Qian, Frank K Soong, and Haifeng Li, · 2014
Cited alongside, same era.
“Voice conversion using deep neural networks with speaker-independent pre-training,”
Seyed Hamidreza Mohammadi and Alexander Kain, · 2014
Cited alongside, same era.
“Sequence error (se) minimization training of neural network for voice conversion,”
Feng-Long Xie, Yao Qian, Yuchen Fan, Frank K Soong, and Haifeng Li, · 2014
Cited alongside, same era.
“Voice Conversion Using Deep Neural Networks With Layer-Wise Generative Training,”
Ling-hui Chen, Zhen-hua Ling, Li-juan Liu, and Li-rong Dai, · 2014
Cited alongside, same era.
Later among the works it cites.
“Stargan: Unified generative adversarial networks for multi-domain image-to-image translation,”
Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo, · 2018
Later among the works it cites.
“Learning attribute representations with localization for flexible fashion search,”
Kenan E Ak, Ashraf A Kassim, Joo Hwee Lim, and Jo Yew Tham, · 2018
Later among the works it cites.
“Efficient multi-attribute similarity learning towards attribute-based fashion search,”
Kenan E Ak, Joo Hwee Lim, Jo Yew Tham, and Ashraf A Kassim, · 2018
Later among the works it cites.
“Multimodal unsupervised image-to-image translation,”
Xun Huang, Ming-Yu Liu, Serge Belongie, and Jan Kautz, · 2018
Later among the works it cites.
“High-resolution image synthesis and semantic manipulation with conditional gans,”
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, and Bryan Catanzaro, · 2018
Later among the works it cites.
“Cycle-Consistent Speech Enhancement,”
Zhong Meng, Jinyu Li, Yifan Gong, and Biing-Hwang (Fred) Juang, · 2018
Later among the works it cites.
“Voice conversion using conditional cyclegan,”
Dongsuk Yook, In-Chul Yoo, and Seungho Yoo, · 2018
Later among the works it cites.
“Timbretron: A wavenet (cyclegan (cqt (audio))) pipeline for musical timbre transfer,”
Sicong Huang, Qiyang Li, Cem Anil, Xuchan Bao, Sageev Oore, and Roger B Grosse, · 2018
Later among the works it cites.
“Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks,”
Takuhiro Kaneko and Hirokazu Kameoka, · 2018
Later among the works it cites.
“Rhythm-flexible voice conversion without parallel data using cycle-gan over phoneme posteriorgram sequences,”
Cheng-chieh Yeh, Po-chun Hsu, Ju-chieh Chou, Hung-yi Lee, and Lin-shan Lee, · 2018
Later among the works it cites.
“Error reduction network for dblstm-based voice conversion,”
Mingyang Zhang, Berrak Sisman, Sai Sirisha Rallabandi, Haizhou Li, and Li Zhao, · 2018
Later among the works it cites.
“Voice conversion based on cross-domain features using variational auto encoders,”
Wen-Chin Huang, Hsin-Te Hwang, Yu-Huai Peng, Yu Tsao, and Hsin-Min Wang, · 2018
Later among the works it cites.
“Many-to-many voice conversion based on bottleneck features with variational autoencoder for non-parallel training data,”
Yanping Li, Kong Aik Lee, Yougen Yuan, Haizhou Li, and Zhen Yang, · 2018
Later among the works it cites.
“Non-parallel voice conversion using variational autoencoders conditioned by phonetic posteriorgrams and d-vectors,”
Yuki Saito, Yusuke Ijima, Kyosuke Nishida, and Shinnosuke Takamichi, · 2018
Later among the works it cites.
“Towards end-to-end prosody transfer for expressive speech synthesis with tacotron,”
RJ Skerry-Ryan, Eric Battenberg, Ying Xiao, Yuxuan Wang, Daisy Stanton, Joel Shor, Ron J Weiss, Rob Clark, and Rif A Saurous, · 2018
Later among the works it cites.
“On the analysis of training data for wavenet-based speech synthesis,”
Jakub Vít, Zdeněk Hanzlíček, and Jindřich Matoušek, · 2018
Later among the works it cites.
“Quality-net: An end-to-end non-intrusive speech quality assessment model based on blstm,”
Szu-Wei Fu, Yu Tsao, Hsin-Te Hwang, and Hsin-Min Wang, · 2018
Later among the works it cites.
Mikołaj Bińkowski, Dougal J Sutherland, Michael Arbel, and Arthur Gretton, · 2018
Later among the works it cites.
“The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods,”
Jaime Lorenzo-Trueba, Junichi Yamagishi, Tomoki Toda, Daisuke Saito, Fernando Villavicencio, Tomi Kinnunen, and Zhenhua Ling, · 2018
Later among the works it cites.
“The nu non-parallel voice conversion system for the voice conversion challenge 2018,”
Yichiao Wu, Patrick Lumban Tobing, Tomoki Hayashi, Kazuhiro Kobayashi, and Tomoki Toda, · 2018
Later among the works it cites.
“sprocket: Open-source voice conversion software,”
Kazuhiro Kobayashi and Tomoki Toda, · 2018
Later among the works it cites.
Mingyang Zhang, Xin Wang, Fuming Fang, Haizhou Li, and Junichi Yamagishi, · 2019
Later among the works it cites.
“Evaluating voice conversion-based privacy protection against informed attackers,” 11 2019
Brij Srivastava, Nathalie Vauquier, Md Sahidullah, Aurélien Bellet, Marc Tommasi, and Emmanuel Vincent, · 2019
Later among the works it cites.
“Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion,”
Joan Serrà, Santiago Pascual, and Carlos Segura Perales, · 2019
Later among the works it cites.
Yi-Chiao Wu, Tomoki Hayashi, Patrick Lumban Tobing, Kazuhiro Kobayashi, and Tomoki Toda, · 2019
Later among the works it cites.
“Statistical voice conversion with quasi-periodic wavenet vocoder,”
Yi-Chiao Wu, Patrick Lumban Tobing, Tomoki Hayashi, Kazuhiro Kobayashi, and Tomoki Toda, · 2019
Later among the works it cites.
“Wavenet factorization with singular value decomposition for voice conversion,”
H. Du, X. Tian, L. Xie, and H. Li, · 2019
Later among the works it cites.
“Refined wavenet vocoder for variational autoencoder based voice conversion,”
Wen-Chin Huang, Yi-Chiao Wu, Hsin-Te Hwang, Patrick Lumban Tobing, Tomoki Hayashi, Kazuhiro Kobayashi, Tomoki Toda, Yu Tsao, and Hsin-Min Wang, · 2019
Later among the works it cites.
“Group Sparse Representation with WaveNet Vocoder Adaptation for Spectrum and Prosody Conversion,”
Berrak Sisman, Mingyang Zhang, and Haizhou Li, · 2019
Later among the works it cites.
“WaveGlow: A Flow-based Generative Network for Speech Synthesis,”
Ryan Prenger, Raffael Valle, and Bryan Catanzaro, · 2019
Later among the works it cites.
“Flowavenet: A generative flow for raw audio,”
Sungwon Kim, Sang-Gil Lee, Jongyoon Song, Jaehyeon Kim, and Sungroh Yoon, · 2019
Later among the works it cites.
“Teacher-student training for robust tacotron-based tts,”
Rui Liu, Berrak Sisman, Jingdong Li, Feilong Bao, Guanglai Gao, and Haizhou Li, · 2019
Later among the works it cites.
Machine Learning for Limited Data Voice Conversion
Berrak Sisman, · 2019
Later among the works it cites.
“A speaker-dependent wavenet for voice conversion with non-parallel data,”
Xiaohai Tian, Eng Siong Chng, and Haizhou Li, · 2019
Later among the works it cites.
“A compact framework for voice conversion using wavenet conditioned on phonetic posteriorgrams,”
Hui Lu, Zhiyong Wu, Runnan Li, Shiyin Kang, Jia Jia, and Helen Meng, · 2019
Later among the works it cites.
“Wavenet factorization with singular value decomposition for voice conversion,”
Hongqiang Du, Xiaohai Tian, Lei Xie, and Haizhou Li, · 2019
Later among the works it cites.
“Jointly trained conversion model and wavenet vocoder for non-parallel voice conversion using mel-spectrograms and phonetic posteriorgrams,”
Songxiang Liu, Yuewen Cao, Xixin Wu, Lifa Sun, Xunying Liu, and Helen Meng, · 2019
Later among the works it cites.
“Unsupervised speech representation learning using wavenet autoencoders,”
Jan Chorowski, Ron Weiss, Samy Bengio, and Aaron Oord, · 2019
Later among the works it cites.
“Towards achieving robust universal neural vocoding,”
Jaime Lorenzo-Trueba, Thomas Drugman, Javier Latorre, Thomas Merritt, Bartosz Putrycz, Roberto Barra-Chicote, Alexis Moinet, and Vatsal Aggarwal, · 2019
Later among the works it cites.
“A comparison of recent neural vocoders for speech signal reconstruction,”
Prachi Govalkar, Johannes Fischer, Frank Zalkow, and Christian Dittmar, · 2019
Later among the works it cites.
“Singing voice synthesis using deep autoregressive neural networks for acoustic modeling,”
Yuan-Hao Yi, Yang Ai, Zhen-Hua Ling, and Li-Rong Dai, · 2019
Later among the works it cites.
“Real-time neural text-to-speech with sequence-to-sequence acoustic model and waveglow or single gaussian wavernn vocoders,”
Takuma Okamoto, Tomoki Toda, Yoshinori Shiga, and Hisashi Kawai, · 2019
Later among the works it cites.
“Parametric resynthesis with neural vocoders,”
Soumi Maiti and Michael I Mandel, · 2019
Later among the works it cites.
“Neural source-filter-based waveform model for statistical parametric speech synthesis,”
Xin Wang, Shinji Takaki, and Junichi Yamagishi, · 2019
Later among the works it cites.
Xin Wang and Junichi Yamagishi, · 2019
Later among the works it cites.
“Neural source-filter waveform models for statistical parametric speech synthesis,”
Xin Wang, Shinji Takaki, and Junichi Yamagishi, · 2019
Later among the works it cites.
“Melgan: Generative adversarial networks for conditional waveform synthesis,”
Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Brébisson, Yoshua Bengio, and Aaron C Courville, · 2019
Later among the works it cites.
“Cross-lingual voice conversion with bilingual phonetic posteriorgrams and average modeling,”
Yi Zhou, Xiaohai Tian, Haihua Xu, Rohan Kumar Das, and Haizhou Li, · 2019
Later among the works it cites.
“Sequence-to-sequence acoustic modeling for voice conversion,”
Jing-Xuan Zhang, Zhen-Hua Ling, Li-Juan Liu, Yuan Jiang, and Li-Rong Dai, · 2019
Later among the works it cites.
“Atts2s-vc: Sequence-to-sequence voice conversion with attention and context preservation mechanisms,”
K. Tanaka, H. Kameoka, T. Kaneko, and N. Hojo, · 2019
Later among the works it cites.
“Few-shot unsupervised image-to-image translation,”
Ming-Yu Liu, Xun Huang, Arun Mallya, Tero Karras, Timo Aila, Jaakko Lehtinen, and Jan Kautz, · 2019
Later among the works it cites.
“Attribute manipulation generative adversarial networks for fashion images,”
Kenan E Ak, Joo Hwee Lim, Jo Yew Tham, and Ashraf A Kassim, · 2019
Later among the works it cites.
Deep learning approaches for attribute manipulation and text-to-image synthesis
Kenan Emir Ak, · 2019
Later among the works it cites.
“Semantically consistent hierarchical text to fashion image synthesis with an enhanced-attentional generative adversarial network,”
Kenan Emir Ak, Joo Hwee Lim, Jo Yew Tham, and Ashraf Kassim, · 2019
Later among the works it cites.
“Cyclegan-vc2: Improved cyclegan-based non-parallel voice conversion,”
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, and Nobukatsu Hojo, · 2019
Later among the works it cites.
“Voice conversion with cyclic recurrent neural network and fine-tuned wavenet vocoder,”
Patrick Lumban Tobing, Yi-Chiao Wu, Tomoki Hayashi, Kazuhiro Kobayashi, and Tomoki Toda, · 2019
Later among the works it cites.
“On the study of generative adversarial networks for cross-lingual voice conversion,”
Berrak Sisman, Mingyang Zhang, Minghui Dong, and Haizhou Li, · 2019
Later among the works it cites.
Wen-Chin Huang, Tomoki Hayashi, Yi-Chiao Wu, Hirokazu Kameoka, and Tomoki Toda, · 2019
Later among the works it cites.
“Improving sequence-to-sequence voice conversion by adding text-supervision,”
Jing-Xuan Zhang, Zhen-Hua Ling, Yuan Jiang, Li-Juan Liu, Chen Liang, and Li-Rong Dai, · 2019
Later among the works it cites.
“Bootstrapping non-parallel voice conversion from speaker-adaptive text-to-speech,”
Hieu-Thi Luong and Junichi Yamagishi, · 2019
Later among the works it cites.
“Parrotron: An end-to-end speech-to-speech conversion model and its applications to hearing-impaired speech and speech separation,”
Fadi Biadsy, Ron J Weiss, Pedro J Moreno, Dimitri Kanvesky, and Ye Jia, · 2019
Later among the works it cites.
Ju-Chieh Chou, Cheng chieh Yeh, and Hung yi Lee, · 2019
Later among the works it cites.
“Generating diverse high-fidelity images with vq-vae-2,”
Ali Razavi, Aaron van den Oord, and Oriol Vinyals, · 2019
Later among the works it cites.
“Group latent embedding for vector quantized variational autoencoder in non-parallel voice conversion.,”
Shaojin Ding and Ricardo Gutierrez-Osuna, · 2019
Later among the works it cites.
“Self-attention generative adversarial networks,”
Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena, · 2019
Later among the works it cites.
“Mosnet: Deep learning based objective assessment for voice conversion,”
Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang, · 2019
Later among the works it cites.
“High fidelity speech synthesis with adversarial networks,”
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C Cobo, and Karen Simonyan, · 2019
Later among the works it cites.
“Sequence-to-sequence acoustic modeling for voice conversion,”
J. Zhang, Z. Ling, L. Liu, Y. Jiang, and L. Dai, · 2019
Later among the works it cites.
“ASVspoof 2019: future horizons in spoofed and fake audio detection,”
Massimiliano Todisco, Xin Wang, Ville Vestman, Md. Sahidullah, Héctor Delgado, Andreas Nautsch, Junichi Yamagishi, Nicholas Evans, Tomi H. Kinnunen, and Kong Aik Lee, · 2019
Later among the works it cites.
“Asvspoof 2019: a large-scale public database of synthetic, converted and replayed speech,” 2019
Xin Wang, Junichi Yamagishi, Massimiliano Todisco, Hector Delgado, Andreas Nautsch, Nicholas Evans, Md Sahidullah, Ville Vestman, Tomi Kinnunen, Kong Aik Lee, Lauri Juvela, Paavo Alku, Yu-Huai Peng, Hsin-Te Hwang, Yu Tsao, Hsin-Min Wang, Sebastien Le Maguer, Markus Becker, Fergus Henderson, Rob Clark, Yu Zhang, Quan Wang, Ye Jia, Kai Onuma, Koji Mushika, Takashi Kaneda, Yuan Jiang, Li-Juan Liu, Yi-Chiao Wu, Wen-Chin Huang, Tomoki Toda, Kou Tanaka, Hirokazu Kameoka, Ingmar Steiner, Driss Matrouf, Jean-Francois Bonastre, Avashna Govender, Srikanth Ronanki, Jing-Xuan Zhang, and Zhen-Hua Ling, · 2019
Later among the works it cites.
“LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu, · 2019
Later among the works it cites.
“Defending your voice: Adversarial attack on voice conversion,”
Chien yu Huang, Yist Y. Lin, Hung yi Lee, and Lin shan Lee, · 2020
Closest in time.
“Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,”
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim, · 2020
Closest in time.
Seung-won Park, Doo-young Kim, and Myun-chul Joe, · 2020
Closest in time.
“Stargan v2: Diverse image synthesis for multiple domains,”
Yunjey Choi, Youngjung Uh, Jaejun Yoo, and Jung-Woo Ha, · 2020
Closest in time.
“Semantically consistent text to fashion image synthesis with an enhanced attentional generative adversarial network,”
Kenan E Ak, Joo Hwee Lim, Jo Yew Tham, and Ashraf A Kassim, · 2020
Closest in time.
“Speech synthesis of children’s reading based on cyclegan model,”
Ning Jia, Chunjun Zheng, and Wei Sun, · 2020
Closest in time.
“Transforming spectrum and prosody for emotional voice conversion with non-parallel training data,”
Kun Zhou, Berrak Sisman, and Haizhou Li, · 2020
Closest in time.
“Converting anyone’s emotion: Towards speaker-independent emotional voice conversion,”
Kun Zhou, Berrak Sisman, Mingyang Zhang, and Haizhou Li, · 2020
Closest in time.
“Wavetts: Tacotron-based tts with joint time-frequency domain loss,”
Rui Liu, Berrak Sisman, Feilong Bao, Guanglai Gao, and Haizhou Li, · 2020
Closest in time.
“Modeling prosodic phrasing with multi-task learning in tacotron-based tts,”
Rui Liu, Berrak Sisman, Feilong Bao, Guanglai Gao, and Haizhou Li, · 2020
Closest in time.
“Expressive tts training with frame and style reconstruction loss,”
Rui Liu, Berrak Sisman, Guanglai Gao, and Haizhou Li, · 2020
Closest in time.
“Transfer learning from speech synthesis to voice conversion with non-parallel training data,” 2020
Mingyang Zhang, Yi Zhou, Li Zhao, and Haizhou Li, · 2020
Closest in time.
“Nautilus: a versatile voice cloning system,”
Hieu-Thi Luong and Junichi Yamagishi, · 2020
Closest in time.
“Multi-target emotional voice conversion with neural vocoders,”
Songxiang Liu, Yuewen Cao, and Helen Meng, · 2020
Closest in time.
“One-shot voice conversion by vector quantization,”
Da-Yi Wu and Hung-yi Lee, · 2020
Closest in time.
“Vqvc+: One-shot voice conversion by vector quantization and u-net architecture,”
Da-Yi Wu, Yen-Hao Chen, and Hung-Yi Lee, · 2020
Closest in time.
“Unsupervised representation disentanglement using cross domain features and adversarial learning in variational autoencoder based voice conversion,”
Wen-Chin Huang, Hao Luo, Hsin-Te Hwang, Chen-Chou Lo, Yu-Huai Peng, Yu Tsao, and Hsin-Min Wang, · 2020
Closest in time.
“Edge-gan: Edge conditioned multi-view face image generation,”
Heqing Zou, Kenan E Ak, and Ashraf A Kassim, · 2020
Closest in time.
“Transferring source style in non-parallel voice conversion,”
Songxiang Liu, Yuewen Cao, Shiyin Kang, Na Hu, Xunying Liu, Dan Su, Dong Yu, and Helen Meng, · 2020
Closest in time.
“Deepconversion: Voice conversion with limited parallel training data,”
Mingyang Zhang, Berrak Sisman, Li Zhao, and Haizhou Li, · 2020
Closest in time.
Jennifer Williams, Joanna Rownicka, Pilar Oplustil, and Simon King, · 2020
Closest in time.
“Non-parallel sequence-to-sequence voice conversion with disentangled linguistic and speaker representations,”
J. Zhang, Z. Ling, and L. Dai, · 2020
Closest in time.
“Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion,”
Yi Zhao, Wen-Chin Huang, Xiaohai Tian, Junichi Yamagishi, Rohan Kumar Das, Tomi Kinnunen, Zhenhua Ling, and Tomoki Toda, · 2020
Closest in time.
“Voxceleb: Large-scale speaker verification in the wild,”
Arsha Nagrani, Joon Son Chung, Weidi Xie, and Andrew Zisserman, · 2020
Closest in time.
“Espnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai, · 2020
Closest in time.
“ASVspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge,”
Zhizheng Wu, Tomi Kinnunen, Nicholas Evans, Junichi Yamagishi, Cemal Hanilçi, Md. Sahidullah, and Aleksandr Sizov, · 2041
Closest in time.