Fetching the paper…
Reading the bibliography…
Voice style transfer, also called voice conversion, seeks to modify one speaker's voice to generate speech as if it came from another (target) speaker.
Mel-cepstral distance measure for objective speech quality assessment
Robert Kubichek · 1993
Earlier work this paper cites.
Using dynamic time warping to find patterns in time series
Donald J Berndt and James Clifford · 1994
Earlier work this paper cites.
Map-based adaptation for speech conversion using adaptation data selection and non-parallel training
Chung-Han Lee and Chung-Hsien Wu · 2006
Earlier work this paper cites.
Speaking aid system for total laryngectomees using voice conversion of body transmitted artificial speech
Keigo Nakamura, Tomoki Toda, Hiroshi Saruwatari, and Kiyohiro Shikano · 2006
Earlier work this paper cites.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Lindasalwa Muda, Mumtaj Begam, and Irraivan Elamvazuthi · 2010
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
XuanLong Nguyen, Martin J Wainwright, and Michael I Jordan · 2010
Earlier work this paper cites.
Applying voice conversion to concatenative singing-voice synthesis
Fernando Villavicencio and Jordi Bonada · 2010
Earlier work this paper cites.
Voice conversion using dynamic frequency warping with amplitude scaling, for parallel or nonparallel corpora
Elizabeth Godoy, Olivier Rosec, and Thierry Chonavel · 2011
Earlier work this paper cites.
Spoken digits recognition using weighted mfcc and improved features for dynamic time warping
Santosh V Chapaneri · 2012
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Earlier work this paper cites.
Isolated speech recognition using mfcc and dtw
Shivanker Dev Dhingra, Geeta Nijhawan, and Poonam Pandit · 2013
Earlier work this paper cites.
Ensemble estimation of multivariate f-divergence
Kevin R Moon and Alfred O Hero · 2014
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
Danilo Rezende and Shakir Mohamed · 2015
Earlier work this paper cites.
Deep variational information bottleneck
Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy · 2016
Earlier work this paper cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Image style transfer using convolutional neural networks
Leon A Gatys, Alexander S Ecker, and Matthias Bethge · 2016
Cited alongside, same era.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2016
Cited alongside, same era.
Voice conversion from non-parallel corpora using variational auto-encoder
Chin-Cheng Hsu, Hsin-Te Hwang, Yi-Chiao Wu, Yu Tsao, and Hsin-Min Wang · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2016
Stargan-vc: Non-parallel many-to-many voice conversion using star generative adversarial networks
Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, and Nobukatsu Hojo · 2018
Later among the works it cites.
Cyclegan-vc: Non-parallel voice conversion using cycle-consistent adversarial networks
Takuhiro Kaneko and Hirokazu Kameoka · 2018
Later among the works it cites.
Disentangling by factorising
Hyunjik Kim and Andriy Mnih · 2018
Later among the works it cites.
Glow: Generative flow with invertible 1x1 convolutions
Durk P Kingma and Prafulla Dhariwal · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Non-parallel voice conversion using variational autoencoders conditioned by phonetic posteriorgrams and d-vectors
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Analysis of the voice conversion challenge 2016 evaluation results
Mirjam Wester, Zhizheng Wu, and Junichi Yamagishi · 2016
Cited alongside, same era.
Coherent online video style transfer
Dongdong Chen, Jing Liao, Lu Yuan, Nenghai Yu, and Gang Hua · 2017
Cited alongside, same era.
Darla: Improving zero-shot transfer in reinforcement learning
Irina Higgins, Arka Pal, Andrei Rusu, Loic Matthey, Christopher Burgess, Alexander Pritzel, Matthew Botvinick, Charles Blundell, and Alexander Lerchner · 2017
Cited alongside, same era.
Real-time neural style transfer for videos
Haozhi Huang, Hao Wang, Wenhan Luo, Lin Ma, Wenhao Jiang, Xiaolong Zhu, Zhifeng Li, and Wei Liu · 2017
Cited alongside, same era.
Arbitrary style transfer in real-time with adaptive instance normalization
Xun Huang and Serge Belongie · 2017
Cited alongside, same era.
Deep photo style transfer
Fujun Luan, Sylvain Paris, Eli Shechtman, and Kavita Bala · 2017
Cited alongside, same era.
An overview of voice conversion systems
Seyed Hamidreza Mohammadi and Alexander Kain · 2017
Cited alongside, same era.
Yuki Saito, Yusuke Ijima, Kyosuke Nishida, and Shinnosuke Takamichi · 2018
Later among the works it cites.
Adaptive wavenet vocoder for residual compensation in gan-based voice conversion
Berrak Sisman, Mingyang Zhang, Sakriani Sakti, Haizhou Li, and Satoshi Nakamura · 2018
Later among the works it cites.
Petar Veličković, William Fedus, William L Hamilton, Pietro Liò, Yoshua Bengio, and R Devon Hjelm · 2018
Later among the works it cites.
Generalized end-to-end loss for speaker verification
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno · 2018
Later among the works it cites.
Unsupervised text style transfer using language models as discriminators
Zichao Yang, Zhiting Hu, Chris Dyer, Eric P Xing, and Taylor Berg-Kirkpatrick · 2018
Later among the works it cites.
One-shot voice conversion by separating speaker and content representations with instance normalization
Ju-chieh Chou and Hung-Yi Lee · 2019
Later among the works it cites.
Multiple-attribute text rewriting
Guillaume Lample, Sandeep Subramanian, Eric Smith, Ludovic Denoyer, Marc’Aurelio Ranzato, and Y-Lan Boureau · 2019
Later among the works it cites.
Challenging common assumptions in the unsupervised learning of disentangled representations
Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem · 2019
Later among the works it cites.
Autovc: Zero-shot voice style transfer with only autoencoder loss
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, and Mark Hasegawa-Johnson · 2019
Later among the works it cites.
Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion
Joan Serrà, Santiago Pascual, and Carlos Segura Perales · 2019
Later among the works it cites.
Learning controllable fair representations
Jiaming Song, Pratyusha Kalluri, Aditya Grover, Shengjia Zhao, and Stefano Ermon · 2019
Later among the works it cites.
Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92)
Junichi Yamagishi, Christophe Veaux, Kirsten MacDonald, et al · 2019
Later among the works it cites.
Unsupervised multi-target domain adaptation: An information theoretic approach
Behnam Gholami, Pritish Sahu, Ognjen Rudovic, Konstantinos Bousmalis, and Vladimir Pavlovic · 2020
Later among the works it cites.