Fetching the paper…
Reading the bibliography…
Voice conversion (VC), as a voice style transfer technology, is becoming increasingly prevalent while raising serious concerns about its illegal use.
Embedding Limitations with Digital-audio Watermarking Method Based on Cochlear Delay Characteristics
Masashi Unoki, Kuniaki Imabeppu, Daiki Hamada, Atsushi Haniu, and Ryota Miyauchi. 2011 · 2011
Earlier work this paper cites.
A dual-channel time-spread echo method for audio watermarking
Yong Xiang, Iynkaran Natgunanathan, Dezhong Peng, Wanlei Zhou, and Shui Yu. 2011 · 2011
Earlier work this paper cites.
Speaker-independent style conversion for HMM-based expressive speech synthesis. In 2013 IEEE International Conference on Acoustics, Speech and Signal Processing . IEEE, 7864–7868
Hiroki Kanagawa, Takashi Nose, and Takao Kobayashi. 2013 · 2013
Earlier work this paper cites.
Time-spread echo-based audio watermarking with optimized imperceptibility and robustness
Guang Hua, Jonathan Goh, and Vrizlynn. L. L. Thing. 2015 · 2014
Earlier work this paper cites.
Auto-Encoding Variational Bayes. In International Conference on Learning Representations (ICLR)
Diederik P. Kingma and Max Welling. 2014 · 2014
Earlier work this paper cites.
Voice expression conversion with factorised HMM-TTS models. In Fifteenth Annual Conference of the International Speech Communication Association
Javier Latorre, Vincent Wan, and Kayoko Yanagisawa. 2014 · 2014
Earlier work this paper cites.
Librispeech: an ASR corpus based on public domain audio books. In Acoustics, Speech and Signal Processing (ICASSP), 2015 IEEE International Conference on . IEEE, 5206–5210
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
SUPERSEDED - CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit
Christophe Veaux, Junichi Yamagishi, and Kirsten MacDonald. 2016 · 2016
Earlier work this paper cites.
Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis
Ye Jia, Yu Zhang, Ron Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, and Yonghui Wu. 2018 · 2018
Earlier work this paper cites.
Glow: Generative flow with invertible 1x1 convolutions
Durk P Kingma and Prafulla Dhariwal. 2018 · 2018
Earlier work this paper cites.
Patchwork-based audio watermarking robust against de-synchronization and recapturing attacks
Zhenghui Liu, Yuankun Huang, and Jiwu Huang. 2018 · 2018
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection for dimension reduction
Leland McInnes, John Healy, and James Melville. 2018 · 2018
Earlier work this paper cites.
SNR-constrained heuristics for optimizing the scaling parameter of robust audio watermarking
Zhaopin Su, Guofu Zhang, Feng Yue, Lejie Chang, Jianguo Jiang, and Xin Yao. 2018 · 2018
Earlier work this paper cites.
Generalized end-to-end loss for speaker verification. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 4879–4883
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno. 2018 · 2018
Earlier work this paper cites.
H.R.3230 - DEEP FAKES Accountability Act
Yvette D. Clarke. 2019 · 2019
Earlier work this paper cites.
WaveGlow: A flow-based generative network for speech synthesis. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3617–3621
Ryan Prenger, Rafael Valle, and Bryan Catanzaro. 2019 · 2019
Earlier work this paper cites.
Speech sanitizer: Speech content desensitization and voice anonymization
Jianwei Qian, Haohua Du, Jiahui Hou, Linlin Chen, Taeho Jung, and Xiang-Yang Li. 2019a · 2019
Cited alongside, same era.
Fraudsters Used AI to Mimic CEO’s Voice in Unusual Cybercrime Case
Catherine Stupp. 2019 · 2019
Cited alongside, same era.
Singing voice conversion with disentangled representations of singer and vocal technique using variational autoencoders. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3277–3281
Yin-Jyun Luo, Chin-Cheng Hsu, Kat Agres, and Dorien Herremans. 2020 · 2020
Cited alongside, same era.
Voice conversion for dubbing using linear predictive coding and hidden markov model
Firra M Mukhneri, Inung Wijayanto, and Sugondo Hadiyoso. 2020 · 2020
Cited alongside, same era.
Unsupervised Speech Decomposition via Triple Information Bottleneck. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning Research, Vol. 119) . PMLR, 7836–7846
Voicemod
2022 · 2022
Later among the works it cites.
Deepfake Face Traceability with Disentangling Reversing Network
Jiaxin Ai, Zhongyuan Wang, Baojin Huang, and Zhen Han. 2022 · 2022
Later among the works it cites.
Provisions on the Administration of Deep Synthesis Internet Information Services (Draft for solicitation of comments)
CAC. 2022 · 2022
Later among the works it cites.
YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone. In International Conference on Machine Learning . PMLR, 2709–2720
Edresson Casanova, Julian Weber, Christopher D Shulby, Arnaldo Candido Junior, Eren Gölge, and Moacir A Ponti. 2022 · 2022
Later among the works it cites.
SpeechSplit2.0: Unsupervised Speech Disentanglement for Voice Conversion without Tuning Autoencoder Bottlenecks. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6332–6336
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson, and David D. Cox. 2020 · 2020
Cited alongside, same era.
Evaluating voice conversion-based privacy protection against informed attackers. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2802–2806
Brij Mohan Lal Srivastava, Nathalie Vauquier, Md Sahidullah, Aurélien Bellet, Marc Tommasi, and Emmanuel Vincent. 2020 · 2020
Cited alongside, same era.
Multi-Subspace Echo Hiding Based on Time-Frequency Similarities of Audio Signals
Shengbei Wang, Weitao Yuan, and Masashi Unoki. 2020 · 2020
Cited alongside, same era.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In International Conference on Machine Learning . PMLR, 9929–9939
Tongzhou Wang and Phillip Isola. 2020 · 2020
Cited alongside, same era.
Speaker recognition based on deep learning: An overview
Zhongxin Bai and Xiao-Lei Zhang. 2021 · 2021
Cited alongside, same era.
Distribution-Preserving Steganography Based on Text-to-Speech Generative Models
Kejiang Chen, Hang Zhou, Hanqing Zhao, Dongdong Chen, Weiming Zhang, and Nenghai Yu. 2021 · 2021
Cited alongside, same era.
I’m a victim of voice cloning: VP Mohadi
Mukudzei Chingwere. 2021 · 2021
Cited alongside, same era.
HiNet: deep image hiding by invertible network. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 4733–4742
Junpeng Jing, Xin Deng, Mai Xu, Jianyi Wang, and Zhenyu Guan. 2021 · 2021
Cited alongside, same era.
Chak Ho Chan, Kaizhi Qian, Yang Zhang, and Mark Hasegawa-Johnson. 2022 · 2022
Later among the works it cites.
S3prl-vc: Open-source voice conversion framework with self-supervised speech representations. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6552–6556
Wen-Chin Huang, Shu-Wen Yang, Tomoki Hayashi, Hung-Yi Lee, Shinji Watanabe, and Tomoki Toda. 2022 · 2022
Later among the works it cites.
Robust disentangled variational speech representation learning for zero-shot voice conversion. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6572–6576
Jiachen Lian, Chunlei Zhang, and Dong Yu. 2022 · 2022
Later among the works it cites.
DeAR: A Deep-learning-based Audio Re-recording Resilient Watermarking
Chang Liu, Jie Zhang, Han Fang, Zehua Ma, Weiming Zhang, and Nenghai Yu. 2022 · 2022
Later among the works it cites.
Human perception of audio deepfakes. In Proceedings of the 1st International Workshop on Deepfake Detection for Audio Multimedia . 85–91
Nicolas M Müller, Karla Pizzi, and Jennifer Williams. 2022 · 2022
Later among the works it cites.
Robust speech watermarking by a jointly trained embedder and detector using a DNN
Kosta Pavlović, Slavko Kovačević, Igor Djurović, and Adam Wojciechowski. 2022 · 2022
Later among the works it cites.
ContentVec: An improved self-supervised speech representation by disentangling speakers. In International Conference on Machine Learning . PMLR, 18003–18017
Kaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni, Cheng-I Lai, David Cox, Mark Hasegawa-Johnson, and Shiyu Chang. 2022 · 2022
Later among the works it cites.
AVQVC: One-shot Voice Conversion by Vector Quantization with applying contrastive learning. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 4613–4617
Huaizhen Tang, Xulong Zhang, Jianzong Wang, Ning Cheng, and Jing Xiao. 2022 · 2022
Later among the works it cites.
Robust Invertible Image Steganography. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 7875–7884
Youmin Xu, Chong Mou, Yujie Hu, Jingfen Xie, and Jian Zhang. 2022 · 2022
Later among the works it cites.
Add 2022: the first audio deep synthesis detection challenge. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 9216–9220
Jiangyan Yi, Ruibo Fu, Jianhua Tao, Shuai Nie, Haoxin Ma, Chenglong Wang, Tao Wang, Zhengkun Tian, Ye Bai, Cunhang Fan, et al · 2022
Later among the works it cites.
SIG-VC: A Speaker Information Guided Zero-Shot Voice Conversion System for Both Human Beings and Machines. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6567–65571
Haozhe Zhang, Zexin Cai, Xiaoyi Qin, and Ming Li. 2022 · 2022
Later among the works it cites.