Fetching the paper…
Reading the bibliography…
With the rapid advancement of technologies like text-to-speech (TTS) and voice conversion (VC), detecting deepfake voices has become increasingly crucial.
Speaker verification against synthetic speech
Lian-Wu Chen, Wu Guo, and Li-Rong Dai · 2010
Earlier work this paper cites.
Speaker verification performance degradation against spoofing and tampering attacks
Jesús Villalba and Eduardo Lleida · 2010
Earlier work this paper cites.
Synthetic speech discrimination using pitch pattern statistics derived from image analysis
Phillip L De Leon, Bryan Stewart, and Junichi Yamagishi · 2012
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Karol J Piczak · 2015
Earlier work this paper cites.
A review of time-scale modification of music signals
Jonathan Driedger and Meinard Müller · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Spoofing detection from a feature representation perspective
Xiaohai Tian, Zhizheng Wu, Xiong Xiao, Eng Siong Chng, and Haizhou Li · 2016
Earlier work this paper cites.
Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al · 2016
Earlier work this paper cites.
Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng · 2017
Earlier work this paper cites.
Jsut corpus: free large-scale japanese speech corpus for end-to-end speech synthesis
Ryosuke Sonobe, Shinnosuke Takamichi, and Hiroshi Saruwatari · 2017
Earlier work this paper cites.
Darts: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2018
Earlier work this paper cites.
Novel technique of customizing the audio fade-out shape
Lucian Lupşa-Tătaru · 2018
Earlier work this paper cites.
Replay detection using cqt-based modified group delay feature and resnewt network in asvspoof 2019
Xingliang Cheng, Mingxing Xu, and Thomas Fang Zheng · 2019
Earlier work this paper cites.
Res2net: A new multi-scale backbone architecture
Shang-Hua Gao, Ming-Ming Cheng, Kai Zhao, Xin-Yu Zhang, Ming-Hsuan Yang, and Philip Torr · 2019
Earlier work this paper cites.
Implementing the fade-in audio effect for real-time computing
Lucian Lupşa-Tătaru · 2019
Earlier work this paper cites.
Magicdata mandarin chinese read speech corpus, 05 2019
Magic Data Technology Co., Ltd · 2019
Earlier work this paper cites.
Optimization of false acceptance/rejection rates and decision threshold for end-to-end text-dependent speaker verification systems
Victoria Mingote, Antonio Miguel, Dayana Ribas, Alfonso Ortega Giménez, and Eduardo Lleida · 2019
Earlier work this paper cites.
For: A dataset for synthetic speech detection
Ricardo Reimao and Vassilios Tzerpos · 2019
Earlier work this paper cites.
Asvspoof 2019: Future horizons in spoofed and fake audio detection
Massimiliano Todisco, Xin Wang, Ville Vestman, Md Sahidullah, Héctor Delgado, Andreas Nautsch, Junichi Yamagishi, Nicholas Evans, Tomi Kinnunen, and Kong Aik Lee · 2019
Earlier work this paper cites.
Voicepop: A pop noise based anti-spoofing system for voice authentication on smartphones
Qian Wang, Xiu Lin, Man Zhou, Yanjiao Chen, Cong Wang, Qi Li, and Xiangyang Luo · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations, 2020
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
Improved rawnet with feature map scaling for text-independent speaker verification using raw waveforms
Jee-weon Jung, Seung-bin Kim, Hye-jin Shim, Ju-ho Kim, and Ha-Jin Yu · 2020
Earlier work this paper cites.
Aishell-3: A multi-speaker mandarin tts corpus and the baselines
Yao Shi, Hui Bu, Xin Xu, Shaoji Zhang, and Ming Li · 2020
Earlier work this paper cites.
Deepsonar: Towards effective and robust detection of ai-synthesized fake voices
Run Wang, Felix Juefei-Xu, Yihao Huang, Qing Guo, Xiaofei Xie, Lei Ma, and Yang Liu · 2020
Earlier work this paper cites.
A history of audio effects
Thomas Wilmering, David Moffat, Alessia Milo, and Mark B Sandler · 2020
Earlier work this paper cites.
Xls-r: Self-supervised cross-lingual speech representation learning at scale
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick Von Platen, Yatharth Saraf, Juan Pino, et al · 2021
Earlier work this paper cites.
Wavefake: A data set to facilitate audio deepfake detection
Joel Frank and Lea Schönherr · 2021
Earlier work this paper cites.
Yihui Fu, Luyao Cheng, Shubo Lv, Yukai Jv, Yuxiang Kong, Zhuo Chen, Yanxin Hu, Lei Xie, Jian Wu, Hui Bu, et al · 2021
Cited alongside, same era.
Partially-connected differentiable architecture search for deepfake and spoofing detection
Wanying Ge, Michele Panariello, Jose Patino, Massimiliano Todisco, and Nicholas Evans · 2021
Cited alongside, same era.
Raw differentiable architecture search for speech deepfake and spoofing detection
Wanying Ge, Jose Patino, Massimiliano Todisco, and Nicholas Evans · 2021
Cited alongside, same era.
Towards end-to-end synthetic speech detection
Guang Hua, Andrew Beng Jin Teoh, and Haijian Zhang · 2021
Cited alongside, same era.
How deep are the fakes? focusing on audio deepfake: A survey
Diff-hiervc: Diffusion-based hierarchical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation, 2023
Ha-Yeong Choi, Sang-Hoon Lee, and Seong-Whan Lee · 2023
Later among the works it cites.
Samo: Speaker attractor multi-center one-class learning for voice anti-spoofing
Siwen Ding, You Zhang, and Zhiyao Duan · 2023
Later among the works it cites.
Bts-e: Audio deepfake detection using breathing-talking-silence encoder
Thien-Phuc Doan, Long Nguyen-Vu, Souhwan Jung, and Kihun Hong · 2023
Later among the works it cites.
Houjian Guo, Chaoran Liu, Carlos Toshinori Ishi, and Hiroshi Ishiguro · 2023
Later among the works it cites.
Large language models for software engineering: A systematic literature review
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zahra Khanjani, Gabrielle Watson, and Vandana P Janeja · 2021
Cited alongside, same era.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech, 2021
Jaehyeon Kim, Jungil Kong, and Juhee Son · 2021
Cited alongside, same era.
Starganv2-vc: A diverse, unsupervised, non-parallel framework for natural-sounding voice conversion, 2021
Yinghao Aaron Li, Ali Zare, and Nima Mesgarani · 2021
Cited alongside, same era.
Hemlata Tak, Jee-weon Jung, Jose Patino, Madhu Kamble, Massimiliano Todisco, and Nicholas Evans · 2021
Cited alongside, same era.
End-to-end anti-spoofing with rawnet2
Hemlata Tak, Jose Patino, Massimiliano Todisco, Andreas Nautsch, Nicholas Evans, and Anthony Larcher · 2021
Cited alongside, same era.
Graph attention networks for anti-spoofing, 2021
Hemlata Tak, Jee weon Jung, Jose Patino, Massimiliano Todisco, and Nicholas Evans · 2021
Cited alongside, same era.
Stc antispoofing systems for the asvspoof2021 challenge
Anton Tomilov, Aleksei Svishchev, Marina Volkova, Artem Chirkovskiy, Alexander Kondratev, and Galina Lavrentyeva · 2021
Cited alongside, same era.
A comparative study on recent neural spoofing countermeasures for synthetic speech detection
Xin Wang and Junich Yamagishi · 2021
Cited alongside, same era.
Later among the works it cites.
Audio deepfakes: A survey
Zahra Khanjani, Gabrielle Watson, and Vandana P Janeja · 2023
Later among the works it cites.
Phase-aware spoof speech detection based on res2net with phase network
Juntae Kim and Sung Min Ban · 2023
Later among the works it cites.
A continual deepfake detection benchmark: Dataset, methods, and essentials
Chuqiao Li, Zhiwu Huang, Danda Pani Paudel, Yabin Wang, Mohamad Shahbazi, Xiaopeng Hong, and Luc Van Gool · 2023
Later among the works it cites.
Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models, 2023
Yinghao Aaron Li, Cong Han, Vinay S. Raghavan, Gavin Mischler, and Nima Mesgarani · 2023
Later among the works it cites.
Deepfake generation and detection: Case study and challenges
Yogesh Patel, Sudeep Tanwar, Rajesh Gupta, Pronaya Bhattacharya, Innocent Ewean Davidson, Royi Nyameko, Srinivas Aluvala, and Vrince Vimal · 2023
Later among the works it cites.
Openvoice: Versatile instant voice cloning
Zengyi Qin, Wenliang Zhao, Xumin Yu, and Xin Sun · 2023
Later among the works it cites.
Ai-synthesized voice detection using neural vocoder artifacts
Chengzhe Sun, Shan Jia, Shuwei Hou, and Siwei Lyu · 2023
Later among the works it cites.
Add 2023: the second audio deepfake detection challenge
Jiangyan Yi, Jianhua Tao, Ruibo Fu, Xinrui Yan, Chenglong Wang, Tao Wang, Chu Yuan Zhang, Xiaohui Zhang, Yan Zhao, Yong Ren, et al · 2023
Later among the works it cites.
Audio deepfake detection: A survey
Jiangyan Yi, Chenglong Wang, Jianhua Tao, Xiaohui Zhang, Chu Yuan Zhang, and Yan Zhao · 2023
Later among the works it cites.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Ziqiang Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, Yu Wu, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Later among the works it cites.
https://keithito.com/LJ-Speech-Dataset/
The LJ speech dataset · 2024
Closest in time.
https://www.kaggle.com/code/jasoncallaway/fra-txt-details , September 2018
Fra.Txt details · 2024
Closest in time.
https://www.wsj.com/articles/fraudsters-use-ai-to-mimic-ceos-voice-in-unusual-cybercrime-case-11567157402 , August 2019
Fraudsters used AI to mimic CEO’s voice in unusual cybercrime case · 2024
Closest in time.
Rawbmamba: End-to-end bidirectional state space model for audio deepfake detection
Yujie Chen, Jiangyan Yi, Jun Xue, Chenglong Wang, Xiaohui Zhang, Shunbo Dong, Siding Zeng, Jianhua Tao, Lv Zhao, and Cunhang Fan · 2024
Closest in time.
Dddm-vc: Decoupled denoising diffusion models with disentangled representation and prior mixup for verified robust voice conversion
Ha-Yeong Choi, Sang-Hoon Lee, and Seong-Whan Lee · 2024
Closest in time.
Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al · 2024
Closest in time.
Towards benchmarking and evaluating deepfake detection
Jingyi Deng, Chenhao Lin, Pengbin Hu, Chao Shen, Qian Wang, Qi Li, and Qiming Li · 2024
Closest in time.
Deepfake generation and detection: A benchmark and survey, 2024
Gan Pei, Jiangning Zhang, Menghan Hu, Zhenyu Zhang, Chengjie Wang, Yunsheng Wu, Guangtao Zhai, Jian Yang, Chunhua Shen, and Dacheng Tao · 2024
Closest in time.
Clad: Robust audio deepfake detection against manipulation attacks with contrastive learning
Haolin Wu, Jing Chen, Ruiying Du, Cong Wu, Kun He, Xingcan Shang, Hao Ren, and Guowen Xu · 2024
Closest in time.
Ctrsvdd: A benchmark dataset and baseline analysis for controlled singing voice deepfake detection, 2024
Yongyi Zang, Jiatong Shi, You Zhang, Ryuichi Yamamoto, Jionghao Han, Yuxun Tang, Shengyuan Xu, Wenxiao Zhao, Jing Guo, Tomoki Toda, and Zhiyao Duan · 2024
Closest in time.
What to remember: Self-adaptive continual learning for audio deepfake detection
XiaoHui Zhang, Jiangyan Yi, Chenglong Wang, Chu Yuan Zhang, Siding Zeng, and Jianhua Tao · 2024
Closest in time.