Fetching the paper…
Reading the bibliography…
The rapid progress of deep speech synthesis models has posed significant threats to society such as malicious manipulation of content.
Steven Davis and Paul Mermelstein. Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences. IEEE transactions on acoustics, speech, and signal processing, 28(4):357–366, 1980
1980
Earlier work this paper cites.
Van der Maaten L, Hinton G. Visualizing data using t-SNE[J]. Journal of machine learning research, 2008, 9(11)
2008
Earlier work this paper cites.
Xinhui Zhou, Daniel Garcia-Romero, Ramani Duraiswami, Carol Espy-Wilson, and Shihab Shamma, “Linear versus mel frequency cepstral coefficients for speaker recognition,” in 2011 IEEE Workshop on Automatic Speech Recognition & Understanding. IEEE, 2011, pp. 559–564
2011
Earlier work this paper cites.
Malik H. Acoustic environment identification and its applications to audio forensics[J]. IEEE Transactions on Information Forensics and Security, 2013, 8(11): 1827-1837
2013
Earlier work this paper cites.
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. Deep learning. nature, 521(7553): 436–444, 2015
2015
Earlier work this paper cites.
Z. Wu, T. Kinnunen, N. Evans, J. Yamagishi, C. Hanilc¸i, and et al., “Asvspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge,” in Proc. of INTERSPEECH, 2015
2015
Earlier work this paper cites.
Z. Wu, A. Khodabakhsh, C. Demiroglu, J. Yamagishi, D. Saito, T. Toda, and S. King, “SAS: A speaker verification spoofing database containing diverse attacks,” in Proc. IEEE Int. Conf. on Acoustics, Speech, and Signal Processing (ICASSP), 2015
2015
Earlier work this paper cites.
Wang D, Zhang X. Thchs-30: A free chinese speech corpus[J]. arXiv preprint arXiv:1512.01882, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio. Deep learning, volume 1. MIT press Cambridge, 2016
2016
Earlier work this paper cites.
M. Todisco, H. Delgado, and N. Evans, “A new feature for automatic speaker verification antispoofing: Constant q cepstral coefficients,” in Processings of Odyssey 2016, 2016
2016
Earlier work this paper cites.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016
2016
Earlier work this paper cites.
Bendale A, Boult T E. Towards open set deep networks[C]//Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 1563-1572
2016
Earlier work this paper cites.
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al. Tacotron: Towards end-to-end speech synthesis. Proc. Interspeech 2017, pages 4006–4010, 2017
2017
Earlier work this paper cites.
T. Kinnunen, M. Sahidullah, H. Delgado, N. Evans M. Todisco, and et al., “The asvspoof 2017 challenge: Assessing the limits of replay spoofing attack detection,” in Proc. of INTERSPEECH, 2017
2017
Earlier work this paper cites.
Bu H, Du J, Na X, et al. Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline[C]//2017 20th conference of the oriental chapter of the international coordinating committee on speech databases and speech I/O systems and assessment (O-COCOSDA). IEEE, 2017: 1-5
2017
Earlier work this paper cites.
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu. Efficient neural audio synthesis. In International Conference on Machine Learning, pages 2410–2419. PMLR, 2018
2018
Earlier work this paper cites.
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al. Natural tts synthesis by conditioning wavenet on mel spectrogram predictions. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 4779–4783. IEEE, 2018
2018
Earlier work this paper cites.
Zakariah M, Khan M K, Malik H. Digital multimedia audio forensics: past, present and future[J]. Multimedia tools and applications, 2018, 77(1): 1009-1040
2018
Earlier work this paper cites.
M. Todisco, H. Delgado, K.-A. Lee, M. Sahidullah, N. W. D.Evans, T. H. Kinnunen, and J. Yamagishi, “Integrated presentation attack detection and automatic speaker verification: Common features and gaussian back-end fusion,” in Interspeech, 2018
2018
Earlier work this paper cites.
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur. X-vectors: Robust dnn embeddings for speaker recognition. In 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 5329–5333. IEEE
2018
Earlier work this paper cites.
Jie Hu, Li Shen, and Gang Sun. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 7132–7141, 2018
2018
Earlier work this paper cites.
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. Fastspeech: Fast, robust and controllable text-to-speech. In NeurIPS, 2019
2019
Earlier work this paper cites.
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. Neural speech synthesis with transformer network. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 6706–6713, 2019
2019
Earlier work this paper cites.
M. Todisco, X. Wang, V. Vestman, Md. Sahidullah, and K. Lee, “Asvspoof 2019: Future horizons in spoofed and fake audio detection,” in Proc. of INTERSPEECH, 2019
2019
Earlier work this paper cites.
Reimao R, Tzerpos V. For: A dataset for synthetic speech detection[C]//2019 International Conference on Speech Technology and Human-Computer Dialogue (SpeD). IEEE, 2019: 1-10
2019
Earlier work this paper cites.
Yu N, Davis L S, Fritz M. Attributing fake images to gans: Learning and analyzing gan fingerprints[C]//Proceedings of the IEEE/CVF international conference on computer vision. 2019: 7556-7566
2019
Cited alongside, same era.
Shmueli B. Multi-class metrics made simple, Part II: The F 1 -score[J]. Retrieved from Towards Data Science: https://towardsdatascience. com/multi-class-metrics-made-simplepart-ii-the-F 1 -score-ebe8b2c2ca1, 2019
2019
Cited alongside, same era.
Paszke A, Gross S, Massa F, et al. Pytorch: An imperative style, high-performance deep learning library[J]. Advances in neural information processing systems, 2019, 32
2019
Cited alongside, same era.
Yoshihashi R, Shao W, Kawakami R, et al. Classification-reconstruction learning for open-set recognition[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019: 4016-4025
2019
Cited alongside, same era.
Guan, Jiyang, Liang, Jian, Ran He. Are You Stealing My Model? Sample Correlation for Fingerprinting Deep Neural Networks. NeurIPSnull. 2022, [8] Yu, Junchi, Cao, Jie, Ran He. Improving Subgraph Recognition with Variational Graph Information Bottleneck. IEEE Conference on Computer Vision and Pattern Recognition
2022
Closest in time.
Shim H, Heo J, Park J H, et al. Graph attentive feature aggregation for text-independent speaker verification[C]//ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022: 7972-7976
2022
Closest in time.
Tak H, Kamble M, Patino J, et al. Rawboost: A raw data boosting and augmentation method applied to automatic speaker verification anti-spoofing[C]//ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022: 6382-6386
2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Prajwal K R, Mukhopadhyay R, Namboodiri V P, et al. Learning individual speaking styles for accurate lip to speech synthesis[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020: 13796-13805
2020
Cited alongside, same era.
Kong J, Kim J, Bae J. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis[J]. Advances in Neural Information Processing Systems, 2020, 33: 17022-17033
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Jung J, Kim S, Shim H, et al. Improved RawNet with Feature Map Scaling for Text-Independent Speaker Verification Using Raw Waveforms[J]. Proc. Interspeech 2020, 2020: 1496-1500
2020
Cited alongside, same era.
Geng C, Huang S, Chen S. Recent advances in open set recognition: A survey[J]. IEEE transactions on pattern analysis and machine intelligence, 2020, 43(10): 3614-3631
2020
Cited alongside, same era.
Baevski A, Zhou Y, Mohamed A, et al. wav2vec 2.0: A framework for self-supervised learning of speech representations[J]. Advances in neural information processing systems, 2020, 33: 12449-12460
2020
Cited alongside, same era.
2021
Cited alongside, same era.
Lei Y, Yang S, Xie L. Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis[C]//2021 IEEE Spoken Language Technology Workshop (SLT). IEEE, 2021: 423-430
2021
Cited alongside, same era.
2022
Closest in time.
A. Hamza et al., Hamza A, Javed A R R, Iqbal F, et al. Deepfake audio detection via MFCC features using machine learning[J]. IEEE Access, 2022, 10: 134018-134028
2022
Closest in time.
Liu Z, Fu Y, Pan Q, et al. Orientational distribution learning with hierarchical spatial attention for open set recognition[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022, 45(7): 8757-8772
2022
Closest in time.
Chen S, Wang C, Chen Z, et al. Wavlm: Large-scale self-supervised pre-training for full stack speech processing[J]. IEEE Journal of Selected Topics in Signal Processing, 2022, 16(6): 1505-1518
2022
Closest in time.
2022
Closest in time.
Xu W, Dong X, Ma L, et al. Rawformer: an efficient vision transformer for low-light raw image enhancement[J]. IEEE Signal Processing Letters, 2022, 29: 2677-2681
2022
Closest in time.
Yan X, Yi J, Tao J, et al. An initial investigation for detecting vocoder fingerprints of fake audio[C]//Proceedings of the 1st International Workshop on Deepfake Detection for Audio Multimedia. 2022: 61-68
2022
Closest in time.
Salvi D, Bestagini P, Tubaro S. Exploring the synthetic speech attribution problem through data-driven detectors[C]//2022 IEEE International Workshop on Information Forensics and Security (WIFS). IEEE, 2022: 1-6
2022
Closest in time.
Liu X, Wang X, Sahidullah M, et al. Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild[J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023
2023
Closest in time.
2023
Closest in time.
Lu J, Zhang Y, Li Z, et al. Detecting Unknown Speech Spoofing Algorithms with Nearest Neighbors[C]//DADA@ IJCAI. 2023: 89-94
2023
Closest in time.
Qin X, Wang X, Chen Y, et al. From Speaker Verification to Deepfake Algorithm Recognition: Our Learned Lessons from ADD2023 Track 3[C]//DADA@ IJCAI. 2023: 107-112
2023
Closest in time.
Zeng X M, Zhang J T, Li K, et al. Deepfake Algorithm Recognition System with Augmented Data for ADD 2023 Challenge[C]//DADA@ IJCAI. 2023: 31-36
2023
Closest in time.
Wang Z, Wang Q, Yao J, et al. The NPU-ASLP System for Deepfake Algorithm Recognition in ADD 2023 Challenge[C]//DADA@ IJCAI. 2023: 64-69
2023
Closest in time.
Tian Y, Chen Y, Tang Y, et al. Deepfake Algorithm Recognition through Multi-model Fusion Based On Manifold Measure[C]//DADA@ IJCAI. 2023: 76-81
2023
Closest in time.
Shaaban O A, Yildirim R, Alguttar A A. Audio Deepfake Approaches[J]. IEEE Access, 2023, 11: 132652-132682
2023
Closest in time.
Altalahin I, AlZu’bi S, Alqudah A, et al. Unmasking the truth: A deep learning approach to detecting deepfake audio through mfcc features[C]//2023 International Conference on Information Technology (ICIT). IEEE, 2023: 511-518
2023
Closest in time.
Kilinc H H, Kaledibi F. Audio Deepfake Detection by using Machine and Deep Learning[C]//2023 10th International Conference on Wireless Networks and Mobile Communications (WINCOM). IEEE, 2023: 1-5
2023
Closest in time.
Yang T, Wang D, Tang F, et al. Progressive open space expansion for open-set model attribution[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023: 15856-15865
2023
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.