Fetching the paper…
Reading the bibliography…
The performance of spoofing countermeasure systems depends fundamentally upon the use of sufficiently representative training data.
“CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit,”
C. Veaux, J. Yamagishi, et al., · 1994
Earlier work this paper cites.
“Perceptual coding of digital audio,”
T. Painter and A. Spanias, · 2000
Earlier work this paper cites.
“A one-class classification approach to generalised speaker verification spoofing countermeasures using local binary patterns,”
F. Alegre, A. Amehraye, et al., · 2013
Earlier work this paper cites.
“The BOSARIS toolkit: Theory, algorithms and code for surviving the new DCF,”
N. Brümmer and E. de Villiers, · 2013
Earlier work this paper cites.
“Speech recognition and keyword spotting for low-resource languages: BABEL project research at CUED,”
M. JF Gales, K. M Knill, et al., · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P Kingma and J. Ba, · 2014
Earlier work this paper cites.
“Batch normalization: Accelerating deep network training by reducing internal covariate shift,”
S. Ioffe and C. Szegedy, · 2015
Earlier work this paper cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville, · 2016
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, et al., · 2017
Earlier work this paper cites.
“Self-normalizing neural networks,”
G. Klambauer, T. Unterthiner, et al., · 2017
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron Van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
A. v. d. Oord, Y. Li, et al., · 2018
Earlier work this paper cites.
“Speaker recognition from raw waveform with SincNet,”
M. Ravanelli and Y. Bengio, · 2018
Earlier work this paper cites.
“Graph attention networks,”
P. Veličković, G. Cucurull, et al., · 2018
Earlier work this paper cites.
“Attentive statistics pooling for deep speaker embedding,”
K. Okabe, T. Koshinaka, et al., · 2018
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,”
J. Devlin, M.-W Chang, et al., · 2019
Earlier work this paper cites.
“Wav2vec: Unsupervised Pre-Training for Speech Recognition,”
S. Schneider, A. Baevski, et al., · 2019
Earlier work this paper cites.
“Graph u-nets,”
H. Gao and S. Ji, · 2019
Earlier work this paper cites.
“Heterogeneous graph attention network,”
Xiao Wang, Houye Ji, Chuan Shi, Bai Wang, Yanfang Ye, Peng Cui, and Philip S Yu, · 2019
Earlier work this paper cites.
“fairseq: A Fast, Extensible Toolkit for Sequence Modeling,”
M. Ott, S. Edunov, et al., · 2019
Earlier work this paper cites.
“Utterance-level aggregation for speaker recognition in the wild,”
W. Xie, A. Nagrani, et al., · 2019
Earlier work this paper cites.
“ASVspoof 2019: Future Horizons in Spoofed and Fake Audio Detection,”
M. Todisco, X. Wang, et al., · 2019
Earlier work this paper cites.
“Wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,”
A. Baevski, Y. Zhou, et al., · 2020
Earlier work this paper cites.
“Learning Robust and Multilingual Speech Representations,”
K. Kawakami, L. Wang, et al., · 2020
Earlier work this paper cites.
“Generalization of Audio Deepfake Detection,”
T. Chen, A. Kumar, et al., · 2020
Cited alongside, same era.
“MLS: A Large-Scale Multilingual Dataset for Speech Research,”
V. Pratap, Q. Xu, et al., · 2020
Cited alongside, same era.
“Common Voice: A Massively-Multilingual Speech Corpus,”
R. Ardila, M. Branson, et al., · 2020
Cited alongside, same era.
“Improving Multi-Scale Aggregation Using Feature Pyramid Module for Robust Speaker Verification of Variable-Duration Utterances,”
Y. Jung, S. M Kye, et al., · 2020
Cited alongside, same era.
“ASVspoof 2019: a large-scale public database of synthetized, converted and replayed speech,”
X. Wang, J. Yamagishi, et al., · 2020
Cited alongside, same era.
“Tandem Assessment of Spoofing Countermeasures and Automatic Speaker Verification: Fundamentals,”
“Data Augmentation with Signal Companding for Detection of Logical Access Attacks,”
R. K. Das, J. Yang, et al., · 2021
Later among the works it cites.
“An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems,”
Y. Zhang, G. Zhu, et al., · 2021
Later among the works it cites.
“UR Channel-Robust Synthetic Speech Detection System for ASVspoof 2021,”
X. Chen, Y. Zhang, et al., · 2021
Later among the works it cites.
“Similarity Analysis of Self-Supervised Speech Representations,”
Y.-A Chung, Y. Belinkov, et al., · 2021
Later among the works it cites.
“HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,”
W.-N. Hsu, B. Bolte, et al., · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Kinnunen, H. Delgado, et al., · 2020
Cited alongside, same era.
“ASVspoof 2019: spoofing countermeasures for the detection of synthesized, converted and replayed speech,”
A. Nautsch, X. Wang, et al., · 2021
Cited alongside, same era.
“ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection,”
J. Yamagishi, X. Wang, et al., · 2021
Cited alongside, same era.
“STC Antispoofing Systems for the ASVspoof 2021 Challenge,”
A. Tomilov, A. Svishchev, et al., · 2021
Cited alongside, same era.
“Pindrop Labs’ Submission to the ASVspoof 2021 Challenge,”
T. Chen, E. Khoury, et al., · 2021
Cited alongside, same era.
“CRIM’s system description for the ASVspoof 2021 Challenge,”
W. H Kang, J. Alam, et al., · 2021
Cited alongside, same era.
“The Biometric Vox System for the ASVspoof 2021 Challenge,”
J. Cáceres, R. Font, et al., · 2021
Cited alongside, same era.
W.-N. Hsu, Y.-H. H. Tsai, et al., · 2021
Later among the works it cites.
“Self-training and pre-training are complementary for speech recognition,”
Q. Xu, A. Baevski, et al., · 2021
Later among the works it cites.
“WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,”
S. Chen, C. Wang, et al., · 2021
Later among the works it cites.
“Explore wav2vec 2.0 for Mispronunciation Detection,”
X. Xu, Y. Kang, et al., · 2021
Later among the works it cites.
“A Study on Fine-Tuning wav2vec2. 0 Model for the Task of Mispronunciation Detection and Diagnosis,”
L. Peng, K. Fu, et al., · 2021
Later among the works it cites.
“Fine-tuning wav2vec2 for speaker recognition,”
N. Vaessen and D. A van Leeuwen, · 2021
Later among the works it cites.
“Exploring wav2vec 2.0 on speaker verification and language identification,”
Z. Fan, M. Li, et al., · 2021
Later among the works it cites.
“Emotion Recognition from Speech Using Wav2vec 2.0 Embeddings,”
L. Pepino, P. Riera, et al., · 2021
Later among the works it cites.
“VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,”
C. Wang, M. Rivière, et al., · 2021
Later among the works it cites.
“VoxLingua107: a dataset for spoken language recognition,”
J. Valk and T. Alumäe, · 2021
Later among the works it cites.
“End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detection,”
H. Tak, J. Jung, et al., · 2021
Later among the works it cites.
“Graph attention networks for speaker verification,”
J.-w. Jung, H.-S. Heo, et al., · 2021
Later among the works it cites.
“A Comparative Study on Recent Neural Spoofing Countermeasures for Synthetic Speech Detection,”
X. Wang and J. Yamagishi, · 2021
Later among the works it cites.
“Optimizing Tandem Speaker Verification and Anti-Spoofing Systems,”
A. Kanervisto, V. Hautamäki, et al., · 2021
Later among the works it cites.
“AASIST: Audio Anti-Spoofing using Integrated Spectro-Temporal Graph Attention Networks,”
J. Jung, H. Heo, et al., · 2022
Closest in time.
“RawBoost: A Raw Data Boosting and Augmentation Method applied to Automatic Speaker Verification Anti-Spoofing,”
H. Tak, M. Kamble, et al., · 2022
Closest in time.
“RawNeXt: Speaker verification system for variable-duration utterances with deep layer aggregation and extended dynamic scaling policies,”
J.-h. Kim, H.-j. Shim, et al., · 2022
Closest in time.
“Graph attentive feature aggregation for text-independent speaker verification,”
H.-j Shim, J. Heo, et al., · 2022
Closest in time.