Fetching the paper…
Reading the bibliography…
The present paper proposes a waveform boundary detection system for audio spoofing attacks containing partially manipulated segments.
“The Kaldi Speech Recognition Toolkit,”
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, et al., · 2011
Earlier work this paper cites.
“Spoofing and Countermeasures for Speaker Verification: A Survey,”
Z. Wu, N. Evans, T. Kinnunen, J. Yamagishi, F. Alegre, and H. Li, · 2015
Earlier work this paper cites.
“Librispeech: An ASR Corpus Based on Public Domain Audio Books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“MUSAN: A Music, Speech, and Noise Corpus,”
D. Snyder, G. Chen, and D. Povey, · 2015
Earlier work this paper cites.
“WaveNet: A Generative Model for Raw Audio,”
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, · 2016
Earlier work this paper cites.
“A Study on Data Augmentation of Reverberant Speech for Robust Speech Recognition,”
T. Ko, V. Peddinti, D. Povey, M. L. Seltzer, and S. Khudanpur, · 2017
Earlier work this paper cites.
“Attention Is All You Need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“Natural TTS Synthesis by Conditioning Wavenet on Mel Spectrogram Predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, et al., · 2018
Earlier work this paper cites.
“A Light CNN for Deep Face Representation with Noisy Labels,”
X. Wu, R. He, Z. Sun, and T. Tan, · 2018
Earlier work this paper cites.
“Aishell-2: Transforming Mandarin ASR Research into Industrial Scale,”
J. Du, X. Na, X. Liu, and H. Bu, · 2018
Earlier work this paper cites.
“One-Shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization,”
J. chieh Chou and H.-Y. Lee, · 2019
Cited alongside, same era.
“Wav2vec: Unsupervised Pre-Training for Speech Recognition,”
S. Schneider, A. Baevski, R. Collobert, and M. Auli, · 2019
Cited alongside, same era.
“HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,”
J. Kong, J. Kim, and J. Bae, · 2020
Cited alongside, same era.
“From Speaker Verification to Multispeaker Speech Synthesis, Deep Transfer with Feedback Constraint,”
Z. Cai, C. Zhang, and M. Li, · 2020
Cited alongside, same era.
“The VOiCES from a Distance Challenge 2019: Analysis of Speaker Verification Results and Remaining Challenges,”
M. K. Nandwana, M. Lomnitz, C. Richey, M. McLaren, D. Castan, L. Ferrer, and A. Lawson, · 2020
Cited alongside, same era.
“The DKU-CMRI System for the ASVspoof 2021 Challenge: Vocoder based Replay Channel Response Estimation,”
X. Wang, X. Qin, T. Zhu, C. Wang, S. Zhang, and M. Li, · 2021
Later among the works it cites.
“Half-Truth: A Partially Fake Audio Detection Dataset,”
J. Yi, Y. Bai, J. Tao, H. Ma, Z. Tian, C. Wang, T. Wang, and R. Fu, · 2021
Later among the works it cites.
“An Initial Investigation for Detecting Partially Spoofed Audio,”
L. Zhang, X. Wang, E. Cooper, J. Yamagishi, J. Patino, and N. Evans, · 2021
Later among the works it cites.
“Multi-task Learning in Utterance-level and Segmental-level Spoof Detection,”
L. Zhang, X. Wang, E. Cooper, and J. Yamagishi, · 2021
Later among the works it cites.
“Unsupervised Speech Recognition,”
A. Baevski, W.-N. Hsu, A. Conneau, and M. Auli, · 2021
Later among the works it cites.
“SIG-VC: A Speaker Information Guided Zero-Shot Voice Conversion System for Both Human Beings and Machines,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Advances in Anti-spoofing: from the Perspective of ASVspoof Challenges,”
M. R. Kamble, H. B. Sailor, H. A. Patil, and H. Li, · 2020
Cited alongside, same era.
“Wav2vec 2.0: A Framework for Self-supervised Learning of Speech Representations,”
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, · 2020
Cited alongside, same era.
“Fragmentvc: Any-to-any Voice Conversion by End-to-end Extracting and Fusing Fine-grained Voice Fragments with Attention,”
Y. Y. Lin, C.-M. Chien, J.-H. Lin, H.-y. Lee, and L.-s. Lee, · 2021
Cited alongside, same era.
“End-to-end Anti-spoofing with RawNet2,”
H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher, · 2021
Cited alongside, same era.
H. Zhang, Z. Cai, X. Qin, and M. Li, · 2022
Closest in time.
“ADD 2022: the first Audio Deep Synthesis Detection Challenge,”
J. Yi, R. Fu, J. Tao, S. Nie, H. Ma, C. Wang, T. Wang, Z. Tian, Y. Bai, C. Fan, S. Liang, S. Wang, S. Zhang, X. Yan, L. Xu, Z. Wen, and H. Li, · 2022
Closest in time.
“Fake Audio Detection Based On Unsupervised Pretraining Models,”
Z. Lv, S. Zhang, K. Tang, and P. Hu, · 2022
Closest in time.
“Partially Fake Audio Detection by Self-Attention-Based Fake Span Discovery,”
H. Wu, H.-C. Kuo, N. Zheng, K.-H. Hung, H.-Y. Lee, Y. Tsao, H.-M. Wang, and H. Meng, · 2022
Closest in time.