Fetching the paper…
Reading the bibliography…
Many datasets have been designed to further the development of fake audio detection.
A. W. Rix, M. P. Hollier, A. P. Hekstra, J. G. Beerends, Perceptual evaluation of speech quality (pesq) the new itu standard for end-to-end speech quality assessment: Part i: Time-delay compensation, Journal of the Audio Engineering Society 50 (10) (2002) 755–764
2002
Earlier work this paper cites.
Y. W. Lau, M. Wagner, D. Tran, Vulnerability of speaker verification to voice mimicking, 2004, pp. 145–148
2004
Earlier work this paper cites.
L. Ma, B. P. Milner, D. Smith, Acoustic environment classification, ACM Transactions on Speech and Language Processing 3 (2) (2006) 1–22
2006
Earlier work this paper cites.
P. C. Loizou, Speech enhancement: Theory and practice, CRC Press, Inc. (2007)
2007
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, J. R. Jensen, A short-time objective intelligibility measure for time-frequency weighted noisy speech, in: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2010, pp. 4214–4217
2010
Earlier work this paper cites.
H. Zhao, H. Malik, Audio recording location identification using acoustic environment signature, IEEE Transactions on Information Forensics and Security 8 (11) (2013) 1746–1759
2013
Earlier work this paper cites.
H. Malik, Acoustic environment identification and its applications to audio forensics, IEEE Transactions on Information Forensics & Security 8 (11) (2013) 1827–1837
2013
Earlier work this paper cites.
R. G. Hautamaki, T. Kinnunen, V. Hautamaki, T. Leino, A. M. Laukkanen, I-vectors meet imitators: on vulnerability of speaker verification systems against voice mimicry, in: Annual Conference of the International Speech Communication Association (Interspeech), 2013
2013
Earlier work this paper cites.
H. Malik, Acoustic environment identification and its applications to audio forensics, IEEE Transactions on Information Forensics & Security 8 (11) (2013) 1827–1837
2013
Earlier work this paper cites.
D. Stowell, D. Giannoulis, E. Benetos, M. Lagrange, M. Plumbley, Detection and classification of acoustic scenes and events, IEEE Transactions on Multimedia 17 (10) (2015) 1733–1746
2015
Earlier work this paper cites.
Z. Wu, N. Evans, T. Kinnunen, J. Yamagishi, F. Alegre, H. Li, Spoofing and countermeasures for speaker verification: A survey, Speech Communication 66 (2015) 130–153
2015
Earlier work this paper cites.
Z. Wu, A. Khodabakhsh, C. Demiroğlu, J. Yamagishi, D. Saito, T. Toda, S. King, Sas: A speaker verification spoofing database containing diverse attacks, 2015, pp. 4440–4444
2015
Earlier work this paper cites.
Z. Wu, T. Kinnunen, N. Evans, J. Yamagishi, C. Hanilc¸i, et al., Asvspoof 2015: the first automatic speaker verification spoofing and countermeasures challenge, in: Annual Conference of the International Speech Communication Association (Interspeech), 2015
2015
Earlier work this paper cites.
M. Zakariah, M. K. Khan, H. Malik, Digital multimedia audio forensics: past, present and future, Multimedia Tools and Applications 77 (2016) 1009–1040
2016
Cited alongside, same era.
T. Kinnunen, M. Sahidullah, H. Delgado, N. E. M. Todisco, et al., The asvspoof 2017 challenge: Assessing the limits of replay spoofing attack detection, in: Annual Conference of the International Speech Communication Association (Interspeech), 2017
2017
Cited alongside, same era.
Kolbaek, Morten, Tan, Zheng-Hua, Jensen, Jesper, Speech intelligibility potential of general and specialized deep neural network based speech enhancement systems, IEEE/ACM Transactions on Audio Speech & Language Processing (2017)
2017
Cited alongside, same era.
A. Owens, A. A. Efros, Audio-visual scene analysis with self-supervised multisensory features, in: European Conference on Computer Vision (ECCV), 2018
2018
Cited alongside, same era.
Z. Wu, R. K. Das1, J. Yang, H. Li, Light convolutional neural network with feature genuinization for detection of synthetic speech attacks, in: Annual Conference of the International Speech Communication Association (Interspeech), 2020
2020
Later among the works it cites.
J. W. Jung, S. B. Kim, H. J. Shim, J. H. Kim, H. J. Yu, Improved rawnet with filter-wise rescaling for text-independent speaker verification using raw waveforms, in: Annual Conference of the International Speech Communication Association (Interspeech), 2020
2020
Later among the works it cites.
J. Yi, Y. Bai, J. Tao, Z. Tian, C. Wang, T. Wang, R. Fu, Half-truth: A partially fake audio detection dataset, in: Annual Conference of the International Speech Communication Association (Interspeech), 2021, pp. 1654–1658
2021
Later among the works it cites.
A. Nautsch, X. Wang, N. W. D. Evans, T. H. Kinnunen, V. Vestman, M. Todisco, H. Delgado, M. Sahidullah, J. Yamagishi, K.-A. Lee, Asvspoof 2019: Spoofing countermeasures for the detection of synthesized, converted and replayed speech, IEEE Transactions on Biometrics, Behavior, and Identity Science 3 (2021) 252–265
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. H. McDermott, A. Torralba, The sound of pixels, in: European Conference on Computer Vision (ECCV), 2018
2018
Cited alongside, same era.
D. Stoller, S. Ewert, S. Dixon, Wave-u-net: A multi-scale neural network for end-to-end audio source separation, 2018
2018
Cited alongside, same era.
A. Pandey, D. Wang, A new framework for cnn-based speech enhancement in the time domain, IEEE/ACM Transactions on Audio, Speech, and Language Processing 27 (7) (2019) 1179–1188
2019
Cited alongside, same era.
C. Gan, H. Zhao, P. Chen, D. G. Cox, A. Torralba, Self-supervised moving vehicle tracking with stereo sound, 2019, pp. 7052–7061
2019
Cited alongside, same era.
G. Lavrentyeva, S. Novoselov, M. Volkova, Y. N. Matveev, M. D. Marsico, Phonespoof: A new dataset for spoofing attack detection in telephone channel, 2019, pp. 2572–2576
2019
Cited alongside, same era.
doi:10.1109/SPED.2019.8906599
R. Reimao, V. Tzerpos, For: A dataset for synthetic speech detection, in: 2019 International Conference on Speech Technology and Human-Computer Dialogue (SpeD), 2019, pp. 1–10 · 2019
Cited alongside, same era.
K. Tan, D. L. Wang, Learning complex spectral mapping with gated convolutional recurrent networks for monaural speech enhancement, IEEE/ACM Transactions on Audio, Speech, and Language Processing 28 (1) (2019) 380–390
2019
Cited alongside, same era.
T. Heittola, A. Mesaros, T. Virtanen, Acoustic scene classification in dcase 2020 challenge: generalization across devices and low complexity solutions, in: Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE 2020), 2020, pp. 56–60
2020
Cited alongside, same era.
2021
Later among the works it cites.
J. Frank, L. Schnherr, Wavefake: A data set to facilitate audio deepfake detection, in: NeurIPS (Benchmark and Dataset Track), 2021
2021
Later among the works it cites.
A. Li, W. Liu, C. Zheng, C. Fan, X. Li, Two heads are better than one: A two-stage complex spectral mapping approach for monaural speech enhancement, IEEE/ACM Transactions on Audio, Speech, and Language Processing 29 (2021) 1829–1843
2021
Later among the works it cites.
J. Yamagishi, X. Wang, M. Todisco, M. Sahidullah, J. Patino, A. Nautsch, X. Liu, K.-A. Lee, T. H. Kinnunen, N. W. D. Evans, H. Delgado, Asvspoof 2021: accelerating progress in spoofed and deepfake speech detection, in: he ASVspoof 2021 Workshop, 2022
2022
Closest in time.
J. Yi, R. Fu, J. Tao, S. Nie, H. Ma, C. Wang, T. Wang, Z. Tian, Y. Bai, C. Fan, S. Liang, S. Wang, S. Zhang, X. Yan, L. Xu, Z. Wen, H. Li, Add 2022: the first audio deep synthesis detection challenge, 2022, pp. 9216–9220
2022
Closest in time.
doi:10.5281/zenodo.6337421
T. Heittola, A. Mesaros, T. Virtanen, TAU Urban Acoustic Scenes 2022 Mobile, Development dataset (Mar. 2022) · 2022
Closest in time.
J. Yi, J. Tao, R. Fu, X. Yan, C. Wang, T. Wang, C. Y. Zhang, X. Zhang, Y. Zhao, Y. Ren, L. Xu, J. Zhou, H. Gu, Z. Wen, S. Liang, Z. Lian, S. Nie, H. Li, Add 2023: the second audio deepfake detection challenge, in: DADA@IJCAI, 2023
2023
Closest in time.
S. Pegg, K. Li, X. Hu, Rtfs-net: Recurrent time-frequency modelling for efficient audio-visual speech separation, in: The Twelfth International Conference on Learning Representations (ICLR), 2024
2024
Closest in time.
Y. Zang, Y. Zhang, M. Heydari, Z. Duan, Singfake: Singing voice deepfake detection, in: Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2024, pp. 1–5
2024
Closest in time.