Fetching the paper…
Reading the bibliography…
Advances in neural speech synthesis have brought us technology that is not only close to human naturalness, but is also capable of instant voice cloning with little data, and is highly accessible with pre-trained models available.
“Secure spread spectrum watermarking for images, audio and video,”
I. J. Cox, J. Kilian, T. Leighton, and T. Shamoon, · 1996
Earlier work this paper cites.
“Techniques for data hiding,”
W. Bender, D. Gruhl, N. Morimoto, and A. Lu, · 1996
Earlier work this paper cites.
“Echo hiding,”
D. Gruhl, A. Lu, and W. Bender, · 1996
Earlier work this paper cites.
“Watermarking schemes evaluation,”
F. Petitcolas, · 2000
Earlier work this paper cites.
Digitale Wasserzeichen fuer Audiodaten
M. Steinebach, · 2004
Earlier work this paper cites.
“Generative adversarial nets,”
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, · 2014
Earlier work this paper cites.
“Spoofing and countermeasures for speaker verification: A survey,”
Z. Wu, N. Evans, T. Kinnunen, J. Yamagishi, F. Alegre, and H. Li, · 2015
Earlier work this paper cites.
“MUSAN: A Music, Speech, and Noise Corpus,”
D. Snyder, G. Chen, and D. Povey, · 2015
Earlier work this paper cites.
“Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. Lopez Moreno, Y. Wu, et al., · 2018
Earlier work this paper cites.
“A light CNN for deep face representation with noisy labels,”
X. Wu, R. He, Z. Sun, and T. Tan, · 2018
Earlier work this paper cites.
“GELP: GAN-excited linear prediction for speech synthesis from mel-spectrogram,”
L. Juvela, B. Bollepalli, J. Yamagishi, and P. Alku, · 2019
Earlier work this paper cites.
“Probability density distillation with generative adversarial networks for high-quality parallel waveform generation,”
R. Yamamoto, E. Song, and J.-M. Kim, · 2019
Cited alongside, same era.
“CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),” 2019
J. Yamagishi, C. Veaux, and K. MacDonald, · 2019
Cited alongside, same era.
“HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,”
J. Kong, J. Kim, and J. Bae, · 2020
Cited alongside, same era.
“End-to-end anti-spoofing with RawNet2,”
H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher, · 2020
Cited alongside, same era.
“Speech is Silver, Silence is Golden: What do ASVspoof-trained Models Really Learn?,”
N. Müller, F. Dieckmann, P. Czempin, R. Canals, K. Böttinger, and J. Williams, · 2021
Cited alongside, same era.
“ADD 2022: The first Audio Deep Synthesis Detection Challenge,”
J. Yi, R. Fu, J. Tao, S. Nie, H. Ma, C. Wang, T. Wang, Z. Tian, Y. Bai, C. Fan, S. Liang, S. Wang, S. Zhang, X. Yan, L. Xu, Z. Wen, and H. Li, · 2022
Later among the works it cites.
“Does Audio Deepfake Detection Generalize?,”
N. M. Müller, P. Czempin, F. Dieckmann, A. Froghyar, and K. Böttinger, · 2022
Later among the works it cites.
“Robust speech watermarking by a jointly trained embedder and detector using a DNN,”
K. Pavlović, S. Kovačević, I. Djurović, and A. Wojciechowski, · 2022
Later among the works it cites.
“Responsible disclosure of generative models using scalable fingerprinting,”
N. Yu, V. Skripniuk, D. Chen, L. S. Davis, and M. Fritz, · 2022
Later among the works it cites.
“Removing batch normalization boosts adversarial training,”
H. Wang, A. Zhang, S. Zheng, X. Shi, M. Li, and Z. Wang, · 2022
Later among the works it cites.
“ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Zhang, W. Wang, and P. Zhang, · 2021
Cited alongside, same era.
“Artificial fingerprinting for generative models: Rooting deepfake attribution in training data,”
N. Yu, V. Skripniuk, S. Abdelnabi, and M. Fritz, · 2021
Cited alongside, same era.
“A comparative study on recent neural spoofing countermeasures for synthetic speech detection,”
X. Wang and J. Yamagishi, · 2021
Cited alongside, same era.
“WaveGrad: Estimating gradients for waveform generation,”
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, · 2021
Cited alongside, same era.
“DiffWave: A versatile diffusion model for audio synthesis,”
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, · 2021
Cited alongside, same era.
“YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone,”
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. Gölge, and M. A. Ponti, · 2022
Cited alongside, same era.
X. Liu, X. Wang, M. Sahidullah, J. Patino, H. Delgado, T. Kinnunen, M. Todisco, J. Yamagishi, N. Evans, A. Nautsch, and K. A. Lee, · 2023
Closest in time.
“Wavmark: Watermarking for audio generation,”
G. Chen, Y. Wu, S. Liu, T. Liu, X. Du, and F. Wei, · 2023
Closest in time.
“Spoofed training data for speech spoofing countermeasure can be efficiently created using neural vocoders,”
X. Wang and J. Yamagishi, · 2023
Closest in time.
“Generalization of spoofing countermeasures: A case study with ASVspoof 2015 and BTAS 2016 corpora,”
D. Paul, M. Sahidullah, and G. Saha, · 2051
Closest in time.
“A comparison of features for synthetic speech detection,”
M. Sahidullah, T. Kinnunen, and C. Hanilçi, · 2091
Closest in time.