Fetching the paper…
Reading the bibliography…
Substantial improvements have been achieved in recent years in voice conversion, which converts the speaker characteristics of an utterance into those of another speaker without changing the linguistic content of the utterance.
“Fooling end-to-end speaker verification with adversarial examples,”
F. Kreuk, Y. Adi, M. Cisse, and J. Keshet, · 1966
Earlier work this paper cites.
“Signal estimation from modified short-time fourier transform,”
D. Griffin and Jae Lim, · 1984
Earlier work this paper cites.
“Vulnerability of speaker verification to voice mimicking,”
Yee Wah Lau, M. Wagner, and D. Tran, · 2004
Earlier work this paper cites.
“A study on spoofing attack in state-of-the-art speaker verification: the telephone speech case,”
Z. Wu, T. Kinnunen, E. S. Chng, H. Li, and E. Ambikairajah, · 2012
Earlier work this paper cites.
“Spoofing and countermeasures for automatic speaker verification,”
Nicholas Evans, Tomi Kinnunen, and Junichi Yamagishi, · 2013
Earlier work this paper cites.
“Intriguing properties of neural networks,”
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus, · 2014
Earlier work this paper cites.
“Explaining and harnessing adversarial examples,”
Ian Goodfellow, Jonathon Shlens, and Christian Szegedy, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P. Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Deep features for automatic spoofing detection,”
Yanmin Qian, Nanxin Chen, and Kai Yu, · 2016
Earlier work this paper cites.
“Wavenet: A generative model for raw audio,”
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu, · 2016
Earlier work this paper cites.
“Adversarial machine learning at scale,”
Alexey Kurakin, Ian J. Goodfellow, and Samy Bengio, · 2017
Earlier work this paper cites.
“Did you hear that? adversarial examples against automatic speech recognition,”
Moustafa Alzantot, Bharathan Balaji, and Mani B. Srivastava, · 2017
Earlier work this paper cites.
“Houdini: Fooling deep structured visual and speech recognition models with adversarial examples,”
Moustapha M Cisse, Yossi Adi, Natalia Neverova, and Joseph Keshet, · 2017
Cited alongside, same era.
“Audio replay attack detection with deep learning frameworks,”
Galina Lavrentyeva, Sergey Novoselov, Egor Malykh, Alexander Kozlov, Oleg Kudashev, and Vadim Shchemelinin, · 2017
Cited alongside, same era.
“You can hear but you cannot steal: Defending against voice impersonation attacks on smartphones,”
S. Chen, K. Ren, S. Piao, C. Wang, Q. Wang, J. Weng, L. Su, and A. Mohaisen, · 2017
Cited alongside, same era.
“Towards evaluating the robustness of neural networks,”
N. Carlini and D. Wagner, · 2017
Cited alongside, same era.
“Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,” 2017
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al., · 2017
Cited alongside, same era.
“Cyclegan-vc2: Improved cyclegan-based non-parallel voice conversion,”
T. Kaneko, H. Kameoka, K. Tanaka, and N. Hojo, · 2019
Later among the works it cites.
“One-Shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization,”
Ju chieh Chou and Hung-Yi Lee, · 2019
Later among the works it cites.
“AutoVC: Zero-shot voice style transfer with only autoencoder loss,”
Kaizhi Qian, Yang Zhang, Shiyu Chang, Xuesong Yang, and Mark Hasegawa-Johnson, · 2019
Later among the works it cites.
“Exposing deepfake videos by detecting face warping artifacts,”
Yuezun Li and Siwei Lyu, · 2019
Later among the works it cites.
“Adversarial attacks against automatic speech recognition systems via psychoacoustic hiding,”
Lea Schönherr, Katharina Kohls, Steffen Zeiler, Thorsten Holz, and Dorothea Kolossa, · 2019
Later among the works it cites.
“Targeted adversarial examples for black box audio systems,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Cited alongside, same era.
“Multi-target voice conversion without parallel data by adversarially learning disentangled audio representations,”
Ju chieh Chou, Cheng chieh Yeh, Hung yi Lee, and Lin shan Lee, · 2018
Cited alongside, same era.
“Stargan-vc: Non-parallel many-to-many voice conversion using star generative adversarial networks,”
Hirokazu Kameoka, Takuhiro Kaneko, Kou Tanaka, and Nobukatsu Hojo, · 2018
Cited alongside, same era.
“Towards deep learning models resistant to adversarial attacks,”
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu, · 2018
Cited alongside, same era.
“Adversarial examples for generative models,”
Jernej Kos, Ian Fischer, and Dawn Song, · 2018
Cited alongside, same era.
“Audio adversarial examples: Targeted attacks on speech-to-text,”
Nicholas Carlini and David Wagner, · 2018
Cited alongside, same era.
“An end-to-end spoofing countermeasure for automatic speaker verification using evolving recurrent neural networks,”
Giacomo Valenti, Héctor Delgado, Massimiliano Todisco, Nicholas Evans, and Laurent Pilati, · 2018
Cited alongside, same era.
R. Taori, A. Kamsetty, B. Chu, and N. Vemuri, · 2019
Later among the works it cites.
“Adversarial attacks on spoofing countermeasures of automatic speaker verification,”
S. Liu, H. Wu, H. Lee, and H. Meng, · 2019
Later among the works it cites.
“ASVspoof 2019: Future Horizons in Spoofed and Fake Audio Detection,”
Massimiliano Todisco, Xin Wang, Ville Vestman, Md. Sahidullah, Héctor Delgado, Andreas Nautsch, Junichi Yamagishi, Nicholas Evans, Tomi H. Kinnunen, and Kong Aik Lee, · 2019
Later among the works it cites.
“Towards Achieving Robust Universal Neural Vocoding,”
Jaime Lorenzo-Trueba, Thomas Drugman, Javier Latorre, Thomas Merritt, Bartosz Putrycz, Roberto Barra-Chicote, Alexis Moinet, and Vatsal Aggarwal, · 2019
Later among the works it cites.
“The deepfake detection challenge dataset,” 2020
Brian Dolhansky, Joanna Bitton, Ben Pflaum, Jikuo Lu, Russ Howes, Menglin Wang, and Cristian Canton Ferrer, · 2020
Closest in time.
“Fakecatcher: Detection of synthetic portrait videos using biological signals,”
U. A. Ciftci, I. Demir, and L. Yin, · 2020
Closest in time.
“Adversarial attacks on gmm i-vector based speaker verification systems,”
X. Li, J. Zhong, X. Wu, J. Yu, X. Liu, and H. Meng, · 2020
Closest in time.
“A comparison of features for synthetic speech detection,”
Md Sahidullah, Tomi Kinnunen, and Cemal Hanilçi, · 2091
Closest in time.