Fetching the paper…
Reading the bibliography…
Audio deepfake detection (ADD) is crucial to combat the misuse of speech synthesized from generative AI models.
Prosody in the comprehension of spoken language: A literature review
Anne Cutler, Delphine Dahan, and Wilma Van Donselaar · 1997
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C Pereira, and William Bialek · 2000
Earlier work this paper cites.
Canonical correlation analysis
David Weenink · 2003
Earlier work this paper cites.
The linguistics of speech
William A Kretzschmar · 2009
Earlier work this paper cites.
Opensmile: the munich versatile and fast open-source audio feature extractor
Florian Eyben, Martin Wöllmer, and Björn Schuller · 2010
Earlier work this paper cites.
Compute canada: advancing computational research
Susan Baldwin · 2012
Earlier work this paper cites.
Paralinguistics in speech and language—state-of-the-art and the challenge
Björn Schuller, Stefan Steidl, Anton Batliner, Felix Burkhardt, Laurence Devillers, Christian MüLler, and Shrikanth Narayanan · 2013
Earlier work this paper cites.
Sensitivity analysis of welch’st-test
Nor Aishah Ahad and Sharipah Soaad Syed Yahaya · 2014
Earlier work this paper cites.
The geneva minimalistic acoustic parameter set (gemaps) for voice research and affective computing
Florian Eyben, Klaus R Scherer, Björn W Schuller, Johan Sundberg, Elisabeth André, Carlos Busso, Laurence Y Devillers, Julien Epps, Petri Laukka, Shrikanth S Narayanan, et al · 2015
Earlier work this paper cites.
The role of language in emotion: Predictions from psychological constructionism
Kristen A Lindquist, Jennifer K MacCormack, and Holly Shablack · 2015
Earlier work this paper cites.
An overview of voice conversion systems
Seyed Hamidreza Mohammadi and Alexander Kain · 2017
Earlier work this paper cites.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C. Lipton · 2018
Earlier work this paper cites.
The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english
Steven R Livingstone and Frank A Russo · 2018
Earlier work this paper cites.
The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods
Jaime Lorenzo-Trueba, Junichi Yamagishi, Tomoki Toda, Daisuke Saito, Fernando Villavicencio, Tomi Kinnunen, and Zhenhua Ling · 2018
Earlier work this paper cites.
X-vectors: Robust dnn embeddings for speaker recognition
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur · 2018
Earlier work this paper cites.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton · 2019
Earlier work this paper cites.
Age-related differences in conversational discourse abilities a comparative study
Natalie Pereira, Ana Paula Bresolin Gonçalves, Mariana Goulart, Marina Amarante Tarrasconi, Renata Kochhann, and Rochele Paz Fonseca · 2019
Earlier work this paper cites.
Asvspoof 2019: Future horizons in spoofed and fake audio detection
Massimiliano Todisco, Xin Wang, Ville Vestman, Md Sahidullah, Héctor Delgado, Andreas Nautsch, Junichi Yamagishi, Nicholas Evans, Tomi Kinnunen, and Kong Aik Lee · 2019
Earlier work this paper cites.
CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit, 2019
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald · 2019
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber · 2020
Earlier work this paper cites.
Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador García, Sergio Gil-López, Daniel Molina, Richard Benjamins, et al · 2020
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
Recurrent convolutional structures for audio spoof and video deepfake detection
Akash Chintha, Bao Thai, Saniat Javid Sohrawardi, Kartavya Bhatt, Andrea Hickerson, Matthew Wright, and Raymond Ptucha · 2020
Earlier work this paper cites.
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck · 2020
Earlier work this paper cites.
Voice conversion challenge 2020—intra-lingual semi-parallel and cross-lingual voice conversion–
Zhao Yi, Wen-Chin Huang, Xiaohai Tian, Junichi Yamagishi, Rohan Kumar Das, Tomi Kinnunen, Zhen-Hua Ling, and Tomoki Toda · 2020
Earlier work this paper cites.
Xls-r: Self-supervised cross-lingual speech representation learning at scale
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, et al · 2021
Earlier work this paper cites.
Exploring wav2vec 2.0 on speaker verification and language identification
Zhiyun Fan, Meng Li, Shiyu Zhou, and Bo Xu · 2021
Earlier work this paper cites.
Fine-tuned XLSR-53 large model for speech recognition in English
Jonatas Grosman · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Cited alongside, same era.
Speech is silver, silence is golden: What do asvspoof-trained models really learn?
Nicolas Müller, Franziska Dieckmann, Pavel Czempin, Roman Canals, Konstantin Böttinger, and Jennifer Williams · 2021
Cited alongside, same era.
Layer-wise analysis of a self-supervised speech representation model
Ankita Pasad, Ju-Chieh Chou, and Karen Livescu · 2021
Cited alongside, same era.
Emotion recognition from speech using wav2vec 2.0 embeddings
Leonardo Pepino, Pablo Riera, and Luciana Ferrer · 2021
Cited alongside, same era.
Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild
Xuechen Liu, Xin Wang, Md Sahidullah, Jose Patino, Héctor Delgado, Tomi Kinnunen, Massimiliano Todisco, Junichi Yamagishi, Nicholas Evans, Andreas Nautsch, et al · 2023
Later among the works it cites.
Deepfakes generation and detection: State-of-the-art, open challenges, countermeasures, and way forward
Momina Masood, Mariam Nawaz, Khalid Mahmood Malik, Ali Javed, Aun Irtaza, and Hafiz Malik · 2023
Later among the works it cites.
Complex-valued neural networks for voice anti-spoofing
Nicolas M Müller, Philip Sperl, and Konstantin Böttinger · 2023
Later among the works it cites.
Comparative layer-wise analysis of self-supervised speech models
Ankita Pasad, Bowen Shi, and Karen Livescu · 2023
Later among the works it cites.
A survey of technologies for automatic dysarthric speech recognition
Zhaopeng Qian, Kejing Xiao, and Chongchong Yu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mirco Ravanelli, Titouan Parcollet, Peter Plantinga, Aku Rouhe, Samuele Cornell, Loren Lugosch, Cem Subakan, Nauman Dawalatabad, Abdelwahab Heba, Jianyuan Zhong, et al · 2021
Cited alongside, same era.
Jui Shah, Yaman Kumar Singla, Changyou Chen, and Rajiv Ratn Shah · 2021
Cited alongside, same era.
End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detection
Hemlata Tak, Jee-Weon Jung, Jose Patino, Madhu Kamble, Massimiliano Todisco, and Nicholas Evans · 2021
Cited alongside, same era.
A survey on neural speech synthesis
Xu Tan, Tao Qin, Frank Soong, and Tie-Yan Liu · 2021
Cited alongside, same era.
Investigating self-supervised front ends for speech spoofing countermeasures
Xin Wang and Junichi Yamagishi · 2021
Cited alongside, same era.
Yingzhi Wang, Abdelmoumene Boumadane, and Abdelwahab Heba · 2021
Cited alongside, same era.
A review of modern audio deepfake detection methods: Challenges and future directions
Zaynab Almutairi and Hebah Elgibreen · 2022
Cited alongside, same era.
Neural text-to-speech synthesis
Xu Tan · 2023
Later among the works it cites.
An overview of affective speech synthesis and conversion in the deep learning era
Andreas Triantafyllopoulos, Björn W Schuller, Gökçe İymen, Metin Sezgin, Xiangheng He, Zijiang Yang, Panagiotis Tzirakis, Shuo Liu, Silvan Mertes, Elisabeth André, et al · 2023
Later among the works it cites.
Spoofed training data for speech spoofing countermeasure can be efficiently created using neural vocoders
Xin Wang and Junichi Yamagishi · 2023
Later among the works it cites.
Learning a self-supervised domain-invariant feature representation for generalized audio deepfake detection
Yuankun Xie, Haonan Cheng, Yutian Wang, and Long Ye · 2023
Later among the works it cites.
Ps3dt: Synthetic speech detection using patched spectrogram transformer
Amit Kumar Singh Yadav, Ziyue Xiang, Kratika Bhagtani, Paolo Bestagini, Stefano Tubaro, and Edward J Delp · 2023
Later among the works it cites.
Seeing is not always believing: Discrepancies in saliency maps
Masahiro Yanagawa and Junya Sato · 2023
Later among the works it cites.
Audio deepfake detection: A survey
Jiangyan Yi, Chenglong Wang, Jianhua Tao, Xiaohui Zhang, Chu Yuan Zhang, and Yan Zhao · 2023
Later among the works it cites.
The impact of silence on speech anti-spoofing
Yuxiang Zhang, Zhuo Li, Jingze Lu, Hua Hua, Wenchao Wang, and Pengyuan Zhang · 2023
Later among the works it cites.
Takanori Ashihara, Marc Delcroix, Takafumi Moriya, Kohei Matsuura, Taichi Asami, and Yusuke Ijima · 2024
Closest in time.
How i broke into a bank account with an ai-generated voice
Joseph Cox · 2024
Closest in time.
A comprehensive survey on automatic speech recognition using neural networks
Amandeep Singh Dhanjal and Williamjeet Singh · 2024
Closest in time.
Audio deepfake detection with self-supervised wavlm and multi-fusion attentive classifier
Yinlin Guo, Haofan Huang, Xi Chen, He Zhao, and Yuehai Wang · 2024
Closest in time.
Frame-to-utterance convergence: A spectra-temporal approach for unified spoofing detection
Awais Khan, Khalid Mahmood Malik, and Shah Nawaz · 2024
Closest in time.
Researchers say the deepfake biden robocall was likely made with tools from ai startup elevenlabs
Kate Knibbs · 2024
Closest in time.
Every breath you don’t take: Deepfake speech detection using breath
Seth Layton, Thiago De Andrade, Daniel Olszewski, Kevin Warren, Carrie Gates, Kevin Butler, and Patrick Traynor · 2024
Closest in time.
One-class knowledge distillation for spoofing speech detection
Jingze Lu, Yuxiang Zhang, Wenchao Wang, Zengqiang Shang, and Pengyuan Zhang · 2024
Closest in time.
Mlaad: The multi-language audio anti-spoofing dataset
Nicolas M Müller, Piotr Kawa, Wei Herng Choong, Edresson Casanova, Eren Gölge, Thorsten Müller, Piotr Syga, Philip Sperl, and Konstantin Böttinger · 2024
Closest in time.
Ai threatens courts with fake evidence, uw prof says
Terry Pender · 2024
Closest in time.
Alexandra Saliba, Yuanchao Li, Ramon Sanabria, and Catherine Lai · 2024
Closest in time.
Does audio deepfake detection rely on artifacts?
Tsu-Hsien Shih, Chin-Yuan Yeh, and Ming-Syan Chen · 2024
Closest in time.
Expressivity and speech synthesis
Andreas Triantafyllopoulos and Björn W Schuller · 2024
Closest in time.
Can large-scale vocoded spoofed data improve speech spoofing countermeasure with a self-supervised front end?
Xin Wang and Junichi Yamagishi · 2024
Closest in time.
A robust audio deepfake detection system via multi-view feature
Yujie Yang, Haochen Qin, Hang Zhou, Chengcheng Wang, Tianyu Guo, Kai Han, and Yunhe Wang · 2024
Closest in time.