Fetching the paper…
Reading the bibliography…
With the proliferation of Audio Language Model (ALM) based deepfake audio, there is an urgent need for generalized detection methods.
“Principles of risk minimization for learning theory,”
Vladimir Vapnik, · 1991
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Long short-term memory,”
Alex Graves and Alex Graves, · 2012
Earlier work this paper cites.
“Deep neural networks for small footprint text-dependent speaker verification,”
Ehsan Variani, Xin Lei, Erik McDermott, Ignacio Lopez Moreno, and Javier Gonzalez-Dominguez, · 2014
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Wavenet: A generative model for raw audio,”
A. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, · 2016
Earlier work this paper cites.
“Phonetic posteriorgrams for many-to-one voice conversion without parallel data training,”
Lifa Sun, Kun Li, Hao Wang, Shiyin Kang, and Helen Meng, · 2016
Earlier work this paper cites.
“Neural discrete representation learning,”
Aaron Van Den Oord, Oriol Vinyals, et al., · 2017
Earlier work this paper cites.
“Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit,”
Christophe Veaux, Junichi Yamagishi, and Kirsten MacDonald, · 2017
Earlier work this paper cites.
“Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,”
J. Shen, R. Pang, Ron J Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, et al., · 2018
Earlier work this paper cites.
“Parallel wavenet: Fast high-fidelity speech synthesis,”
Aaron Oord, Yazhe Li, Igor Babuschkin, Karen Simonyan, Oriol Vinyals, Koray Kavukcuoglu, George Driessche, Edward Lockhart, Luis Cobo, Florian Stimberg, et al., · 2018
Earlier work this paper cites.
“Neural speech synthesis with transformer network,”
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu, · 2019
Earlier work this paper cites.
“Fastspeech: Fast, robust and controllable text to speech,”
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Liu, · 2019
Earlier work this paper cites.
“Waveglow: A flow-based generative network for speech synthesis,”
R. Prenger, R. Valle, and B. Catanzaro, · 2019
Earlier work this paper cites.
“Melgan: Generative adversarial networks for conditional waveform synthesis,”
K. Kumar, Thibault Kumar, R., L. Gestin, W. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. Courville, · 2019
Earlier work this paper cites.
“Libritts: A corpus derived from librispeech for text-to-speech,”
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, · 2019
Earlier work this paper cites.
“Audiocaps: Generating captions for audios in the wild,”
Chris Dongjoo Kim, Byeongchang Kim, Hyunmin Lee, and Gunhee Kim, · 2019
Earlier work this paper cites.
“Stc antispoofing systems for the asvspoof2019 challenge,”
G. Lavrentyeva, Andzhukaev Novoselov, S., M. Volkova, A. Gorlanov, and A. Kozlov, · 2019
Earlier work this paper cites.
“ASVspoof 2019: Future Horizons in Spoofed and Fake Audio Detection,”
Massimiliano Todisco, Xin Wang, Ville Vestman, Md. Sahidullah, Héctor Delgado, Andreas Nautsch, Junichi Yamagishi, Nicholas Evans, Tomi H. Kinnunen, and Kong Aik Lee, · 2019
Earlier work this paper cites.
“Fastspeech 2: Fast and high-quality end-to-end text to speech,”
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T. Liu, · 2020
Earlier work this paper cites.
“Glow-tts: A generative flow for text-to-speech via monotonic alignment search,”
Jaehyeon Kim, Sungwon Kim, Jungil Kong, and Sungroh Yoon, · 2020
Earlier work this paper cites.
“Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,”
R. Yamamoto, E. Song, and J. Kim, · 2020
Earlier work this paper cites.
“Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,”
J. Kong, J. Kim, and J. Bae, · 2020
Earlier work this paper cites.
“Diffwave: A versatile diffusion model for audio synthesis,”
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro, · 2020
Earlier work this paper cites.
“Wavegrad: Estimating gradients for waveform generation,”
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, and William Chan, · 2020
Earlier work this paper cites.
“Seanet: A multi-modal speech enhancement network,”
Dominik Roblek, Karolis Misiunas, Marco Tagliasacchi, and Pen Li, · 2020
Earlier work this paper cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al., · 2020
Earlier work this paper cites.
Jiaqi Su, Zeyu Jin, and Adam Finkelstein, · 2020
Earlier work this paper cites.
“Neural networks fail to learn periodic functions and how to fix it,”
Liu Ziyin, Tilman Hartwig, and Masahito Ueda, · 2020
Cited alongside, same era.
“Asvspoof 2019: spoofing countermeasures for the detection of synthesized, converted and replayed speech,”
Andreas Nautsch, Xin Wang, Nicholas Evans, Tomi H Kinnunen, Ville Vestman, Massimiliano Todisco, Héctor Delgado, Md Sahidullah, Junichi Yamagishi, and Kong Aik Lee, · 2021
Cited alongside, same era.
“Grad-tts: A diffusion probabilistic model for text-to-speech,”
Vadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova, and Mikhail Kudinov, · 2021
Cited alongside, same era.
“Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,”
Jaehyeon Kim, Jungil Kong, and Juhee Son, · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
Xin Wang and Junichi Yamagishi, · 2023
Later among the works it cites.
“The ustc-nercslip system for the track 1.2 of audio deepfake detection (add 2023) challenge,”
Haochen Wu, Zhuhai Li, Luzhen Xu, Zhentao Zhang, Wenting Zhao, Bin Gu, Yang Ai, Yexin Lu, Jie Zhang, Zhenhua Ling, et al., · 2023
Later among the works it cites.
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian, · 2023
Later among the works it cites.
“Audiolm: a language modeling approach to audio generation,”
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al., · 2023
Later among the works it cites.
“Neural codec language models are zero-shot text to speech synthesizers,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Soundstream: An end-to-end neural audio codec,”
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi, · 2021
Cited alongside, same era.
“W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,”
Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu, · 2021
Cited alongside, same era.
“WaveFake: A Data Set to Facilitate Audio Deepfake Detection,”
Joel Frank and Lea Schönherr, · 2021
Cited alongside, same era.
“Upsampling artifacts in neural audio synthesis,”
Jordi Pons, Santiago Pascual, Giulio Cengarle, and Joan Serrà, · 2021
Cited alongside, same era.
“AISHELL-3: A Multi-Speaker Mandarin TTS Corpus,”
Yao Shi, Hui Bu, Xin Xu, Shaoji Zhang, and Ming Li, · 2021
Cited alongside, same era.
“Xls-r: Self-supervised cross-lingual speech representation learning at scale,”
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, et al., · 2021
Cited alongside, same era.
“Superb: Speech processing universal performance benchmark,”
Shu wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung yi Lee, · 2021
Cited alongside, same era.
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Later among the works it cites.
“Musiclm: Generating music from text,”
Andrea Agostinelli, Timo I Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, et al., · 2023
Later among the works it cites.
“Speak foreign languages with your own voice: Cross-lingual neural codec language modeling,”
Ziqiang Zhang, Long Zhou, Chengyi Wang, Sanyuan Chen, Yu Wu, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al., · 2023
Later among the works it cites.
“Viola: Unified codec language models for speech recognition, synthesis, and translation,”
Tianrui Wang, Long Zhou, Ziqiang Zhang, Yu Wu, Shujie Liu, Yashesh Gaur, Zhuo Chen, Jinyu Li, and Furu Wei, · 2023
Later among the works it cites.
“Speechx: Neural codec language model as a versatile speech transformer,”
Xiaofei Wang, Manthan Thakker, Zhuo Chen, Naoyuki Kanda, Sefik Emre Eskimez, Sanyuan Chen, Min Tang, Shujie Liu, Jinyu Li, and Takuya Yoshioka, · 2023
Later among the works it cites.
“Lauragpt: Listen, attend, understand, and regenerate audio with gpt,”
Qian Chen, Yunfei Chu, Zhifu Gao, Zerui Li, Kai Hu, Xiaohuan Zhou, Jin Xu, Ziyang Ma, Wen Wang, Siqi Zheng, et al., · 2023
Later among the works it cites.
“Audiodec: An open-source streaming high-fidelity neural audio codec,”
Yi-Chiao Wu, Israel D Gebru, Dejan Marković, and Alexander Richard, · 2023
Later among the works it cites.
“Hifi-codec: Group-residual vector quantization for high fidelity audio codec,”
Dongchao Yang, Songxiang Liu, Rongjie Huang, Jinchuan Tian, Chao Weng, and Yuexian Zou, · 2023
Later among the works it cites.
“Funcodec: A fundamental, reproducible and integrable open-source toolkit for neural speech codec,”
Zhihao Du, Shiliang Zhang, Kai Hu, and Siqi Zheng, · 2023
Later among the works it cites.
“Audiopalm: A large language model that can speak and listen,”
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al., · 2023
Later among the works it cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al., · 2023
Later among the works it cites.
“Timit-tts: A text-to-speech dataset for multimodal synthetic media detection,”
Davide Salvi, Brian Hosler, Paolo Bestagini, Matthew C Stamm, and Stefano Tubaro, · 2023
Later among the works it cites.
“Add 2023: the second audio deepfake detection challenge,”
Jiangyan Yi, Jianhua Tao, Ruibo Fu, Xinrui Yan, Chenglong Wang, Tao Wang, Chuyuan Zhang, Xiaohui Zhang, Zhao Yan, Yong Ren, Le Xu, Junzuo Zhou, Hao Gu, Zhengqi Wen, Shan Liang, Zheng Lian, and Haizhou Li, · 2023
Later among the works it cites.
“Multi-Dataset Co-Training with Sharpness-Aware Optimization for Audio Anti-spoofing,”
Hye jin Shim, Jee weon Jung, and Tomi Kinnunen, · 2023
Later among the works it cites.
“Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,”
Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Yanqing Liu, Yichong Leng, Kaitao Song, Siliang Tang, et al., · 2024
Closest in time.
“Simple and controllable music generation,”
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez, · 2024
Closest in time.
“High-fidelity audio compression with improved rvqgan,”
Rithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar, and Kundan Kumar, · 2024
Closest in time.
“Speechtokenizer: Unified speech tokenizer for speech language models,”
Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, and Xipeng Qiu, · 2024
Closest in time.
“Mlaad: The multi-language audio anti-spoofing dataset,”
Nicolas M Müller, Piotr Kawa, Wei Herng Choong, Edresson Casanova, Eren Gölge, Thorsten Müller, Piotr Syga, Philip Sperl, and Konstantin Böttinger, · 2024
Closest in time.
“Asvspoof 5: Crowdsourced speech data, deepfakes, and adversarial attacks at scale,”
Xin Wang, Hector Delgado, Hemlata Tak, Jee-weon Jung, Hye-jin Shim, Massimiliano Todisco, Ivan Kukanov, Xuechen Liu, Md Sahidullah, Tomi Kinnunen, et al., · 2024
Closest in time.
“Towards generalisable and calibrated audio deepfake detection with self-supervised representations,”
Octavian Pascu, Adriana Stan, Dan Oneata, Elisabeta Oneata, and Horia Cucu, · 2024
Closest in time.
Orchid Chetia Phukan, Gautam Siddharth Kashyap, Arun Balaji Buduru, and Rajesh Sharma, · 2024
Closest in time.
“Audio deepfake detection with self-supervised xls-r and sls classifier,”
Qishan Zhang, Shuangbing Wen, and Tao Hu, · 2024
Closest in time.