Fetching the paper…
Reading the bibliography…
Recent advances in Text-to-Speech (TTS) and Voice-Conversion (VC) using generative Artificial Intelligence (AI) technology have made it possible to generate high-quality and realistic human-like audio.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen. 1989 · 1989
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio
Aaron Van Den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, Koray Kavukcuoglu, et al · 2016
Earlier work this paper cites.
The LJ Speech Dataset
Keith Ito and Linda Johnson. 2017 · 2017
Earlier work this paper cites.
ASVspoof: the automatic speaker verification spoofing and countermeasures challenge
Zhizheng Wu, Junichi Yamagishi, Tomi Kinnunen, Cemal Hanilçi, Mohammed Sahidullah, Aleksandr Sizov, Nicholas Evans, Massimiliano Todisco, and Hector Delgado. 2017 · 2017
Earlier work this paper cites.
Efficient neural audio synthesis. In International Conference on Machine Learning . PMLR, 2410–2419
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu. 2018 · 2018
Earlier work this paper cites.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions. In IEEE international conference on acoustics, speech and signal processing (ICASSP) . 4779–4783
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, et al · 2018
Earlier work this paper cites.
Melgan: Generative adversarial networks for conditional waveform synthesis
Kundan Kumar, Rithesh Kumar, Thibault De Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre De Brebisson, Yoshua Bengio, and Aaron C Courville. 2019 · 2019
Earlier work this paper cites.
STC antispoofing systems for the ASVspoof2019 challenge
Galina Lavrentyeva, Sergey Novoselov, Andzhukaev Tseren, Marina Volkova, Artem Gorlanov, and Alexandr Kozlov. 2019 · 2019
Earlier work this paper cites.
Neural speech synthesis with transformer network. In Proceedings of the AAAI conference on artificial intelligence , Vol. 33. 6706–6713
Naihan Li, Shujie Liu, Yanqing Liu, Sheng Zhao, and Ming Liu. 2019 · 2019
Earlier work this paper cites.
Waveglow: A flow-based generative network for speech synthesis. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 3617–3621
Ryan Prenger, Rafael Valle, and Bryan Catanzaro. 2019 · 2019
Earlier work this paper cites.
Fastspeech: Fast, robust and controllable text to speech
Yi Ren, Yangjun Ruan, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2019 · 2019
Earlier work this paper cites.
ASVspoof 2019: Future horizons in spoofed and fake audio detection
Massimiliano Todisco, Xin Wang, Ville Vestman, Md Sahidullah, Héctor Delgado, Andreas Nautsch, Junichi Yamagishi, Nicholas Evans, Tomi Kinnunen, and Kong Aik Lee. 2019 · 2019
Earlier work this paper cites.
Libritts: A corpus derived from librispeech for text-to-speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. 2019 · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Earlier work this paper cites.
Wavegrad: Estimating gradients for waveform generation
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, and William Chan. 2020 · 2020
Earlier work this paper cites.
Zero-shot multi-speaker text-to-speech with state-of-the-art neural speaker embeddings. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 6184–6188
Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang, Xin Wang, Nanxin Chen, and Junichi Yamagishi. 2020 · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020a · 2020
Earlier work this paper cites.
Diffwave: A versatile diffusion model for audio synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. 2020b · 2020
Earlier work this paper cites.
ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Xin Wang, Junichi Yamagishi, Massimiliano Todisco, Héctor Delgado, Andreas Nautsch, Nicholas Evans, Md Sahidullah, Ville Vestman, Tomi Kinnunen, Kong Aik Lee, et al · 2020
Cited alongside, same era.
Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 6199–6203
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim. 2020 · 2020
Cited alongside, same era.
Wavefake: A data set to facilitate audio deepfake detection
Joel Frank and Lea Schönherr. 2021 · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Cited alongside, same era.
Clap learning audio concepts from natural language supervision. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 1–5
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang. 2023 · 2023
Later among the works it cites.
Improved deepfake detection using whisper features
Piotr Kawa, Marcin Plata, Michał Czuba, Piotr Szymański, and Piotr Syga. 2023 · 2023
Later among the works it cites.
Prompttts 2: Describing and generating voices with text prompt
Yichong Leng, Zhifang Guo, Kai Shen, Xu Tan, Zeqian Ju, Yanqing Liu, Yufei Liu, Dongchao Yang, Leying Zhang, Kaitao Song, et al · 2023
Later among the works it cites.
Asvspoof 2021: Towards spoofed and deepfake speech detection in the wild
Xuechen Liu, Xin Wang, Md Sahidullah, Jose Patino, Héctor Delgado, Tomi Kinnunen, Massimiliano Todisco, Junichi Yamagishi, Nicholas Evans, Andreas Nautsch, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. In International Conference on Machine Learning . PMLR, 5530–5540
Jaehyeon Kim, Jungil Kong, and Juhee Son. 2021 · 2021
Cited alongside, same era.
ASVspoof 2019: spoofing countermeasures for the detection of synthesized, converted and replayed speech
Andreas Nautsch, Xin Wang, Nicholas Evans, Tomi H Kinnunen, Ville Vestman, Massimiliano Todisco, Héctor Delgado, Md Sahidullah, Junichi Yamagishi, and Kong Aik Lee. 2021 · 2021
Cited alongside, same era.
Hemlata Tak, Jee-weon Jung, Jose Patino, Madhu Kamble, Massimiliano Todisco, and Nicholas Evans. 2021a · 2021
Cited alongside, same era.
Investigating self-supervised front ends for speech spoofing countermeasures
Xin Wang and Junichi Yamagishi. 2021 · 2021
Cited alongside, same era.
Multi-band melgan: Faster waveform generation for high-quality text-to-speech. In 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 492–498
Geng Yang, Shan Yang, Kai Liu, Peng Fang, Wei Chen, and Lei Xie. 2021 · 2021
Cited alongside, same era.
Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone. In International Conference on Machine Learning . PMLR, 2709–2720
Edresson Casanova, Julian Weber, Christopher D Shulby, Arnaldo Candido Junior, Eren Gölge, and Moacir A Ponti. 2022 · 2022
Cited alongside, same era.
Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks. In IEEE international conference on acoustics, speech and signal processing (ICASSP) . 6367–6371
Jee-weon Jung, Hee-Soo Heo, Hemlata Tak, Hye-jin Shim, Joon Son Chung, Bong-Jin Lee, Ha-Jin Yu, and Nicholas Evans. 2022 · 2022
Cited alongside, same era.
Attack agnostic dataset: Towards generalization and stabilization of audio deepfake detection
Piotr Kawa, Marcin Plata, and Piotr Syga. 2022a · 2022
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision. In International conference on machine learning . PMLR, 28492–28518
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023 · 2023
Later among the works it cites.
Ai-synthesized voice detection using neural vocoder artifacts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 904–912
Chengzhe Sun, Shan Jia, Shuwei Hou, and Siwei Lyu. 2023 · 2023
Later among the works it cites.
Audiobox: Unified audio generation with natural language prompts
Apoorv Vyas, Bowen Shi, Matthew Le, Andros Tjandra, Yi-Chiao Wu, Baishan Guo, Jiemin Zhang, Xinyue Zhang, Robert Adkins, William Ngan, et al · 2023
Later among the works it cites.
Neural codec language models are zero-shot text to speech synthesizers
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al · 2023
Later among the works it cites.
Add 2023: the second audio deepfake detection challenge
Jiangyan Yi, Jianhua Tao, Ruibo Fu, Xinrui Yan, Chenglong Wang, Tao Wang, Chu Yuan Zhang, Xiaohui Zhang, Yan Zhao, Yong Ren, et al · 2023
Later among the works it cites.
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Philip Anastassiou, Jiawei Chen, Jitong Chen, Yuanzhe Chen, Zhuo Chen, Ziyi Chen, Jian Cong, Lelai Deng, Chuang Ding, Lu Gao, et al · 2024
Closest in time.
XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Edresson Casanova, Kelly Davis, Eren Gölge, Görkem Göknar, Iulian Gulea, Logan Hart, Aya Aljafari, Joshua Meyer, Reuben Morais, Samuel Olayemi, et al · 2024
Closest in time.
Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Yanqing Liu, Yichong Leng, Kaitao Song, Siliang Tang, et al · 2024
Closest in time.
Voicebox: Text-guided multilingual universal speech generation at scale
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, et al · 2024
Closest in time.
CFAD: A Chinese dataset for fake audio detection
Haoxin Ma, Jiangyan Yi, Chenglong Wang, Xinrui Yan, Jianhua Tao, Tao Wang, Shiming Wang, and Ruibo Fu. 2024 · 2024
Closest in time.
Mlaad: The multi-language audio anti-spoofing dataset
Nicolas M Müller, Piotr Kawa, Wei Herng Choong, Edresson Casanova, Eren Gölge, Thorsten Müller, Piotr Syga, Philip Sperl, and Konstantin Böttinger. 2024 · 2024
Closest in time.
The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio
Yuankun Xie, Yi Lu, Ruibo Fu, Zhengqi Wen, Zhiyong Wang, Jianhua Tao, Xin Qi, Xiaopeng Wang, Yukun Liu, Haonan Cheng, et al · 2024
Closest in time.
FlashSpeech: Efficient Zero-Shot Speech Synthesis
Zhen Ye, Zeqian Ju, Haohe Liu, Xu Tan, Jianyi Chen, Yiwen Lu, Peiwen Sun, Jiahao Pan, Weizhen Bian, Shulin He, et al · 2024
Closest in time.
Singfake: Singing voice deepfake detection. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 12156–12160
Yongyi Zang, You Zhang, Mojtaba Heydari, and Zhiyao Duan. 2024 · 2024
Closest in time.