Fetching the paper…
Reading the bibliography…
The recent surge in AI-generated songs presents exciting possibilities and challenges.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
I Loshchilov · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Hongyi Zhang · 2017
Earlier work this paper cites.
Stc antispoofing systems for the asvspoof2019 challenge
Galina Lavrentyeva, Sergey Novoselov, Andzhukaev Tseren, Marina Volkova, Artem Gorlanov, and Alexandr Kozlov · 2019
Earlier work this paper cites.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Earlier work this paper cites.
Wildmix dataset and spectro-temporal transformer model for monoaural audio source separation
Amir Zadeh, Tianjun Ma, Soujanya Poria, and Louis-Philippe Morency · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al · 2020
Earlier work this paper cites.
Voice conversion challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion
Yi Zhao, Wen-Chin Huang, Xiaohai Tian, Junichi Yamagishi, Rohan Kumar Das, Tomi Kinnunen, Zhenhua Ling, and Tomoki Toda · 2020
Earlier work this paper cites.
End-to-end speaker segmentation for overlap-aware resegmentation
Hervé Bredin and Antoine Laurent · 2021
Earlier work this paper cites.
AST: Audio Spectrogram Transformer
Yuan Gong, Yu-An Chung, and James Glass · 2021
Earlier work this paper cites.
Efficient training of visual transformers with small datasets
Yahui Liu, Enver Sangineto, Wei Bi, Nicu Sebe, Bruno Lepri, and Marco Nadai · 2021
Earlier work this paper cites.
End-to-end anti-spoofing with rawnet2
Hemlata Tak, Jose Patino, Massimiliano Todisco, Andreas Nautsch, Nicholas Evans, and Anthony Larcher · 2021
Earlier work this paper cites.
Prosody and voice factorization for few-shot speaker adaptation in the challenge m2voc 2021
Tao Wang, Ruibo Fu, Jiangyan Yi, Jianhua Tao, Zhengqi Wen, Chunyu Qiang, and Shiming Wang · 2021
Earlier work this paper cites.
Fake speech detection using residual network with transformer encoder
Zhenyu Zhang, Xiaowei Yi, and Xianfeng Zhao · 2021
Earlier work this paper cites.
Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks
Jee-weon Jung, Hee-Soo Heo, Hemlata Tak, Hye-jin Shim, Joon Son Chung, Bong-Jin Lee, Ha-Jin Yu, and Nicholas Evans · 2022
Cited alongside, same era.
Hemlata Tak, Massimiliano Todisco, Xin Wang, Jee-weon Jung, Junichi Yamagishi, and Nicholas Evans · 2022
Cited alongside, same era.
Single and multi-speaker cloned voice detection: from perceptual to learned features
Sarah Barrington, Romit Barua, Gautham Koorma, and Hany Farid · 2023
Cited alongside, same era.
Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction
Han Cai, Junyan Li, Muyan Hu, Chuang Gan, and Song Han · 2023
Cited alongside, same era.
Low quality deepfake detection via unseen artifacts
Saheb Chhabra, Kartik Thakral, Surbhi Mittal, Mayank Vatsa, and Richa Singh · 2023
Cited alongside, same era.
Singing voice graph modeling for singfake detection
Xuanjun Chen, Haibin Wu, Jyh-Shing Roger Jang, and Hung-yi Lee · 2024
Closest in time.
Chatbot arena: An open platform for evaluating llms by human preference, 2024
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica · 2024
Closest in time.
fvcore: A light-weight core library that provides the most common and essential functionality shared in various computer vision frameworks developed at facebook ai research
FAIR · 2024
Closest in time.
Suno-api
gcui art · 2024
Closest in time.
Genius song lyrics with language information
Carlos G. D. C. J · 2024
Closest in time.
Investigating end-to-end asr architectures for long form audio transcription
Nithin Rao Koluguri, Samuel Kriman, Georgy Zelenfroind, Somshubra Majumdar, Dima Rekesh, Vahid Noroozi, Jagadeesh Balam, and Boris Ginsburg · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models
Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou · 2023
Cited alongside, same era.
Hozier would consider strike over ai threat to music
Victoria Derbyshire, Ellie Jacobs, and Tim Dodd · 2023
Cited alongside, same era.
Online detection of ai-generated images
David C Epstein, Ishan Jain, Oliver Wang, and Richard Zhang · 2023
Cited alongside, same era.
Self-supervised representations for singing voice conversion
Tejas Jayashankar, Jilong Wu, Leda Sari, David Kant, Vimal Manohar, and Qing He · 2023
Cited alongside, same era.
Improved deepfake detection using whisper features
Piotr Kawa, Marcin Plata, Michał Czuba, Piotr Szymański, and Piotr Syga · 2023
Cited alongside, same era.
Camel: Communicative agents for” mind” exploration of large language model society
Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem · 2023
Cited alongside, same era.
Towards universal fake image detectors that generalize across generative models
Utkarsh Ojha, Yuheng Li, and Yong Jae Lee · 2023
Cited alongside, same era.
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Billie eilish and nicki minaj want stop to ’predatory’ music ai
Liv McMahon · 2024
Closest in time.
Masked modeling duo: Towards a universal audio pre-training framework
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, and Kunio Kashino · 2024
Closest in time.
Udiowrapper
Marc Riera · 2024
Closest in time.
Cst-former: Transformer with channel-spectro-temporal attention for sound event localization and detection
Yusun Shul and Jung-Woo Choi · 2024
Closest in time.
Journeybench: A challenging one-stop vision-language understanding benchmark of generated images
Zhecan Wang, Junzhang Liu, Chia-Wei Tang, Hani Alomari, Anushka Sivakumar, Rui Sun, Wenhao Li, Md Atabuzzaman, Hammad Ayyubi, Haoxuan You, et al · 2024
Closest in time.
Pytorch image models
Ross Wightman · 2024
Closest in time.
Fsd: An initial chinese dataset for fake song detection
Yuankun Xie, Jingjing Zhou, Xiaolin Lu, Zhenghao Jiang, Yuxin Yang, Haonan Cheng, and Long Ye · 2024
Closest in time.
Transcending forgery specificity with latent space augmentation for generalizable deepfake detection
Zhiyuan Yan, Yuhao Luo, Siwei Lyu, Qingshan Liu, and Baoyuan Wu · 2024
Closest in time.
Gpt4tools: Teaching large language model to use tools via self-instruction
Rui Yang, Lin Song, Yanwei Li, Sijie Zhao, Yixiao Ge, Xiu Li, and Ying Shan · 2024
Closest in time.
Singfake: Singing voice deepfake detection
Yongyi Zang, You Zhang, Mojtaba Heydari, and Zhiyao Duan · 2024
Closest in time.