Fetching the paper…
Reading the bibliography…
Text-To-Music (TTM) models have recently revolutionized the automatic music generation research field.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016, pp. 770–778
2016
Earlier work this paper cites.
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 776–780
2017
Earlier work this paper cites.
W. Dai, C. Dai, S. Qu, J. Li, and S. Das, “Very deep convolutional neural networks for raw waveforms,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 421–425
2017
Earlier work this paper cites.
N. Yu, L. Davis, and M. Fritz, “Attributing fake images to gans: Learning and analyzing gan fingerprints,” in International Conference on Computer Vision (ICCV) , 2019
2019
Earlier work this paper cites.
S. Mandelli, P. Bestagini, L. Verdoliva, and S. Tubaro, “Facing device attribution problem for stabilized video sequences,” IEEE Transactions on Information Forensics and Security (TIFS) , vol. 15, pp. 14–27, 2019
2019
Earlier work this paper cites.
J.-P. Briot, G. Hadjeres, and F.-D. Pachet, Deep learning techniques for music generation . Springer, 2020, vol. 1
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in neural information processing systems , vol. 33, pp. 17 022–17 033, 2020
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher, “End-to-end anti-spoofing with rawnet2,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021
2021
Earlier work this paper cites.
D. Salvi, P. Bestagini, and S. Tubaro, “Exploring the synthetic speech attribution problem through data-driven detectors,” in IEEE International Workshop on Information Forensics and Security (WIFS) , 2022
2022
Earlier work this paper cites.
P.-Y. Huang, H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, and C. Feichtenhofer, “Masked autoencoders that listen,” Advances in Neural Information Processing Systems , vol. 35, pp. 28 708–28 720, 2022
2022
Earlier work this paper cites.
D. Hendrycks and K. Gimpel, “A baseline for detecting misclassified and out-of-distribution examples in neural networks,” in International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” Transactions on Machine Learning Research , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “Audioldm: Text-to-audio generation with latent diffusion models,” in International Conference on Machine Learning . PMLR, 2023
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
M. Feffer, Z. C. Lipton, and C. Donahue, “Deepdrake ft. bts-gan and taylorvc: An exploratory analysis of musical deepfakes and hosting platforms,” in HCMIR@ ISMIR , 2023
2023
Cited alongside, same era.
Z. Sha, Z. Li, N. Yu, and Y. Zhang, “De-fake: Detection and attribution of fake images generated by text-to-image generation models,” in Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security , 2023, pp. 3418–3432
2023
Cited alongside, same era.
R. Corvi, D. Cozzolino, G. Zingarini, G. Poggi, K. Nagano, and L. Verdoliva, “On the detection of synthetic images generated by diffusion models,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
“Suno — suno.com,” https://suno.com/ , [Accessed 12-09-2024]
2024
Closest in time.
“Udio — AI Music Generator - Official Website — udio.com,” https://www.udio.com/ , [Accessed 12-09-2024]
2024
Closest in time.
L. Abady, J. Wang, B. Tondi, and M. Barni, “A siamese-based verification system for open-set architecture attribution of synthetic images,” Pattern Recognition Letters , vol. 180, pp. 75–81, 2024
2024
Closest in time.
A. Wißmann, S. Zeiler, R. M. Nickel, and D. Kolossa, “Whodunit: Detection and attribution of synthetic images by leveraging model-specific fingerprints,” in ACM International Workshop on Multimedia AI against Disinformation (MAD) , 2024
2024
Closest in time.
H. Wu, Y. Tseng, and H.-y. Lee, “Codecfake: Enhancing anti-spoofing models against deepfake audios from codec-based speech synthesis systems,” in Interspeech , 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
K. Bhagtani, E. R. Bartusiak, A. K. S. Yadav, P. Bestagini, and E. J. Delp, “Synthesized speech attribution using the patchout spectrogram attribution transformer,” in ACM Workshop on Information Hiding and Multimedia Security (IH&MMSec) , 2023
2023
Cited alongside, same era.
I. Manco, B. Weck, S. Doh, M. Won, Y. Zhang, D. Bogdanov, Y. Wu, K. Chen, P. Tovstogan, E. Benetos et al. , “The song describer dataset: a corpus of audio captions for music-and-language evaluation,” in Machine Learning for Audio Workshop at NeurIPS , 2023
2023
Cited alongside, same era.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Cited alongside, same era.
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity audio compression with improved rvqgan,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, “Simple and controllable music generation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
H. Liu, Y. Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y. Wang, W. Wang, Y. Wang, and M. D. Plumbley, “Audioldm 2: Learning holistic audio generation with self-supervised pretraining,” IEEE/ACM Trans. Audio Speech Lang. Process , 2024
2024
Cited alongside, same era.
2024
Closest in time.
Y. Zang, Y. Zhang, M. Heydari, and Z. Duan, “Singfake: Singing voice deepfake detection,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 12 156–12 160
2024
Closest in time.
Y. Xie, J. Zhou, X. Lu, Z. Jiang, Y. Yang, H. Cheng, and L. Ye, “Fsd: An initial chinese dataset for fake song detection,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 4605–4609
2024
Closest in time.
X. Chen, H. Wu, J.-S. R. Jang, and H.-y. Lee, “Singing voice graph modeling for singfake detection,” in Interspeech , 2024
2024
Closest in time.
A. Guragain, T. Liu, Z. Pan, H. B. Sailor, and Q. Wang, “Speech foundation model ensembles for the controlled singing voice deepfake detection (ctrsvdd) challenge 2024,” in 2024 IEEE Spoken Language Technology Workshop . IEEE, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
M. Civit, V. Drai-Zerbib, D. Lizcano, and M. J. Escalona, “Sunocaps: A novel dataset of text-prompt based ai-generated music with emotion annotations,” Data in Brief , vol. 55, p. 110743, 2024
2024
Closest in time.
2024
Closest in time.