Fetching the paper…
Reading the bibliography…
The growing popularity of generative music models underlines the need for perceptually relevant, objective music quality metrics.
“Effectively Unbiased FID and Inception Score and where to find them,” June 2020,
M. J. Chong and D. Forsyth, · 1911
Earlier work this paper cites.
“TRILL - Towards Learning a Universal Non-Semantic Representation of Speech,”
J. Shor, A. Jansen, R. Maor, O. Lang, O. Tuval, F. d. C. Quitry, M. Tagliasacchi, I. Shavitt, D. Emanuel, and Y. Haviv, · 2002
Earlier work this paper cites.
“CNN Architectures for Large-Scale Audio Classification,” Jan. 2017,
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, M. Slaney, R. J. Weiss, and K. Wilson, · 2017
Earlier work this paper cites.
“FMA: A Dataset For Music Analysis,” Sept. 2017,
M. Defferrard, K. Benzi, P. Vandergheynst, and X. Bresson, · 2017
Earlier work this paper cites.
“The MUSDB18 corpus for music separation,” Dec. 2017
Z. Rafii, A. Liutkus, F.-R. Stöter, S. I. Mimilakis, and R. Bittner, · 2017
Earlier work this paper cites.
“Fr\’echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms,” Jan. 2019,
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, · 2019
Earlier work this paper cites.
“OpenL3 - Look, Listen, and Learn More: Design Choices for Deep Audio Embeddings,”
A. L. Cramer, H.-H. Wu, J. Salamon, and J. P. Bello, · 2019
Earlier work this paper cites.
“Panns: Large-scale pretrained audio neural networks for audio pattern recognition,”
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, · 2020
Earlier work this paper cites.
“CDPAM: Contrastive learning for perceptual audio similarity,” Feb. 2021,
P. Manocha, Z. Jin, R. Zhang, and A. Finkelstein, · 2021
Earlier work this paper cites.
“Pedalboard,” July 2021
P. Sobot, · 2021
Earlier work this paper cites.
“It’s Raw! Audio Generation with State-Space Models,” Feb. 2022,
K. Goel, A. Gu, C. Donahue, and C. Ré, · 2022
Cited alongside, same era.
“CLAP: Learning Audio Concepts From Natural Language Supervision,” June 2022,
B. Elizalde, S. Deshmukh, M. A. Ismail, and H. Wang, · 2022
Cited alongside, same era.
“MuLan: A Joint Embedding of Music Audio and Natural Language,” Aug. 2022,
Q. Huang, A. Jansen, J. Lee, R. Ganti, J. Y. Li, and D. P. W. Ellis, · 2022
Cited alongside, same era.
“Evaluating generative audio systems and their metrics,” Aug. 2022,
A. Vinay and A. Lerch, · 2022
Cited alongside, same era.
“A Study on the Evaluation of Generative Models,” June 2022,
“AudioLDM: Text-to-Audio Generation with Latent Diffusion Models,” Feb. 2023,
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, · 2023
Closest in time.
“Multi-Source Diffusion Models for Simultaneous Music Generation and Separation,” Feb. 2023,
G. Mariani, I. Tallini, E. Postolache, M. Mancusi, L. Cosmo, and E. Rodolà, · 2023
Closest in time.
“Moisesdb: A dataset for source separation beyond 4-stems,” 2023
I. Pereira, F. Araújo, F. Korzeniowski, and R. Vogl, · 2023
Closest in time.
“The Role of ImageNet Classes in Fr\’echet Inception Distance,” Feb. 2023,
T. Kynkäänniemi, T. Karras, M. Aittala, T. Aila, and J. Lehtinen, · 2023
Closest in time.
“Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Betzalel, C. Penso, A. Navon, and E. Fetaya, · 2022
Cited alongside, same era.
“High Fidelity Neural Audio Compression,” Oct. 2022,
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, · 2022
Cited alongside, same era.
“Simple and Controllable Music Generation,” June 2023,
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez, · 2023
Cited alongside, same era.
“MusicLM: Generating Music From Text,” Jan. 2023,
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, M. Sharifi, N. Zeghidour, and C. Frank, · 2023
Cited alongside, same era.
“Diffsound: Discrete Diffusion Model for Text-to-sound Generation,” Apr. 2023,
D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y. Zou, and D. Yu, · 2023
Cited alongside, same era.
“Noise2Music: Text-conditioned Music Generation with Diffusion Models,” Mar. 2023,
Q. Huang, D. S. Park, T. Wang, T. I. Denk, A. Ly, N. Chen, Z. Zhang, Z. Zhang, J. Yu, C. Frank, J. Engel, Q. V. Le, W. Chan, Z. Chen, and W. Han, · 2023
Cited alongside, same era.
Y. Wu*, K. Chen*, T. Zhang*, Y. Hui*, T. Berg-Kirkpatrick, and S. Dubnov, · 2023
Closest in time.
“MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training,” June 2023,
Y. Li, R. Yuan, G. Zhang, Y. Ma, X. Chen, H. Yin, C. Lin, A. Ragni, E. Benetos, N. Gyenge, R. Dannenberg, R. Liu, W. Chen, G. Xia, Y. Shi, W. Huang, Y. Guo, and J. Fu, · 2023
Closest in time.
“High-Fidelity Audio Compression with Improved RVQGAN,” June 2023,
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, · 2023
Closest in time.
“Mubert,” 2023,
Mubert-Inc, · 2023
Closest in time.
“Natural language supervision for general-purpose audio representations,”
B. Elizalde, S. Deshmukh, and H. Wang, · 2024
Closest in time.