Fetching the paper…
Reading the bibliography…
Audio embeddings are crucial tools in understanding large catalogs of music.
“The HTK book,”
S. Young, G. Evermann, M. Gales, T. Hain, D. Kershaw, X. Liu, G. Moore, J. Odell, D. Ollason, D. Povey, et al., · 2002
Earlier work this paper cites.
“Musical genre classification of audio signals,”
G. Tzanetakis and P. Cook, · 2002
Earlier work this paper cites.
“Locality pursuit embedding,”
W. Min, K. Lu, and X. He, · 2004
Earlier work this paper cites.
“An experimental comparison of audio tempo induction algorithms,”
F. Gouyon, A. Klapuri, S. Dixon, M. Alonso, G. Tzanetakis, C. Uhle, and P. Cano, · 2006
Earlier work this paper cites.
“Evaluation of algorithms using games: The case of music tagging.,”
E. Law, K. West, M. I. Mandel, M. Bay, and J. S. Downie, · 2009
Earlier work this paper cites.
“The million song dataset,”
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere, · 2011
Earlier work this paper cites.
“Streamlined tempo estimation based on autocorrelation and cross-correlation with pulses,”
G. Percival and G. Tzanetakis, · 2014
Earlier work this paper cites.
“Swing ratio estimation,”
U. Marchand and G. Peeters, · 2015
Earlier work this paper cites.
“Two data sets for tempo estimation and key detection in electronic dance music annotated from user corrections,”
P. Knees, Á. Faraldo Pérez, H. Boyer, R. Vogl, S. Böck, F. Hörschläger, M. Le Goff, et al., · 2015
Earlier work this paper cites.
“Neural audio synthesis of musical notes with wavenet autoencoders,”
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan, · 2017
Earlier work this paper cites.
“mixup: Beyond empirical risk minimization,”
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, · 2018
Earlier work this paper cites.
“Unsupervised learning of semantic audio representations,”
A. Jansen, M. Plakal, R. Pandya, D. P. Ellis, S. Hershey, J. Liu, R. C. Moore, and R. A. Saurous, · 2018
Earlier work this paper cites.
“A crowdsourced experiment for tempo estimation of electronic dance music.,”
H. Schreiber and M. Müller, · 2018
Earlier work this paper cites.
“A single-step approach to musical tempo estimation using a convolutional neural network,”
H. Schreiber and M. Müller, · 2018
Earlier work this paper cites.
“Are nearby neighbors relatives? testing deep music embeddings,”
J. Kim, J. Urbano, C. C. S. Liem, and A. Hanjalic, · 2019
Cited alongside, same era.
“Unsupervised learning of deep features for music segmentation,”
M. C. McCallum, · 2019
Cited alongside, same era.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, · 2019
Cited alongside, same era.
“The Harmonix set: Beats, downbeats, and functional segment annotations of western popular music,”
O. Nieto, M. McCallum, M. E. P. Davies, A. Robertson, A. Stark, and E. Egozy, · 2019
Cited alongside, same era.
“Multi-task learning of tempo and beat: Learning one to improve the other.,”
S. Böck, M. E. P. Davies, and P. Knees, · 2019
Cited alongside, same era.
“Metric learning vs classification for disentangled music representation learning,”
“Unsupervised contrastive learning of sound event representations,”
E. Fonseca, D. Ortego, K. McGuinness, N. E. O’Connor, and X. Serra, · 2021
Later among the works it cites.
“Towards learning universal audio representations,”
L. Wang, P. Luc, Y. Wu, A. Recasens, L. Smaira, A. Brock, A. Jaegle, J.-B. Alayrac, S. Dieleman, J. Carreira, et al., · 2022
Later among the works it cites.
“Music representation learning based on editorial metadata from discogs,”
P. Alonso-Jiménez, X. Serra, and D. Bogdanov, · 2022
Later among the works it cites.
“Supervised and unsupervised learning of audio representations for music understanding,”
M. C. McCallum, F. Korzeniowski, S. Oramas, F. Gouyon, and A. F. Ehmann, · 2022
Later among the works it cites.
“Towards proper contrastive self-supervised learning strategies for music audio representation,”
J. Choi, S. Jang, H. Cho, and S. Chung, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Lee, N. J. Bryan, J. Salamon, Z. Jin, and J. Nam, · 2020
Cited alongside, same era.
“Disentangled multidimensional metric learning for music similarity,”
J. Lee, N. J. Bryan, J. Salamon, Z. Jin, and J. Nam, · 2020
Cited alongside, same era.
“A simple framework for contrastive learning of visual representations,”
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, · 2020
Cited alongside, same era.
“Deconstruct, analyse, reconstruct: How to improve tempo, beat, and downbeat estimation,”
S. Böck and M. E. P. Davies, · 2020
Cited alongside, same era.
“Contrastive learning of general-purpose audio representations,”
A. Saeed, D. Grangier, and N. Zeghidour, · 2021
Cited alongside, same era.
“Contrastive learning of musical representations,”
J. Spijkervet and J. A. Burgoyne, · 2021
Cited alongside, same era.
“Data augmenting contrastive learning of speech representations in the time domain,”
E. Kharitonov, M. Rivière, G. Synnaeve, L. Wolf, P.-E. Mazaré, M. Douze, and E. Dupoux, · 2021
Cited alongside, same era.
D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, · 2022
Later among the works it cites.
“Contrastive learning with positive-negative frame mask for music representation,”
D. Yao, Z. Zhao, S. Zhang, J. Zhu, Y. Zhu, R. Zhang, and X. He, · 2022
Later among the works it cites.
“S3T: Self-supervised pre-training with swin transformer for music classification,”
H. Zhao, C. Zhang, B. Zhu, Z. Ma, and K. Zhang, · 2022
Later among the works it cites.
“MAST: Multiscale audio spectrogram transformers,”
S. Ghosh, A. Seth, S. Umesh, and D. Manocha, · 2022
Later among the works it cites.
“Pre-training strategies using contrastive learning and playlist information for music classification and similarity,”
P. Alonso-Jiménez, X. Favory, H. Foroughmand, G. Bourdalas, X. Serra, T. Lidy, and D. Bogdanov, · 2023
Later among the works it cites.
“MERT: Acoustic music understanding model with large-scale self-supervised training,”
Y. Li, R. Yuan, G. Zhang, Y. Ma, X. Chen, H. Yin, C. Lin, A. Ragni, E. Benetos, N. Gyenge, et al., · 2023
Later among the works it cites.
“Improving self-supervised learning for audio representations by feature diversity and decorrelation,”
B. Nguyen, S. Uhlich, and F. Cardinaux, · 2023
Later among the works it cites.
“Similar but faster: manipulation of tempo in music audio embeddings for tempo prediction and search,”
M. C. McCallum, F. Henkel, J. Kim, S. Sandberg, and M. E. P. Davies, · 2024
Closest in time.