Fetching the paper…
Reading the bibliography…
This paper investigates foundation models tailored for music informatics, a domain currently challenged by the scarcity of labeled data and generalization issues.
“Evaluation of algorithms using games: The case of music tagging.,”
Edith Law et al., · 2009
Earlier work this paper cites.
“Transfer learning by supervised pre-training for audio-based music classification,”
Aäron Van Den Oord et al., · 2014
Earlier work this paper cites.
“Mir_eval: A transparent implementation of common mir metrics,”
Colin Raffel et al., · 2014
Earlier work this paper cites.
“Gtzan-rhythm: Extending the gtzan test-set with beat, downbeat and swing annotations,”
Ugo Marchand et al., · 2015
Earlier work this paper cites.
“Two data sets for tempo estimation and key detection in electronic dance music annotated from user corrections,”
Peter Knees et al., · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma et al., · 2015
Earlier work this paper cites.
“Joint beat and downbeat tracking with recurrent neural networks,”
Sebastian Böck et al., · 2016
Earlier work this paper cites.
“Transfer learning for music classification and regression tasks,”
Keunwoo Choi et al., · 2017
Earlier work this paper cites.
“FMA: A dataset for music analysis,”
Michaël Defferrard et al., · 2017
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin et al., · 2018
Earlier work this paper cites.
“Genre-agnostic key classification with convolutional neural networks,”
Filip Korzeniowski et al., · 2018
Earlier work this paper cites.
“The Harmonix Set: Beats, downbeats, and functional segment annotations of western popular music,”
Oriol Nieto et al., · 2019
Earlier work this paper cites.
“A bi-directional transformer for musical chord recognition,”
Jonggwon Park et al., · 2019
Earlier work this paper cites.
“Language models are few-shot learners,”
Tom Brown et al., · 2020
Earlier work this paper cites.
“Semi-supervised learning using teacher-student models for vocal melody extraction,”
Sangeun Kum et al., · 2020
Cited alongside, same era.
“A simple framework for contrastive learning of visual representations,”
Ting Chen et al., · 2020
Cited alongside, same era.
“Jukebox: A generative model for music,”
Prafulla Dhariwal et al., · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski et al., · 2020
Cited alongside, same era.
“Pushing the limits of semi-supervised learning for automatic speech recognition,”
Yu Zhang et al., · 2020
Cited alongside, same era.
“Semi-supervised learning using teacher-student models for vocal melody extraction,”
“Semi-supervised music tagging transformer,”
Minz Won et al., · 2021
Later among the works it cites.
“Contrastive learning of musical representations,”
Janne Spijkervet et al., · 2021
Later among the works it cites.
“Codified audio language modeling learns useful representations for music information retrieval,”
Rodrigo Castellon et al., · 2021
Later among the works it cites.
“Music classification: beyond supervised learning, towards real-world applications,”
Minz Won et al., · 2021
Later among the works it cites.
“Self-supervised learning with random-projection quantizer for speech recognition,”
Chung-Cheng Chiu et al., · 2022
Later among the works it cites.
“Supervised and unsupervised learning of audio representations for music understanding,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sangeun Kum et al., · 2020
Cited alongside, same era.
“Evaluation of cnn-based automatic music tagging models,”
Minz Won et al., · 2020
Cited alongside, same era.
“Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters,”
Jeff Rasley et al., · 2020
Cited alongside, same era.
“On the opportunities and risks of foundation models,”
Rishi Bommasani et al., · 2021
Cited alongside, same era.
“An image is worth 16x16 words: Transformers for image recognition at scale,”
Alexey Dosovitskiy et al., · 2021
Cited alongside, same era.
“Vivit: A video vision transformer,”
Anurag Arnab et al., · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu et al., · 2021
Cited alongside, same era.
Matthew C McCallum et al., · 2022
Later among the works it cites.
“High fidelity neural audio compression,”
Alexandre Défossez et al., · 2022
Later among the works it cites.
“Modeling beats and downbeats with a time-frequency transformer,”
Yun-Ning Hung et al., · 2022
Later among the works it cites.
“To catch a chorus, verse, intro, or anything else: Analyzing a song with structural functions,”
Ju-Chiang Wang et al., · 2022
Later among the works it cites.
“Flashattention: Fast and memory-efficient exact attention with io-awareness,”
Tri Dao et al., · 2022
Later among the works it cites.
“Mert: Acoustic music understanding model with large-scale self-supervised training,”
Yizhi Li et al., · 2023
Closest in time.
“All-in-one metrical and functional structure analysis with neighborhood attentions on demixed audio,”
Taejun Kim et al., · 2023
Closest in time.
“Pre-training strategies using contrastive learning and playlist information for music classification and similarity,”
Pablo Alonso-Jiménez et al., · 2023
Closest in time.