Fetching the paper…
Reading the bibliography…
Music two-tower multimodal systems integrate audio and text modalities into a joint audio-text space, enabling direct comparison between songs and their corresponding labels.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 1877–1901
1901
Earlier work this paper cites.
J. C. Lena and R. A. Peterson, “Classification as culture: Types and trajectories of music genres,” American Sociological Review , vol. 73, no. 5, pp. 697–718, 2008
2008
Earlier work this paper cites.
E. Law, K. West, M. I. Mandel, M. Bay, and J. S. Downie, “Evaluation of algorithms using games: The case of music tagging,” in International Society for Music Information Retrieval Conference , 2009
2009
Earlier work this paper cites.
2009
Earlier work this paper cites.
Z. Fu, G. Lu, K. M. Ting, and D. Zhang, “A survey of audio-based music classification and annotation,” IEEE Transactions on Multimedia , vol. 13, no. 2, pp. 303–319, 2011
2011
Earlier work this paper cites.
Şefki Kolozali, M. Barthet, G. Fazekas, and M. B. Sandler, “Knowledge representation issues in musical instrument ontology design,” in International Society for Music Information Retrieval Conference , 2011
2011
Earlier work this paper cites.
T. Mikolov, K. Chen, G. S. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” in International Conference on Learning Representations , 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. Manning, “GloVe: Global vectors for word representation,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) , A. Moschitti, B. Pang, and W. Daelemans, Eds. Doha, Qatar: Association for Computational Linguistics, Oct. 2014, pp. 1532–1543
2014
Earlier work this paper cites.
C. C. Aggarwal, Data Classification: Algorithms and Applications , 1st ed. Chapman & Hall/CRC, 2014
2014
Earlier work this paper cites.
K. He, Y. Wang, and J. Hopcroft, “A powerful generative model using random weights for the deep image representation,” in Proceedings of the 30th International Conference on Neural Information Processing Systems , ser. NIPS’16. Red Hook, NY, USA: Curran Associates Inc., 2016, p. 631–639
2016
Earlier work this paper cites.
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
A. van Venrooij and V. Schmutz, “Categorical ambiguity in cultural fields: The effects of genre fuzziness in popular music,” Poetics , vol. 66, pp. 1–18, 2018
2018
Earlier work this paper cites.
R. Hennequin, J. Royo-Letelier, and M. Moussallam, “Audio based disambiguation of music genre tags,” in International Society for Music Information Retrieval Conference , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley, “PANNs: Large-scale pretrained audio neural networks for audio pattern recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 2880–2894, 2019
2019
Earlier work this paper cites.
J. Choi, J. Lee, J. Park, and J. Nam, “Zero-shot learning for audio-based music classification and tagging,” in International Society for Music Information Retrieval Conference , 2019
2019
Earlier work this paper cites.
H. Xie and V. Tuomas, “Zero-shot audio classification based on class label embeddings,” 2019 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) , pp. 264–267, 2019
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in North American Chapter of the Association for Computational Linguistics , 2019
2019
Cited alongside, same era.
Y. Tian, D. Krishnan, and P. Isola, “Contrastive multiview coding,” in European Conference on Computer Vision , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “HTS-AT: A hierarchical token-semantic audio transformer for sound classification and detection,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP , 2022
2022
Later among the works it cites.
I. Manco, E. Benetos, E. Quinton, and G. Fazekas, “Contrastive audio-language learning for music,” in Proceedings of the 23rd International Society for Music Information Retrieval Conference (ISMIR) , 2022
2022
Later among the works it cites.
Q. Huang, A. Jansen, J. Lee, R. Ganti, J. Y. Li, and D. P. W. Ellis, “MuLan: A joint embedding of music audio and natural language,” in International Society for Music Information Retrieval Conference , 2022
2022
Later among the works it cites.
D. Dogan, H. Xie, T. Heittola, and T. Virtanen, “Zero-shot audio classification using image embeddings,” 2022 30th European Signal Processing Conference (EUSIPCO) , pp. 1–5, 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Zhen, P. Hu, X. Wang, and D. Peng, “Deep supervised cross-modal retrieval,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 10 386–10 395
2019
Cited alongside, same era.
H. Xie, O. J. Räsänen, and T. Virtanen, “Zero-shot audio classification with factored linear and nonlinear acoustic-semantic projections,” ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 326–330, 2020
2020
Cited alongside, same era.
H. Xie and T. Virtanen, “Zero-shot audio classification via semantic embeddings,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 1233–1242, 2020
2020
Cited alongside, same era.
C. Emanuele, D. Ghisi, V. Lostanlen, F. Lévy, J. Fineberg, and Y. Maresz, “TinySOL: An audio dataset of isolated musical notes (5.0),” 2020
2020
Cited alongside, same era.
M. Won, A. Ferraro, D. Bogdanov, and X. Serra, “Evaluation of cnn-based automatic music tagging models,” in Proc. of 17th Sound and Music Computing , 2020
2020
Cited alongside, same era.
M. Alian and A. Awajan, “Factors affecting sentence similarity and paraphrasing identification,” Int. J. Speech Technol. , vol. 23, no. 4, p. 851–859, dec 2020
2020
Cited alongside, same era.
N. Ndou, R. Ajoodha, and A. Jadhav, “Music genre classification: A review of deep-learning and traditional machine-learning approaches,” in 2021 IEEE International IOT, Electronics and Mechatronics Conference (IEMTRONICS) , 2021, pp. 1–6
2021
Cited alongside, same era.
M. Won, J. Spijkervet, and K. Choi, Music Classification: Beyond Supervised Learning, Towards Real-world Applications . https://music-classification.github.io/tutorial, November 2021
2021
Cited alongside, same era.
X. Wang, B. Ke, X. Li, F. Liu, M. Zhang, X. Liang, and Q. Xiao, “Modality-balanced embedding for video retrieval,” in Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , ser. SIGIR ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 2578–2582
2022
Later among the works it cites.
X. Wang, B. Ke, X. Li, F. Liu, M. Zhang, X. Liang, Q.-E. Xiao, and Y. Yu, “Modality-balanced embedding for video retrieval,” Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , 2022
2022
Later among the works it cites.
D. Oneaţă and H. Cucu, “Improving multimodal speech recognition by data augmentation and speech representations,” 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , pp. 4578–4587, 2022
2022
Later among the works it cites.
J. Choi, S. Jang, H. Cho et al. , “Towards proper contrastive self-supervised learning strategies for music audio representation,” in 2022 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 2022, pp. 1–6
2022
Later among the works it cites.
S. Pratt, I. Covert, R. Liu, and A. Farhadi, “What does a platypus look like? generating customized prompts for zero-shot image classification,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , October 2023, pp. 15 691–15 701
2023
Later among the works it cites.
Y. Ge, J. Ren, A. Gallagher, Y. Wang, M.-H. Yang, H. Adam, L. Itti, B. Lakshminarayanan, and J. Zhao, “Improving zero-shot generalization and robustness of multi-modal models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2023, pp. 11 093–11 101
2023
Later among the works it cites.
S. Goyal, A. Kumar, S. Garg, Z. Kolter, and A. Raghunathan, “Finetune like you pretrain: Improved finetuning of zero-shot vision models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2023, pp. 19 338–19 347
2023
Later among the works it cites.
B. Elizalde, S. Deshmukh, M. A. Ismail, and H. Wang, “CLAP: learning audio concepts from natural language supervision,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Doh, M. Won, K. Choi, and J. Nam, “Toward universal text-to-music retrieval,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2023
2023
Later among the works it cites.
M. J. Mirza, L. Karlinsky, W. Lin, H. Possegger, M. Kozinski, R. Feris, and H. Bischof, “Lafter: Label-free tuning of zero-shot classifier using language and unlabeled image collections,” in Advances in Neural Information Processing Systems , A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curran Associates, Inc., 2023, pp. 5765–5777
2023
Later among the works it cites.
X. Du, Z. Yu, J. Lin, B. Zhu, and Q. Kong, “Joint music and language attention models for zero-shot music tagging,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 1126–1130
2024
Closest in time.