Fetching the paper…
Reading the bibliography…
As one of the most intuitive interfaces known to humans, natural language has the potential to mediate many tasks that involve human-computer interaction, especially in application-focused fields like Music Information Retrieval.
J. Bromley, I. Guyon, Y. Lecun, E. Sickinger, R. Shah, A. Bell, and L. Holmdel, “Signature verification using a" siamese" time delay neural network,” Advances in neural information processing systems , vol. 6, 1993
1993
Earlier work this paper cites.
A. Ghias, J. Logan, D. Chamberlin, and B. C. Smith, “Query by humming,” in Proceedings of the third ACM international conference on Multimedia . Association for Computing Machinery (ACM), 1995, pp. 231–236
1995
Earlier work this paper cites.
B. Whitman and R. Rifkin, “Musical query-by-description as a multiclass learning problem,” in Proceedings of 2002 IEEE Workshop on Multimedia Signal Processing, MMSP 2002 . Institute of Electrical and Electronics Engineers Inc., 2002, pp. 153–156
2002
Earlier work this paper cites.
G. Tzanetakis and P. Cook, “Musical genre classification of audio signals,” IEEE Transactions on Speech and Audio Processing , vol. 10, no. 5, pp. 293–302, 7 2002
2002
Earlier work this paper cites.
J. Ha, L. J. Stephen, D. Sally, and J. Cunningham, “Challenges in cross-cultural/multilingual music information seeking.” in ISMIR , 2005
2005
Earlier work this paper cites.
M. Müller, F. Kurth, D. Damm, C. Fremerey, and M. Clausen, “Lyrics-Based Audio Retrieval and Multimodal Navigation in Music Collections,” Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) , vol. 4675 LNCS, pp. 112–123, 2007
2007
Earlier work this paper cites.
P. Knees, “Search & Select - Intuitively Retrieving Music from Large Collections,” in ISMIR , 2007
2007
Earlier work this paper cites.
E. Law, K. West, M. Mandel, M. Bay, and J. Stephen Downie, “Evaluation of algorithms using games: The case of music tagging,” in Proceedings of the 10th ISMIR Conference , 2009
2009
Earlier work this paper cites.
C. C. Liem, M. Müller, D. Eck, G. Tzanetakis, and A. Hanjalic, “The need for music information retrieval with user-centered and multimodal strategies,” in MM’11 - Proceedings of the 2011 ACM Multimedia Conference and Co-Located Workshops - MIRUM 2011 Workshop, MIRUM’11 , 2011, pp. 1–6
2011
Earlier work this paper cites.
A. van den Oord, S. Dieleman, and B. Schrauwen, “Deep content-based music recommendation,” in Advances in Neural Information Processing Systems , 2013
2013
Earlier work this paper cites.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” in 1st International Conference on Learning Representations, ICLR 2013 - Workshop Track Proceedings . International Conference on Learning Representations, ICLR, 2013
2013
Earlier work this paper cites.
S. Oramas, M. Sordo, L. Espinosa-Anke, and X. Serra, “A semantic-based approach for artist similarity,” in ISMIR , 2015
2015
Earlier work this paper cites.
S. Oramas, L. Espinosa-Anke, M. Sordo, H. Saggion, and X. Serra, “Information extraction for knowledge base construction in the music domain,” Data and Knowledge Engineering , 2016
2016
Earlier work this paper cites.
S. Oramas, L. Espinosa-Anke, A. Lawlor, X. Serra, and H. Saggion, “Exploring Customer Reviews for Music Genre Classification and Evolutionary Studies,” in 17th International Society for Music Information Retrieval Conference , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
K. He, Y. Wang, and J. Hopcroft, “A powerful generative model using random weights for the deep image representation,” in Advances in Neural Information Processing Systems , 2016
2016
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural Machine Translation of Rare Words with Subword Units,” in 54th Annual Meeting of the Association for Computational Linguistics, ACL 2016 - Long Papers , vol. 3. Association for Computational Linguistics (ACL), 2016, pp. 1715–1725
2016
Earlier work this paper cites.
K. Tsukuda, “Lyric Jumper: A Lyrics-Based Music Exploratory Web Service by Modeling Lyrics Generative Process.” in ISMIR , 2017, pp. 544–551
2017
Earlier work this paper cites.
B. Jeon, C. Kim, A. Kim, D. Kim, J. Park, and J. W. Ha, “Music emotion recognition via end-To-end multimodal neural networks,” in CEUR Workshop Proceedings , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, A. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , vol. 2017-Decem. Neural information processing systems foundation, 6 2017, pp. 5999–6009
2017
Earlier work this paper cites.
S. Oramas, F. Barbieri, O. Nieto, and X. Serra, “Multimodal Deep Learning for Music Genre Classification,” Transactions of the International Society for Music Information Retrieval , vol. 1, no. 1, pp. 4–21, 9 2018
2018
Earlier work this paper cites.
R. Delbouys, R. Hennequin, F. Piccoli, J. Royo-Letelier, and M. Moussallam, “Music mood detection based on audio and lyrics with deep neural net,” in Proceedings of the 19th International Society for Music Information Retrieval Conference, ISMIR 2018 , 2018
2018
Earlier work this paper cites.
J. Andreas, D. Klein, and S. Levine, “Learning with Latent Language,” in NAACL HLT 2018 - 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , vol. 1. Association for Computational Linguistics (ACL), 11 2018, pp. 2166–2179
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
B. Li and A. Kumar, “Query by Video: Cross-Modal Music Retrieval,” in 20th International Society for Music Information Retrieval Conference (ISMIR) , 2019, pp. 604–611
2019
Cited alongside, same era.
M. Müller, A. Arzt, S. Balke, M. Dorfer, and G. Widmer, “Cross-Modal Music Retrieval and Applications: An Overview of Key Methodologies,” IEEE Signal Processing Magazine , vol. 36, no. 1, pp. 52–62, 1 2019
2019
Cited alongside, same era.
K. Watanabe and M. Goto, “Query-by-Blending: a Music Exploration System Blending Latent Vector Representations of Lyric Word, Song Audio, and Artist,” in Proceedings of the 20th International Society for Music Information Retrieval Conference , 2019
2019
Cited alongside, same era.
C. Hosey, L. Vujović, B. St. Thomas, J. Garcia-Gathright, and J. Thom, “Just give me what I want: How people use and evaluate music search,” in Conference on Human Factors in Computing Systems - Proceedings . Association for Computing Machinery, 5 2019
2019
M. Won, S. Oramas, O. Nieto, F. Gouyon, and X. Serra, “Multimodal Metric Learning for Tag-based Music Retrieval,” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021
2021
Later among the works it cites.
M. Won, J. Salamon, N. J. Bryan, G. J. Mysore, and X. Serra, “Emotion Embedding Spaces for Matching Music to Stories,” in International Society for Music Information Retrieval Conference (ISMIR) , 11 2021
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning Transferable Visual Models From Natural Language Supervision,” in International Conference on Machine Learning . PMLR, 2021
2021
Later among the works it cites.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig, “Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision,” in International Conference on Machine Learning . PMLR, 2021, pp. 4904–4916
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
J. Choi, J. Lee, J. Park, and J. Nam, “Zero-shot learning for audio-based music classification and tagging,” in Proceedings of the 20th International Society for Music Information Retrieval Conference, ISMIR 2019 , 2019
2019
Cited alongside, same era.
J. Lu, D. Batra, D. Parikh, and S. Lee, “ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks,” in Advances in Neural Information Processing Systems , 2019, pp. 13–23
2019
Cited alongside, same era.
N. Saunshi, P. Orestis, S. Arora, K. Mikhail, and K. Hrishikesh, “A Theoretical Analysis of Contrastive Unsupervised Representation Learning,” in International Conference on Machine Learning . PMLR, 2019
2019
Cited alongside, same era.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” 2019
2019
Cited alongside, same era.
N. Reimers and I. Gurevych, “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,” in EMNLP-IJCNLP 2019 - 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing, Proceedings of the Conference . Association for Computational Linguistics, 8 2019, pp. 3982–3992
2019
Cited alongside, same era.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Association for Computational Linguistics (ACL), 2019, p. 119–132
2019
Cited alongside, same era.
J. Lee, N. J. Bryan, J. Salamon, Z. Jin, and J. Nam, “Disentangled Multidimensional Metric Learning for Music Similarity,” in ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing , 8 2020
2020
Cited alongside, same era.
P. Knees, M. Schedl, and M. Goto, “Intelligent User Interfaces for Music Discovery,” Transactions of the International Society for Music Information Retrieval , vol. 3, no. 1, pp. 165–179, 10 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
A.-M. Oncescu, A. S. Koepke, J. F. Henriques, Z. Akata, and S. Albanie, “Audio Retrieval with Natural Language Queries,” in Interspeech , 5 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation,” in 2021 International Joint Conference on Neural Networks (IJCNN) . IEEE, 3 2021
2021
Later among the works it cites.
A. Saeed, D. Grangier, and N. Zeghidour, “Contrastive Learning of General-Purpose Audio Representations,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 10 2021
2021
Later among the works it cites.
H. Al-Tahan, Y. M. I. C. On, and U. 2021, “CLAR: Contrastive Learning of Auditory Representations,” in International Conference on Artificial Intelligence and Statistics . PMLR, 2021
2021
Later among the works it cites.
J. Spijkervet and J. A. Burgoyne, “Contrastive Learning of Musical Representations,” in ISMIR , 2021
2021
Later among the works it cites.
E. Fonseca, D. Ortego, K. McGuinness, N. E. O’Connor, and X. Serra, “Unsupervised Contrastive Learning of Sound Event Representations,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 11 2021
2021
Later among the works it cites.
M. Zolfaghari, Y. Zhu, P. Gehler, and T. Brox, “CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video Representations,” in International Conference on Computer Vision (ICCV) . Institute of Electrical and Electronics Engineers (IEEE), 9 2021, pp. 1430–1439
2021
Later among the works it cites.
L. Ericsson, H. Gouk, C. C. Loy, and T. M. Hospedales, “Self-Supervised Representation Learning: Introduction, Advances and Challenges,” IEEE Signal Processing Magazine , vol. 39, no. 3, pp. 42–62, 10 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Won, K. Choi, and X. Serra, “Semi-Supervised Music Tagging Transformer,” in International Society for Music Information Retrieval Conference (ISMIR) , 11 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
L. A. Hendricks, J. Mellor, R. Schneider, J.-B. Alayrac, and A. Nematzadeh, “Decoupling the Role of Data, Attention, and Losses in Multimodal Transformers,” Transactions of the Association for Computational Linguistics , vol. 9, pp. 570–585, 1 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
I. Manco, E. Benetos, E. Quinton, and G. Fazekas, “Learning music audio representations via weak language supervision,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022
2022
Closest in time.
2022
Closest in time.
H.-H. Wu, P. Seetharaman, K. Kumar, and J. P. Bello, “Wav2CLIP: Learning Robust Audio Representations From CLIP,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 10 2022
2022
Closest in time.
A. Guzhov, F. Raue, J. Hees, and A. Dengel, “AudioCLIP: Extending CLIP to Image, Text and Audio,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022
2022
Closest in time.
H. Xie, O. Räsänen, K. Drossos, and T. Virtanen, “Unsupervised Audio-Caption Aligning Learns Correspondences between Individual Sound Events and Textual Phrases,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022
2022
Closest in time.
L. Wang, P. Luc, Y. Wu, A. Recasens, L. Smaira, A. Brock, A. Jaegle, J.-B. Alayrac, S. Dieleman, J. Carreira, and A. van den Oord, “Towards Learning Universal Audio Representations,” in ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5 2022, pp. 4593–4597
2022
Closest in time.