Fetching the paper…
Reading the bibliography…
Self-supervised learning (SSL) has recently emerged as a promising paradigm for training generalisable models on large-scale data in the fields of vision, text, and speech.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
musicnn: Pre-trained convolutional neural networks for music audio tagging
Pons, J. and Serra, X. (2019) · 1909
Earlier work this paper cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
Baevski, A., Schneider, S., and Auli, M. (2019) · 1910
Earlier work this paper cites.
Deep learning for symbolic mathematics
Lample, G. and Charton, F. (2019) · 1912
Earlier work this paper cites.
Calculation of a constant q spectral transform
Brown, J. C. (1991) · 1991
Earlier work this paper cites.
Musical genre classification of audio signals
Tzanetakis, G. and Cook, P. (2002) · 2002
Earlier work this paper cites.
Jukebox: A generative model for music
Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I. (2020) · 2005
Earlier work this paper cites.
Evaluation of algorithms using games: The case of music tagging
Law, E., West, K., Mandel, M. I., Bay, M., and Downie, J. S. (2009) · 2009
Earlier work this paper cites.
The million song dataset
Bertin-Mahieux, T., Ellis, D. P., Whitman, B., and Lamere, P. (2011) · 2011
Earlier work this paper cites.
1000 songs for emotional analysis of music
Soleymani, M., Caro, M. N., Schmidt, E. M., Sha, C.-Y., and Yang, Y.-H. (2013) · 2013
Earlier work this paper cites.
Mir_eval: A transparent implementation of common mir metrics
Raffel, C., McFee, B., Humphrey, E. J., Salamon, J., Nieto, O., Liang, D., Ellis, D. P., and Raffel, C. C. (2014) · 2014
Earlier work this paper cites.
Two data sets for tempo estimation and key detection in electronic dance music annotated from user corrections
Knees, P., Faraldo Pérez, Á., Boyer, H., Vogl, R., Böck, S., Hörschläger, F., Le Goff, M., et al. (2015) · 2015
Earlier work this paper cites.
Swing ratio estimation
Marchand, U. and Peeters, G. (2015) · 2015
Earlier work this paper cites.
Transfer learning for music classification and regression tasks
Choi, K., Fazekas, G., Sandler, M., and Cho, K. (2017) · 2017
Earlier work this paper cites.
Neural audio synthesis of musical notes with wavenet autoencoders
Engel, J., Resnick, C., Roberts, A., Dieleman, S., Norouzi, M., Eck, D., and Simonyan, K. (2017) · 2017
Earlier work this paper cites.
End-to-end musical key estimation using a convolutional neural network
Korzeniowski, F. and Widmer, G. (2017) · 2017
Earlier work this paper cites.
The MUSDB18 corpus for music separation
Rafii, Z., Liutkus, A., Stöter, F.-R., Mimilakis, S. I., and Bittner, R. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Deep clustering for unsupervised learning of visual features
Caron, M., Bojanowski, P., Joulin, A., and Douze, M. (2018) · 2018
Earlier work this paper cites.
Vocalset: A singing voice dataset
Wilkins, J., Seetharaman, P., Wahl, A., and Pardo, B. (2018) · 2018
Cited alongside, same era.
The mtg-jamendo dataset for automatic music tagging
Bogdanov, D., Won, M., Tovstogan, P., Porter, A., and Serra, X. (2019) · 2019
Cited alongside, same era.
Data usage in mir: history & future recommendations
Chen, W., Keast, J., Moody, J., Moriarty, C., Villalobos, F., Winter, V., Zhang, X., Lyu, X., Freeman, E., Wang, J., et al. (2019) · 2019
Cited alongside, same era.
Universality and diversity in human song
Mehr, S. A., Singh, M., Knox, D., Ketter, D. M., Pickens-Jones, D., Atwood, S., Lucas, C., Jacoby, N., Egner, A. A., Hopkins, E. J., et al. (2019) · 2019
Cited alongside, same era.
Effectiveness of self-supervised pre-training for asr
Baevski, A. and Mohamed, A. (2020) · 2020
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M. (2020) · 2020
Music representation learning based on editorial metadata from discogs
Alonso-Jiménez, P., Serra, X., and Bogdanov, D. (2022) · 2022
Later among the works it cites.
Data2vec: A general framework for self-supervised learning in speech, vision and language
Baevski, A., Hsu, W.-N., Xu, Q., Babu, A., Gu, J., and Auli, M. (2022) · 2022
Later among the works it cites.
Audiolm: a language modeling approach to audio generation
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Teboul, O., Grangier, D., Tagliasacchi, M., and Zeghidour, N. (2022) · 2022
Later among the works it cites.
Familiarity of background music modulates the cortical tracking of target speech at the “cocktail party”
Brown, J. A. and Bidelman, G. M. (2022) · 2022
Later among the works it cites.
High fidelity neural audio compression
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y. (2022) · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Tailored perception: Individuals’ speech and music perception strategies fit their perceptual abilities
Jasmin, K., Dick, F., Holt, L. L., and Tierney, A. (2020) · 2020
Cited alongside, same era.
Music4all: A new music database and its applications
Santana, I. A. P., Pinhelli, F., Donini, J., Catharin, L., Mangolin, R. B., Feltrim, V. D., Domingues, M. A., et al. (2020) · 2020
Cited alongside, same era.
Codified audio language modeling learns useful representations for music information retrieval
Castellon, R., Donahue, C., and Liang, P. (2021) · 2021
Cited alongside, same era.
Beatnet: Crnn and particle filtering for online joint beat downbeat and meter tracking
Heydari, M., Cwitkowitz, F., and Duan, Z. (2021) · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A. (2021) · 2021
Cited alongside, same era.
Music demixing challenge 2021
Mitsufuji, Y., Fabbro, G., Uhlich, S., Stöter, F.-R., Défossez, A., Kim, M., Choi, W., Yu, C.-Y., and Cheuk, K.-W. (2022) · 2021
Cited alongside, same era.
Later among the works it cites.
Eva: Exploring the limits of masked visual representation learning at scale
Fang, Y., Wang, W., Xie, B., Sun, Q., Wu, L., Wang, X., Huang, T., Wang, X., and Cao, Y. (2022) · 2022
Later among the works it cites.
Mulan: A joint embedding of music audio and natural language
Huang, Q., Jansen, A., Lee, J., Ganti, R., Li, J. Y., and Ellis, D. P. (2022) · 2022
Later among the works it cites.
Map-music2vec: A simple and effective baseline for self-supervised music audio representation learning
Li, Y., Yuan, R., Zhang, G., MA, Y., Lin, C., Chen, X., Ragni, A., Yin, H., Hu, Z., He, H., et al. (2022) · 2022
Later among the works it cites.
Supervised and unsupervised learning of audio representations for music understanding
McCallum, M. C., Korzeniowski, F., Oramas, S., Gouyon, F., and Ehmann, A. F. (2022) · 2022
Later among the works it cites.
Transfer learning with deep neural embeddings for music classification tasks
Modrzejewski, M., Szachewicz, P., and Rokita, P. (2023) · 2022
Later among the works it cites.
The cocktail fork problem: Three-stem audio separation for real-world soundtracks
Petermann, D., Wichern, G., Wang, Z.-Q., and Le Roux, J. (2022) · 2022
Later among the works it cites.
Towards learning universal audio representations
Wang, L., Luc, P., Wu, Y., Recasens, A., Smaira, L., Brock, A., Jaegle, A., Alayrac, J.-B., Dieleman, S., Carreira, J., et al. (2022b) · 2022
Later among the works it cites.
Deformable cnn and imbalance-aware feature learning for singing technique classification
Yamamoto, Y., Nam, J., and Terasawa, H. (2022) · 2022
Later among the works it cites.
Transfer learning with jukebox for music source separation
Zai El Amri, W., Tautz, O., Ritter, H., and Melnik, A. (2022) · 2022
Later among the works it cites.
On the effectiveness of speech self-supervised learning for music
Ma, Y., Yuan, R., Li, Y., Zhang, G., Chen, X., Yin, H., Lin, C., Benetos, E., Ragni, A., Gyenge, N., et al. (2023) · 2023
Closest in time.
Hybrid transformers for music source separation
Rouard, S., Massa, F., and Défossez, A. (2023) · 2023
Closest in time.
Neural codec language models are zero-shot text to speech synthesizers
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., et al. (2023) · 2023
Closest in time.
Deep learning and music adversaries
Kereliuk, C., Sturm, B. L., and Larsen, J. (2015) · 2071
Closest in time.