Fetching the paper…
Reading the bibliography…
In this paper, we introduce Jointist, an instrument-aware multi-instrument framework that is capable of transcribing, recognizing, and separating multiple musical instruments from an audio clip.
M. Piszczalski and B. A. Galler, “Automatic music transcription,” Computer Music Journal , pp. 24–31, 1977
1977
Earlier work this paper cites.
A. Klapuri and A. Eronen, “Automatic transcription of music,” in Proceedings of the Stockholm Music Acoustics Conference . Citeseer, 1998, pp. 6–9
1998
Earlier work this paper cites.
M. Goto, H. Hashiguchi, T. Nishimura, and R. Oka, “RWC Music Database: Popular, classical and jazz music databases.” in Prc. ISMIR , vol. 2, 2002, pp. 287–288
2002
Earlier work this paper cites.
J. Paulus and A. Klapuri, “Model-based event labeling in the transcription of percussive audio signals,” in Proc. Int. Conf. Digital Audio Effects (DAFX) . Citeseer, 2003, pp. 73–77
2003
Earlier work this paper cites.
S. Pauws, “Musical key extraction from audio.” in 5th International Society for Music Information Retrieval Conference , 2004
2004
Earlier work this paper cites.
S. W. Hainsworth and M. D. Macleod, “Particle filtering applied to musical tempo tracking,” EURASIP Journal on Advances in Signal Processing , vol. 2004, no. 15, pp. 1–11, 2004
2004
Earlier work this paper cites.
G. Ozcan, C. Isikhan, and A. Alpkocak, “Melody extraction on midi music files,” in Seventh IEEE International Symposium on Multimedia (ISM’05) . Ieee, 2005, pp. 8–pp
2005
Earlier work this paper cites.
F. Gouyon, A computational approach to rhythm description-Audio features for the computation of rhythm periodicity functions and their use in tempo induction and music content processing . Universitat Pompeu Fabra, 2006
2006
Earlier work this paper cites.
M. Mauch, C. Cannam, M. Davies, S. Dixon, C. Harte, S. Kolozali, D. Tidhar, and M. Sandler, “OMRAS2 metadata project 2009,” in ISMIR Late Breaking and Demo , 2009
2009
Earlier work this paper cites.
E. Law, K. West, M. I. Mandel, M. Bay, and J. S. Downie, “Evaluation of algorithms using games: The case of music tagging.” in Conference of the International Society for Music Information Retrieval Conference (ISMIR) , 2009
2009
Earlier work this paper cites.
J. A. Burgoyne, J. Wild, and I. Fujinaga, “An expert ground truth set for audio chord recognition and music analysis.” in ISMIR , vol. 11, 2011, pp. 633–638
2011
Earlier work this paper cites.
A. Holzapfel, M. E. Davies, J. R. Zapata, J. L. Oliveira, and F. Gouyon, “Selective sampling for beat tracking evaluation,” IEEE Transactions on Audio, Speech, and Language Processing , vol. 20, no. 9, pp. 2539–2548, 2012
2012
Earlier work this paper cites.
E. Benetos, S. Dixon, D. Giannoulis, H. Kirchhoff, and A. Klapuri, “Automatic music transcription: challenges and future directions,” Journal of Intelligent Information Systems , vol. 41, no. 3, pp. 407–434, 2013
2013
Earlier work this paper cites.
F. Krebs, S. Böck, and G. Widmer, “Rhythmic pattern modeling for beat and downbeat tracking in musical audio.” in ISMIR . Citeseer, 2013, pp. 227–232
2013
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” The journal of machine learning research , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” Proceedings of the 3rd International Conference on Learning Representations (ICLR) , 2014
2014
Earlier work this paper cites.
S. Sigtia, E. Benetos, and S. Dixon, “An end-to-end neural network for polyphonic piano music transcription,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, pp. 927–939, 2015
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International conference on machine learning . PMLR, 2015, pp. 448–456
2015
Earlier work this paper cites.
E. J. Humphrey and J. P. Bello, “Four timely insights on automatic chord estimation.” in Proc. ISMIR , vol. 10, 2015, pp. 673–679
2015
Earlier work this paper cites.
U. Marchand and G. Peeters, “Swing ratio estimation,” in Digital Audio Effects 2015 (Dafx15) , 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
C. Raffel, Learning-based methods for comparing sequences, with applications to audio-to-midi alignment and matching . Columbia University, 2016
2016
Earlier work this paper cites.
S. Durand, J. P. Bello, B. David, and G. Richard, “Feature adapted convolutional neural networks for downbeat tracking,” in 2016 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2016, pp. 296–300
2016
Earlier work this paper cites.
R. Vogl, M. Dorfer, G. Widmer, and P. Knees, “Drum transcription via joint beat and drum modeling using convolutional recurrent neural networks.” in 18th International Society for Music Information Retrieval Conference, ISMIR 2017 , 2017, pp. 150–157
2017
Earlier work this paper cites.
K. Choi, G. Fazekas, M. Sandler, and K. Cho, “Transfer learning for music classification and regression tasks,” in 18th International Society for Music Information Retrieval Conference, ISMIR 2017 . International Society for Music Information Retrieval, 2017, pp. 141–149
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
J. Thickstun, Z. Harchaoui, D. P. Foster, and S. M. Kakade, “Invariances and data augmentation for supervised music transcription,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 2241–2245, 2017
2017
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Cited alongside, same era.
A. Jansson, E. Humphrey, N. Montecchio, R. Bittner, A. Kumar, and T. Weyde, “Singing voice separation with deep u-net convolutional networks,” 18th International Society for Music Information Retrieval Conference, ISMIR 2017 , 2017
2017
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of Machine Learning Research , vol. 21, pp. 1–67, 2020
2020
Later among the works it cites.
E. Manilow, P. Seetharaman, and B. Pardo, “Simultaneous separation and transcription of mixtures with multiple polyphonic and percussive instruments,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 771–775
2020
Later among the works it cites.
K. Tanaka, T. Nakatsuka, R. Nishikimi, K. Yoshii, and S. Morishima, “Multi-instrument music transcription based on deep spherical clustering of spectrograms and pitchgrams,” in International Society for Music Information Retrieval Conference (ISMIR), Montreal, Canada , 2020
2020
Later among the works it cites.
M. Won, S. Chun, O. Nieto, and X. Serrc, “Data-driven harmonic filters for audio representation learning,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 536–540
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2017
Cited alongside, same era.
E. Benetos, S. Dixon, Z. Duan, and S. Ewert, “Automatic music transcription: An overview,” IEEE Signal Processing Magazine , vol. 36, no. 1, pp. 20–30, 2018
2018
Cited alongside, same era.
Y. Wu and W. Li, “Automatic audio chord recognition with midi-trained deep feature and blstm-crf sequence decoding model,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 2, pp. 355–366, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, I. Simon, C. Hawthorne, N. Shazeer, A. M. Dai, M. D. Hoffman, M. Dinculescu, and D. Eck, “Music transformer: Generating music with long-term structure,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Cited alongside, same era.
J. W. Kim and J. P. Bello, “Adversarial learning for improved onsets and frames music transcription,” International Society forMusic Information Retrieval Conference , pp. 670–677, 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
S. Böck and M. E. Davies, “Deconstruct, analyse, reconstruct: How to improve tempo, beat, and downbeat estimation.” in 21th International Society for Music Information Retrieval Conference, ISMIR 2020 , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Q. Kong, B. Li, X. Song, Y. Wan, and Y. Wang, “High-resolution piano transcription with pedals by regressing onset and offset times,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 3707–3717, 2021
2021
Later among the works it cites.
K. W. Cheuk, D. Herremans, and L. Su, “Reconvat: A semi-supervised automatic music transcription framework for low-resource real-world data,” in Proceedings of the 29th ACM International Conference on Multimedia , 2021, pp. 3918–3926
2021
Later among the works it cites.
Y.-N. Hung, G. Wichern, and J. Le Roux, “Transcription is all you need: Learning to separate musical mixtures with score as supervision,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 46–50
2021
Later among the works it cites.
2021
Later among the works it cites.
L. Lin, Q. Kong, J. Jiang, and G. Xia, “A unified model for zero-shot music source separation, transcription and synthesis,” in International Society for Music Information Retrieval (ISMIR) , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
M. Won, K. Choi, and X. Serra, “Semi-supervised music tagging transformer,” Conference of the International Society for Music Information Retrieval Conference (ISMIR) , 2021
2021
Later among the works it cites.
M. Won, J. Spijkervet, and K. Choi, Music Classification: Beyond Supervised Learning, Towards Real-world Applications . https://music-classification.github.io/tutorial, 2021. [Online]. Available: https://music-classification.github.io/tutorial
2021
Later among the works it cites.
K. W. Cheuk, Y.-J. Luo, E. Benetos, and D. Herremans, “Revisiting the onsets and frames model with additive attention,” in Proceedings of the International Joint Conference on Neural Networks . IEEE, 2021, p. In press
2021
Later among the works it cites.
M. F. Matthew E. P. Davies, Sebastian B ock, Tempo, Beat and Downbeat Estimation . https://tempobeatdownbeat.github.io/tutorial/intro.html, 2021. [Online]. Available: https://tempobeatdownbeat.github.io/tutorial/intro.html
2021
Later among the works it cites.
2021
Later among the works it cites.
R. Guo, I. Simpson, C. Kiefer, T. Magnusson, and D. Herremans, “Musiac: An extensible generative framework for music infilling applications with multi-level control,” in International Conference on Computational Intelligence in Music, Sound, Art and Design (Part of EvoStar) . Springer, 2022, pp. 341–356
2022
Closest in time.
Y.-Y. Yang, M. Hira, Z. Ni, A. Astafurov, C. Chen, C. Puhrsch, D. Pollack, D. Genzel, D. Greenberg, E. Z. Yang et al. , “Torchaudio: Building blocks for audio and speech processing,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 6982–6986
2022
Closest in time.
Y.-N. Hung, J.-C. Wang, X. Song, W.-T. Lu, and M. Won, “Modeling beats and downbeats with a time-frequency transformer,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 401–405
2022
Closest in time.
J.-C. Wang, Y.-N. Hung, and J. B. Smith, “To catch a chorus, verse, intro, or anything else: Analyzing a song with structural functions,” in Proc. ICASSP , 2022, pp. 416–420
2022
Closest in time.