Fetching the paper…
Reading the bibliography…
Data is the lifeblood of modern machine learning systems, including for those in Music Information Retrieval (MIR).
1906
Earlier work this paper cites.
V. Emiya, N. Bertin, B. David, and R. Badeau, “Maps-a piano database for multipitch estimation and automatic transcription of music,” 2010
2010
Earlier work this paper cites.
L. Naveda, F. Gouyon, C. Guedes, and M. Leman, “Microtiming patterns and interactions with musical properties in samba music,” Journal of New Music Research , vol. 40, no. 3, pp. 225–238, 2011
2011
Earlier work this paper cites.
N. Boulanger-Lewandowski, Y. Bengio, and P. Vincent, “Modeling temporal dependencies in high-dimensional sequences: Application to polyphonic music generation and transcription,” in International Conference on Machine Learning , 2012
2012
Earlier work this paper cites.
C. Raffel, B. McFee, E. J. Humphrey, J. Salamon, O. Nieto, D. Liang, D. P. Ellis, and C. C. Raffel, “mir_eval: A transparent implementation of common MIR metrics,” in In Proceedings of the 15th International Society for Music Information Retrieval Conference, ISMIR , 2014
2014
Earlier work this paper cites.
J. Schlüter and T. Grill, “Exploring data augmentation for improved singing voice detection with neural networks.” in ISMIR , 2015, pp. 121–126
2015
Earlier work this paper cites.
C. Raffel, Learning-based methods for comparing sequences, with applications to audio-to-midi alignment and matching . Columbia University, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
S. Uhlich, M. Porcu, F. Giron, M. Enenkl, T. Kemp, N. Takahashi, and Y. Mitsufuji, “Improving music source separation based on deep neural networks through data augmentation and network blending,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2017, pp. 261–265
2017
Earlier work this paper cites.
C.-Z. A. Huang, T. Cooijmans, A. Roberts, A. Courville, and D. Eck, “Counterpoint by convolution,” in Proceedings of 18st International Conference on Music Information Retrieval, ISMIR , 2017
2017
Earlier work this paper cites.
J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan, “Neural audio synthesis of musical notes with wavenet autoencoders,” in International Conference on Machine Learning . PMLR, 2017, pp. 1068–1077
2017
Earlier work this paper cites.
M. Miron, J. Janer Mestres, and E. Gómez Gutiérrez, “Generating data to train convolutional neural networks for classical music source separation,” in Lokki T, Pätynen J, Välimäki V, editors. Proceedings of the 14th Sound and Music Computing Conference; 2017 Jul 5-8; Espoo, Finland. Aalto: Aalto University; 2017. p. 227-33. Aalto University, 2017
2017
Earlier work this paper cites.
Z. Rafii, A. Liutkus, F.-R. Stöter, S. I. Mimilakis, and R. Bittner, “The MUSDB18 corpus for music separation,” Dec. 2017. [Online]. Available: https://doi.org/10.5281/zenodo.1117372
2017
Earlier work this paper cites.
J. Thickstun, Z. Harchaoui, and S. Kakade, “Learning features of music from scratch,” 2017
2017
Earlier work this paper cites.
R. I.-R. BS.1770-4, “Algorithms to measure audio programme loudness and true-peak audio level,” 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. Cartwright and J. P. Bello, “Increasing drum transcription vocabulary using data synthesis,” in Proc. International Conference on Digital Audio Effects (DAFx) , 2018, pp. 72–79
2018
Earlier work this paper cites.
Q. Xi, R. M. Bittner, J. Pauwels, X. Ye, and J. P. Bello, “Guitarset: A dataset for guitar transcription.” in ISMIR , 2018, pp. 453–460
2018
Earlier work this paper cites.
B. Li, X. Liu, K. Dinesh, Z. Duan, and G. Sharma, “Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,” IEEE Transactions on Multimedia , vol. 21, no. 2, pp. 522–535, 2018
2018
Earlier work this paper cites.
N. Kalchbrenner, E. Elsen, K. Simonyan, S. Noury, N. Casagrande, E. Lockhart, F. Stimberg, A. Oord, S. Dieleman, and K. Kavukcuoglu, “Efficient neural audio synthesis,” in International Conference on Machine Learning . PMLR, 2018, pp. 2410–2419
2018
Earlier work this paper cites.
M. Sogorski, T. Geisel, and V. Priesemann, “Correlated microtiming deviations in jazz and rock music,” PloS one , vol. 13, no. 1, p. e0186361, 2018
2018
Earlier work this paper cites.
K. A. Pati, S. Gururani, and A. Lerch, “Assessment of student music performances using deep neural networks,” Applied Sciences , vol. 8, no. 4, p. 507, 2018
2018
Earlier work this paper cites.
D. Bogdanov, M. Won, P. Tovstogan, A. Porter, and X. Serra, “The mtg-jamendo dataset for automatic music tagging,” 2019
2019
Earlier work this paper cites.
C. Hawthorne, A. Stasyuk, A. Roberts, I. Simon, C.-Z. A. Huang, S. Dieleman, E. Elsen, J. Engel, and D. Eck, “Enabling factorized piano music modeling and generation with the MAESTRO dataset,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=r1lYRjC9F7
2019
Earlier work this paper cites.
C. Payne, “Musenet,” OpenAI Blog , 2019
2019
Earlier work this paper cites.
E. Manilow, G. Wichern, P. Seetharaman, and J. Le Roux, “Cutting music source separation some Slakh: A dataset to study the impact of training data quality and quantity,” in Proc. IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) . IEEE, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
L. Prétet, R. Hennequin, J. Royo-Letelier, and A. Vaglio, “Singing voice separation: A study on training data,” in ICASSP 2019-2019 ieee international conference on acoustics, speech and signal processing (icassp) . IEEE, 2019, pp. 506–510
2019
Cited alongside, same era.
T. Karras, S. Laine, and T. Aila, “A style-based generator architecture for generative adversarial networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 4401–4410
2019
Cited alongside, same era.
2019
Cited alongside, same era.
B. Wang and Y.-H. Yang, “Performancenet: Score-to-audio music generation with multi-band convolutional residual network,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 33, no. 01, 2019, pp. 1174–1181
2019
Cited alongside, same era.
R. Castellon, C. Donahue, and P. Liang, “Codified audio language modeling learns useful representations for music information retrieval,” in Proceedings of 22st International Conference on Music Information Retrieval, ISMIR , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “Sdr–half-baked or well done?” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 626–630
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Later among the works it cites.
Y. Zhang, H. Ling, J. Gao, K. Yin, J.-F. Lafleche, A. Barriuso, A. Torralba, and S. Fidler, “Datasetgan: Efficient labeled data factory with minimal human effort,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 10 145–10 155
2021
Later among the works it cites.
S. I. Nikolenko et al. , Synthetic data for deep learning . Springer, 2021
2021
Later among the works it cites.
E. Wood, T. Baltrušaitis, C. Hewitt, S. Dziadzio, T. J. Cashman, and J. Shotton, “Fake it till you make it: Face analysis in the wild using synthetic data alone,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 3681–3691
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Ens and P. Pasquier, “Building the metamidi dataset: Linking symbolic and audio musical data,” in Proceedings of 22st International Conference on Music Information Retrieval, ISMIR , 2021
2021
Later among the works it cites.
D. Foster, S. Dixon et al. , “Filosax: A dataset of annotated jazz saxophone recordings,” in Proceedings of 22st International Conference on Music Information Retrieval, ISMIR , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
E. Manilow, P. O’Reilly, P. Seetharaman, and B. Pardo, “Source separation by steering pretrained music models,” in ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 126–130
2022
Closest in time.
L. Wang, P. Luc, Y. Wu, A. Recasens, L. Smaira, A. Brock, A. Jaegle, J.-B. Alayrac, S. Dieleman, J. Carreira et al. , “Towards learning universal audio representations,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 4593–4597
2022
Closest in time.
A. Jahanian, X. Puig, Y. Tian, and P. Isola, “Generative models as a data source for multiview representation learning,” in International Conference on Learning Representations , 2022
2022
Closest in time.
Y. Wu, E. Manilow, Y. Deng, R. Swavely, K. Kastner, T. Cooijmans, A. Courville, C.-Z. A. Huang, and J. Engel, “MIDI-DDSP: Detailed control of musical performance via hierarchical modeling,” in International Conference on Learning Representations , 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
H.-W. Dong, C. Zhou, T. Berg-Kirkpatrick, and J. McAuley, “Deep performer: Score-to-audio music performance synthesis,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 951–955
2022
Closest in time.
“JSB-Chorales-dataset,” https://github.com/czhuang/JSB-Chorales-dataset , 2022, [Online; accessed 01-May-2022]
2022
Closest in time.
“Coconet-pytorch,” https://github.com/lukewys/coconet-pytorch , 2022, [Online; accessed 01-May-2022]
2022
Closest in time.
J. P. Gardner, I. Simon, E. Manilow, C. Hawthorne, and J. Engel, “MT3: Multi-task multitrack music transcription,” in International Conference on Learning Representations , 2022
2022
Closest in time.
2022
Closest in time.
M. Kawamura, T. Nakamura, D. Kitamura, H. Saruwatari, Y. Takahashi, and K. Kondo, “Differentiable digital signal processing mixture model for synthesis parameter extraction from mixture of harmonic sounds,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 941–945
2022
Closest in time.
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “Zero-shot audio source separation through query-based learning from weakly-labeled data,” in Proceedings of the AAAI Conference on Artificial Intelligence , 2022
2022
Closest in time.
J. Tremblay, T. To, and S. Birchfield, “Falling things: A synthetic dataset for 3d object detection and pose estimation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops , 2018, pp. 2038–2041
2041
Closest in time.