Fetching the paper…
Reading the bibliography…
Learning musical structures and composition patterns is necessary for both music generation and understanding, but current methods do not make uniform use of learned features to generate and comprehend music simultaneously.
T. Eiter and H. Mannila, “Computing discrete fréchet distance,” 1994
1994
Earlier work this paper cites.
W. Chai and B. Vercoe, “Melody retrieval on the web,” in Multimedia Computing and Networking 2002 , vol. 4673. SPIE, 2001, pp. 226–241
2001
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
Y. Xiong, F. Su, and Q. Wang, “Automatic music mood classification by learning cross-media relevance between audio and lyrics,” in 2017 IEEE international conference on multimedia and expo (ICME) . IEEE, 2017, pp. 961–966
2017
Earlier work this paper cites.
C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, I. Simon, C. Hawthorne, N. Shazeer, A. M. Dai, M. D. Hoffman, M. Dinculescu, and D. Eck, “Music transformer: Generating music with long-term structure,” in International Conference on Learning Representations , 2018
2018
Earlier work this paper cites.
J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of naacL-HLT , vol. 1, 2019, p. 2
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. Hawthorne, A. Stasyuk, A. Roberts, I. Simon, C.-Z. A. Huang, S. Dieleman, E. Elsen, J. Engel, and D. Eck, “Enabling factorized piano music modeling and generation with the MAESTRO dataset,” in International Conference on Learning Representations , 2019
2019
Earlier work this paper cites.
D. Jeong, T. Kwon, Y. Kim, K. Lee, and J. Nam, “Virtuosonet: A hierarchical rnn-based system for modeling expressive piano performance.” in ISMIR , 2019, pp. 908–915
2019
Earlier work this paper cites.
D. Jeong, T. Kwon, Y. Kim, and J. Nam, “Graph neural network for music score data and modeling expressive piano performance,” in International conference on machine learning . PMLR, 2019, pp. 3060–3070
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
Y.-Q. Lim, C. S. Chan, and F. Y. Loo, “Style-conditioned music generation,” in 2020 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 2020, pp. 1–6
2020
Cited alongside, same era.
Y.-S. Huang and Y.-H. Yang, “Pop music transformer: Beat-based modeling and generation of expressive pop piano compositions,” in Proceedings of the 28th ACM international conference on multimedia , 2020, pp. 1180–1188
2020
Cited alongside, same era.
Y. Ren, J. He, X. Tan, T. Qin, Z. Zhao, and T.-Y. Liu, “Popmag: Pop music accompaniment generation,” in Proceedings of the 28th ACM international conference on multimedia , 2020, pp. 1198–1206
2020
Cited alongside, same era.
M. Zeng, X. Tan, R. Wang, Z. Ju, T. Qin, and T.-Y. Liu, “Musicbert: Symbolic music understanding with large-scale pre-training,” in Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 , 2021, pp. 791–800
2021
Later among the works it cites.
2021
Later among the works it cites.
W.-Y. Hsiao, J.-Y. Liu, Y.-C. Yeh, and Y.-H. Yang, “Compound word transformer: Learning to compose full-song music over dynamic directed hypergraphs,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 1, 2021, pp. 178–186
2021
Later among the works it cites.
H.-T. Hung, J. Ching, S. Doh, N. Kim, J. Nam, and Y.-H. Yang, “Emopia: A multi-modal pop piano dataset for emotion recognition and emotion-based music generation,” in International Society for Music Information Retrieval Conference, ISMIR 2021 . International Society for Music Information Retrieval, 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , 2020, pp. 7871–7880
2020
Cited alongside, same era.
2020
Cited alongside, same era.
F. Foscarin, A. Mcleod, P. Rigaux, F. Jacquemard, and M. Sakai, “Asap: a dataset of aligned scores and performances for piano transcription,” in International Society for Music Information Retrieval Conference , no. CONF, 2020, pp. 534–541
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Later among the works it cites.
Q. Kong, B. Li, J. Chen, and Y. Wang, “Giantmidi-piano: A large-scale midi dataset for classical piano music,” 2022
2022
Later among the works it cites.
J. Zhao, G. Ru, Y. Yu, Y. Wu, D. Li, and W. Li, “Multimodal music emotion recognition with hierarchical cross-modal attention network,” in 2022 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 2022, pp. 1–6
2022
Later among the works it cites.
Z. Shen, L. Yang, Z. Yang, and H. Lin, “More than simply masking: Exploring pre-training strategies for symbolic music understanding,” in Proceedings of the 2023 ACM International Conference on Multimedia Retrieval , 2023, pp. 540–544
2023
Later among the works it cites.
Y. Li, H. Fan, R. Hu, C. Feichtenhofer, and K. He, “Scaling language-image pre-training via masking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 23 390–23 400
2023
Later among the works it cites.
2024
Closest in time.