Fetching the paper…
Reading the bibliography…
Recent years have seen many audio-domain text-to-music generation models that rely on large amounts of text-audio pairs for training.
2010
Earlier work this paper cites.
T. Bertin-Mahieux, D. P. Ellis, B. Whitman, and P. Lamere, “The million song dataset,” in Proc. ISMIR , 2011
2011
Earlier work this paper cites.
C. Raffel, “Learning-based methods for comparing sequences, with applications to audio-to-MIDI alignment and matching,” Ph.D. dissertation, Columbia University, 2016
2016
Earlier work this paper cites.
O. Mogren, “C-RNN-GAN: Continuous recurrent neural networks with adversarial training,” in Proc. NeurIPS Workshop on Constructive Machine Learning , 2016
2016
Earlier work this paper cites.
C. Hawthorne, A. Stasyuk, A. Roberts, I. Simon, C.-Z. A. Huang, S. Dieleman, E. Elsen, J. Engel, and D. Eck, “Enabling factorized piano music modeling and generation with the MAESTRO dataset,” in Proc. ICLR , 2019
2019
Earlier work this paper cites.
N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using siamese bert-networks,” in Proc. EMNLP , 2019
2019
Earlier work this paper cites.
S.-L. Wu and Y.-H. Yang, “The Jazz Transformer on the front line: Exploring the shortcomings of AI-composed music through quantitative measures,” in Proc. ISMIR , 2020
2020
Earlier work this paper cites.
J. Ens and P. Pasquier, “Building the MetaMIDI dataset: Linking symbolic and audio musical data,” in Proc. ISMIR , 2021
2021
Earlier work this paper cites.
H.-T. Hung, J. Ching, S. Doh, N. Kim, J. Nam, and Y.-H. Yang, “EMOPIA: A multi-modal pop piano dataset for emotion recognition and emotion-based music generation,” in Proc. ISMIR , 2021
2021
Earlier work this paper cites.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Cited alongside, same era.
S. Wu, D. Yu, X. Tan, and M. Sun, “Clamp: Contrastive language-music pre-training for cross-modal symbolic music information retrieval,” in Proc. ISMIR , 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. von Rütte, L. Biggio, Y. Kilcher, and T. Hofmann, “FIGARO: Controllable music generation using learned and expert features,” in Proc. ICLR , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
H.-W. Dong, K. Chen, S. Dubnov, J. McAuley, and T. Berg-Kirkpatrick, “Multitrack music transformer,” in Proc. ICASSP , 2023
2023
Cited alongside, same era.
S. Doh, K. Choi, J. Lee, and J. Nam, “LP-MusicCaps: LLM-based pseudo music captioning,” in Proc. ISMIR , 2023
2023
Cited alongside, same era.
Y. Wu, K. Chen, T. Zhang, Y. Hui, M. Nezhurina, T. Berg-Kirkpatrick, and S. Dubnov, “Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,” in Proc. ICASSP , 2023
2023
Cited alongside, same era.
2025
Closest in time.
BigScience, T. L. Scao, A. Fan, C. Wolf, T. Rush, S. Biderman, G. B. Black, S. Curtis, D. Elbayad, T. Gao et al. , “BLOOM: A 176b-parameter open-access multilingual language model,” https://huggingface.co/bigscience/bloom , 2022, accessed: 2025-03-25
2025
Closest in time.
——, “all-minilm-l6-v2: A lightweight sentence embedding model,” https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2 , 2021, accessed: 2025-03-25
2025
Closest in time.
“Melobytes ABC to MIDI Converter,” https://melobytes.com/en/app/abc2midi , accessed: 2025-03-25
2025
Closest in time.