Fetching the paper…
Reading the bibliography…
Generating music from text descriptions is a user-friendly mode since the text is a relatively easy interface for user engagement.
K. Choi, C. Hawthorne, I. Simon, M. Dinculescu, and J. Engel, “Encoding musical style with transformer autoencoders,” in International Conference on Machine Learning . PMLR, 2020, pp. 1899–1908
1908
Earlier work this paper cites.
P. R. Cook, Music, Cognition, and Computerized Sound: An Introduction to Psychoacoustics . MIT press, 2001
2001
Earlier work this paper cites.
2014
Earlier work this paper cites.
H. H. Mao, T. Shin, and G. Cottrell, “Deepj: Style-specific music generation,” in 2018 IEEE 12th International Conference on Semantic Computing (ICSC) . IEEE, 2018, pp. 377–382
2018
Earlier work this paper cites.
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT (1) . Association for Computational Linguistics, 2019, pp. 4171–4186
2019
Earlier work this paper cites.
Y. Zhang, Z. Wang, D. Wang, and G. Xia, “BUTTER: A representation learning framework for bi-directional music-sentence retrieval and generation,” in Proceedings of the 1st Workshop on NLP for Music and Audio (NLP4MusA) . Association for Computational Linguistics, 16 Oct. 2020, pp. 54–58
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
Z. Wang, K. Chen, J. Jiang, Y. Zhang, M. Xu, S. Dai, and G. Xia, “POP909: A pop-song dataset for music arrangement generation,” in Proceedings of International Society for Music Information Retrieval Conference (ISMIR) , 2020, pp. 38–45
2020
Earlier work this paper cites.
Y.-S. Huang and Y.-H. Yang, “Pop music transformer: Beat-based modeling and generation of expressive pop piano compositions,” in Proceedings of ACM International Conference on Multimedia (MM) , 2020, pp. 1180–1188
2020
Earlier work this paper cites.
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret, “Transformers are rnns: Fast autoregressive transformers with linear attention,” in International Conference on Machine Learning . PMLR, 2020, pp. 5156–5165
2020
Earlier work this paper cites.
2021
Cited alongside, same era.
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” in International Conference on Machine Learning . PMLR, 2021, pp. 8821–8831
2021
Cited alongside, same era.
2021
Cited alongside, same era.
M. Zeng, X. Tan, R. Wang, Z. Ju, T. Qin, and T.-Y. Liu, “MusicBERT: Symbolic music understanding with large-scale pre-training,” in Findings of the Association for Computational Linguistics (ACL Findings) , 2021, pp. 791–800
2021
Cited alongside, same era.
X. Zhang, J. Zhang, Y. Qiu, L. Wang, and J. Zhou, “Structure-enhanced pop music generation via harmony-aware learning,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 1204–1213
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans et al. , “Photorealistic text-to-image diffusion models with deep language understanding,” Advances in Neural Information Processing Systems , vol. 35, pp. 36 479–36 494, 2022
2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Ens and P. Pasquier, “Building the metamidi dataset: Linking symbolic and audio musical data.” in ISMIR , 2021, pp. 182–188
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Cited alongside, same era.
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High-resolution image synthesis with latent diffusion models,” in CVPR . IEEE, 2022, pp. 10 674–10 685
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Y. Zhu, K. Olszewski, Y. Wu, P. Achlioptas, M. Chai, Y. Yan, and S. Tulyakov, “Quantized gan for complex music generation from dance videos,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXVII . Springer, 2022, pp. 182–199
2022
Cited alongside, same era.
L. N. Ferreira, L. Mou, J. Whitehead, and L. H. Lelis, “Controlling perceived emotion in symbolic music generation with monte carlo tree search,” in Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment , vol. 18, no. 1, 2022, pp. 163–170
2022
Cited alongside, same era.
Later among the works it cites.
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
OpenAI, “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023
2023
Closest in time.
2023
Closest in time.
D. von Rütte, L. Biggio, Y. Kilcher, and T. Hofmann, “Figaro: Controllable music generation using learned and expert features,” in The Eleventh International Conference on Learning Representations , 2023
2023
Closest in time.
P. Liu, W. Yuan, J. Fu, Z. Jiang, H. Hayashi, and G. Neubig, “Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,” ACM Computing Surveys , vol. 55, no. 9, pp. 1–35, 2023
2023
Closest in time.
S. Di, Z. Jiang, S. Liu, Z. Wang, L. Zhu, Z. He, H. Liu, and S. Yan, “Video background music generation with controllable music transformer,” in Proceedings of the 29th ACM International Conference on Multimedia , ser. MM ’21. New York, NY, USA: Association for Computing Machinery, 2021, pp. 2037–2045
2045
Closest in time.