Fetching the paper…
Reading the bibliography…
Text-to-song generation, the task of creating vocals and accompaniment from textual inputs, poses significant challenges due to domain complexity and data scarcity.
The million song dataset
Bertin-Mahieux, T., Ellis, D. P., Whitman, B., and Lamere, P · 2011
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Audio set: An ontology and human-labeled dataset for audio events
Gemmeke, J. F., Ellis, D. P., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M · 2017
Earlier work this paper cites.
Cnn architectures for large-scale audio classification, 2017
Hershey, S., Chaudhuri, S., Ellis, D. P. W., Gemmeke, J. F., Jansen, A., Moore, R. C., Plakal, M., Platt, D., Saurous, R. A., Seybold, B., Slaney, M., Weiss, R. J., and Wilson, K · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Learning to recognize musical genre from audio
Defferrard, M., Mohanty, S. P., Carroll, S. F., and Salathé, M · 2018
Earlier work this paper cites.
The mtg-jamendo dataset for automatic music tagging
Bogdanov, D., Won, M., Tovstogan, P., Porter, A., and Serra, X · 2019
Earlier work this paper cites.
Music transformer: Generating music with long-term structure
Huang, C.-Z. A., Vaswani, A., Uszkoreit, J., Simon, I., Hawthorne, C., Shazeer, N., Dai, A. M., Hoffman, M. D., Dinculescu, M., and Eck, D · 2019
Earlier work this paper cites.
Fréchet audio distance: A metric for evaluating music enhancement algorithms, 2019
Kilgour, K., Zuluaga, M., Roblek, D., and Sharifi, M · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Earlier work this paper cites.
Jukebox: A generative model for music
Dhariwal, P., Jun, H., Payne, C., Kim, J. W., Radford, A., and Sutskever, I · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Ji, S., Luo, J., and Yang, X · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Variational diffusion models
Kingma, D., Salimans, T., Poole, B., and Ho, J · 2021
Earlier work this paper cites.
Soundstream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2021
Earlier work this paper cites.
Scaling instruction-finetuned language models, 2022
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Castro-Ros, A., Pellat, M., Robinson, K., Valter, D., Narang, S., Mishra, G., Yu, A., Zhao, V., Huang, Y., Dai, A., Yu, H., Petrov, S., Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J · 2022
Earlier work this paper cites.
High fidelity neural audio compression
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2022
Cited alongside, same era.
Riffusion-stable diffusion for real-time music generation
Forsgren, S. and Martiros, H · 2022
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision, 2022
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Seed-music: A unified framework for high quality and controlled music generation, 2024
Bai, Y., Chen, H., Chen, J., Chen, Z., Deng, Y., Dong, X., Hantrakul, L., Hao, W., Huang, Q., Huang, Z., Jia, D., La, F., Le, D., Li, B., Li, C., Li, H., Li, X., Liu, S., Lu, W.-T., Lu, Y., Shaw, A., Spijkervet, J., Sun, Y., Wang, B., Wang, J.-C., Wang, Y., Wang, Y., Xu, L., Yang, Y., Yao, C., Zhang, S., Zhang, Y., Zhang, Y., Zhao, H., Zhao, Z., Zhong, D., Zhou, S., and Zou, P · 2024
Later among the works it cites.
Xtts: a massively multilingual zero-shot text-to-speech model, 2024
Casanova, E., Davis, K., Gölge, E., Göknar, G., Gulea, I., Hart, L., Aljafari, A., Meyer, J., Morais, R., Olayemi, S., and Weber, J · 2024
Later among the works it cites.
Musicldm: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies
Chen, K., Wu, Y., Liu, H., Nezhurina, M., Berg-Kirkpatrick, T., and Dubnov, S · 2024
Later among the works it cites.
Simple and controllable music generation
Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., and Défossez, A · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., et al · 2023
Cited alongside, same era.
Lp-musiccaps: Llm-based pseudo music captioning
Doh, S., Choi, K., Lee, J., and Nam, J · 2023
Cited alongside, same era.
Distil-whisper: Robust knowledge distillation via large-scale pseudo labelling, 2023
Gandhi, S., von Platen, P., and Rush, A. M · 2023
Cited alongside, same era.
Funasr: A fundamental end-to-end speech recognition toolkit
Gao, Z., Li, Z., Wang, J., Luo, H., Shi, X., Chen, M., Li, Y., Zuo, L., Du, Z., Xiao, Z., and Zhang, S · 2023
Cited alongside, same era.
Noise2music: Text-conditioned music generation with diffusion models
Huang, Q., Park, D. S., Wang, T., Denk, T. I., Ly, A., Chen, N., Zhang, Z., Zhang, Z., Yu, J., Frank, C., et al · 2023
Cited alongside, same era.
Hybrid transformers for music source separation
Rouard, S., Massa, F., and Défossez, A · 2023
Cited alongside, same era.
Mo \ \backslash ˆ usai: Text-to-music generation with long-context latent diffusion
Schneider, F., Kamal, O., Jin, Z., and Schölkopf, B · 2023
Cited alongside, same era.
Evans, Z., Parker, J. D., Carr, C., Zukowski, Z., Taylor, J., and Pons, J · 2024
Later among the works it cites.
Text-to-song: Towards controllable music generation incorporating vocals and accompaniment, 2024
Hong, Z., Huang, R., Cheng, X., Wang, Y., Li, R., You, F., Zhao, Z., and Zhang, Z · 2024
Later among the works it cites.
High-fidelity audio compression with improved rvqgan
Kumar, R., Seetharaman, P., Luebs, A., Kumar, I., and Kumar, K · 2024
Later among the works it cites.
Songcreator: Lyrics-based universal song generation, 2024
Lei, S., Zhou, Y., Tang, B., Lam, M. W. Y., Liu, F., Liu, H., Wu, J., Kang, S., Wu, Z., and Meng, H · 2024
Later among the works it cites.
Audioldm 2: Learning holistic audio generation with self-supervised pretraining
Liu, H., Yuan, Y., Liu, X., Mei, X., Kong, Q., Tian, Q., Wang, Y., Wang, W., Wang, Y., and Plumbley, M. D · 2024
Later among the works it cites.
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Lyth, D. and King, S · 2024
Later among the works it cites.
Drop the beat! freestyler for accompaniment conditioned rapping voice generation, 2024
Ning, Z., Wang, S., Jiang, Y., Yao, J., He, L., Pan, S., Ding, J., and Xie, L · 2024
Later among the works it cites.
Codec does matter: Exploring the semantic shortcoming of codec for audio language model
Ye, Z., Sun, P., Lei, J., Lin, H., Tan, X., Dai, Z., Kong, Q., Chen, J., Pan, J., Liu, Q., Guo, Y., and Xue, W · 2024
Later among the works it cites.
Musicgen-stem: Multi-stem music generation and edition through autoregressive modeling, 2025
Rouard, S., Roman, R. S., Adi, Y., and Roebel, A · 2025
Closest in time.
Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound
Tjandra, A., Wu, Y.-C., Guo, B., Hoffman, J., Ellis, B., Vyas, A., Shi, B., Chen, S., Le, M., Zacharov, N., Wood, C., Lee, A., and Hsu, W.-N · 2025
Closest in time.
Wu, S., Guo, Z., Yuan, R., Jiang, J., Doh, S., Xia, G., Nam, J., Li, X., Yu, F., and Sun, M · 2025
Closest in time.
Yue: Scaling open foundation models for long-form music generation, 2025
Yuan, R., Lin, H., Guo, S., Zhang, G., Pan, J., Zang, Y., Liu, H., Liang, Y., Ma, W., Du, X., Du, X., Ye, Z., Zheng, T., Jiang, Z., Ma, Y., Liu, M., Tian, Z., Zhou, Z., Xue, L., Qu, X., Li, Y., Wu, S., Shen, T., Ma, Z., Zhan, J., Wang, C., Wang, Y., Chi, X., Zhang, X., Yang, Z., Wang, X., Liu, S., Mei, L., Li, P., Wang, J., Yu, J., Pang, G., Li, X., Wang, Z., Zhou, X., Yu, L., Benetos, E., Chen, Y., Lin, C., Chen, X., Xia, G., Zhang, Z., Zhang, C., Chen, W., Zhou, X., Qiu, X., Dannenberg, R., Liu, J., Yang, J., Huang, W., Xue, W., Tan, X., and Guo, Y · 2025
Closest in time.