Fetching the paper…
Reading the bibliography…
We tackle the task of long-form music generation--particularly the challenging \textbf{lyrics-to-song} problem--by introducing YuE, a family of open foundation models based on the LLaMA2 architecture.
A Generative Theory of Tonal Music, reissue, with a new preface
F. Lerdahl and R. S. Jackendoff · 1996
Earlier work this paper cites.
Jukebox: A generative model for music
P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever · 2005
Earlier work this paper cites.
Jukebox: A generative model for music
P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever · 2005
Earlier work this paper cites.
The perception of structural boundaries in melody lines of western popular music
M. J. Bruderer, M. F. McKinney, and A. Kohlrausch · 2009
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
S. Hershey, S. Chaudhuri, D. P. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, et al · 2017
Earlier work this paper cites.
C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, N. Shazeer, I. Simon, C. Hawthorne, A. M. Dai, M. D. Hoffman, M. Dinculescu, and D. Eck · 2018
Earlier work this paper cites.
The mtg-jamendo dataset for automatic music tagging
D. Bogdanov, M. Won, P. Tovstogan, A. Porter, and X. Serra · 2019
Earlier work this paper cites.
Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi · 2019
Earlier work this paper cites.
Universality and diversity in human song
S. A. Mehr, M. Singh, D. Knox, D. M. Ketter, D. Pickens-Jones, S. Atwood, C. Lucas, N. Jacoby, A. A. Egner, E. J. Hopkins, et al · 2019
Earlier work this paper cites.
Wav2Vec: Unsupervised pre-training for speech recognition
S. Schneider, A. Baevski, R. Collobert, and M. Auli · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli · 2020
Earlier work this paper cites.
Hifisinger: Towards high-fidelity neural singing voice synthesis
J. Chen, X. Tan, J. Luan, T. Qin, and T.-Y. Liu · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann · 2020
Earlier work this paper cites.
Audio-based music structure analysis: Current trends, open challenges, and applications
O. Nieto, G. J. Mysore, C.-i. Wang, J. B. Smith, J. Schlüter, T. Grill, and B. McFee · 2020
Earlier work this paper cites.
W2V-Bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Y.-A. Chung, Y. Zhang, W. Han, C.-C. Chiu, J. Qin, R. Pang, and Y. Wu · 2021
Earlier work this paper cites.
Efficient training of audio transformers with patchout
K. Koutini, J. Schlüter, H. Eghbal-Zadeh, and G. Widmer · 2021
Earlier work this paper cites.
Data2Vec: A general framework for self-supervised learning in speech, vision and language
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli · 2022
Earlier work this paper cites.
High fidelity neural audio compression
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi · 2022
Earlier work this paper cites.
Bytecover2: Towards dimensionality reduction of latent embedding for efficient cover song identification
X. Du, K. Chen, Z. Wang, B. Zhu, and Z. Ma · 2022
Earlier work this paper cites.
Diffsinger: Singing voice synthesis via shallow diffusion mechanism
J. Liu, C. Li, Y. Ren, F. Chen, and Z. Zhao · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Earlier work this paper cites.
ViSinger: Variational inference with adversarial learning for end-to-end singing voice synthesis
Y. Zhang, J. Cong, H. Xue, L. Xie, P. Zhu, and M. Bi · 2022
Earlier work this paper cites.
MusicLM: Generating music from text
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, et al · 2023
Cited alongside, same era.
Audiolm: A language modeling approach to audio generation
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, D. Roblek, O. Teboul, D. Grangier, M. Tagliasacchi, and N. Zeghidour · 2023
Cited alongside, same era.
Musicldm: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies
K. Chen, Y. Wu, H. Liu, M. Nezhurina, T. Berg-Kirkpatrick, and S. Dubnov · 2023
Cited alongside, same era.
Singsong: Generating musical accompaniments from singing
C. Donahue, A. Caillon, A. Roberts, E. Manilow, P. Esling, A. Agostinelli, M. Verzetti, I. Simon, O. Pietquin, N. Zeghidour, et al · 2023
Cited alongside, same era.
Bytecover3: Accurate cover song identification on short queries
X. Du, Z. Wang, X. Liang, H. Liang, B. Zhu, and Z. Ma · 2023
Accompanied singing voice synthesis with fully text-controlled melody
R. Li, Z. Hong, Y. Wang, L. Zhang, R. Huang, S. Zheng, and Z. Zhao · 2024
Later among the works it cites.
SemantiCodec: An ultra low bitrate semantic audio codec for general sound
H. Liu, X. Xu, Y. Yuan, M. Wu, W. Wang, and M. D. Plumbley · 2024
Later among the works it cites.
AudioLDM 2: Learning holistic audio generation with self-supervised pretraining
H. Liu, Y. Yuan, X. Liu, X. Mei, Q. Kong, Q. Tian, Y. Wang, W. Wang, Y. Wang, and M. D. Plumbley · 2024
Later among the works it cites.
Foundation models for music: A survey
Y. Ma, A. Øland, A. Ragni, B. M. Del Sette, C. Saitis, C. Donahue, C. Lin, C. Plachouras, E. Benetos, E. Shatri, et al · 2024
Later among the works it cites.
Scaling transformers for low-bitrate high-quality speech coding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
UniSinger: Unified end-to-end singing voice synthesis with cross-modality information matching
Z. Hong, C. Cui, R. Huang, L. Zhang, J. Liu, J. He, and Z. Zhao · 2023
Cited alongside, same era.
Noise2Music: Text-conditioned music generation with diffusion models
Q. Huang, D. S. Park, T. Wang, T. I. Denk, A. Ly, N. Chen, Z. Zhang, Z. Zhang, J. Yu, C. Frank, et al · 2023
Cited alongside, same era.
All-in-one metrical and functional structure analysis with neighborhood attentions on demixed audio
T. Kim and J. Nam · 2023
Cited alongside, same era.
MERT: Acoustic music understanding model with large-scale self-supervised training
Y. Li, R. Yuan, G. Zhang, Y. Ma, X. Chen, H. Yin, C. Lin, A. Ragni, E. Benetos, N. Gyenge, et al · 2023
Cited alongside, same era.
AudioLDM: Text-to-audio generation with latent diffusion models
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley · 2023
Cited alongside, same era.
MT4SSL: Boosting self-supervised speech representation learning by integrating multiple targets
Z. Ma, Z. Zheng, C. Tang, Y. Wang, and X. Chen · 2023
Cited alongside, same era.
H. Siuzdak · 2023
Cited alongside, same era.
J. D. Parker, A. Smirnov, J. Pons, C. Carr, Z. Zukowski, Z. Evans, and X. Liu · 2024
Later among the works it cites.
Mupt: A generative symbolic music pretrained transformer
X. Qu, Y. Bai, Y. Ma, Z. Zhou, K. M. Lo, J. Liu, R. Yuan, L. Min, X. Liu, T. Zhang, et al · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2024
Later among the works it cites.
Introducing meta llama 3: The most capable openly available llm to date, 2024
T. L. Team · 2024
Later among the works it cites.
Prompt-Singer: Controllable singing-voice-synthesis with natural language prompt
Y. Wang, R. Hu, R. Huang, Z. Hong, R. Li, W. Liu, F. You, T. Jin, and Z. Zhao · 2024
Later among the works it cites.
A foundation model for music informatics
M. Won, Y.-N. Hung, and D. Le · 2024
Later among the works it cites.
TokSing: Singing voice synthesis based on discrete tokens
Y. Wu, J. Shi, Y. Tang, S. Yang, Q. Jin, et al · 2024
Later among the works it cites.
Bigcodec: Pushing the limits of low-bitrate neural speech codec, 2024
D. Xin, X. Tan, S. Takamichi, and H. Saruwatari · 2024
Later among the works it cites.
Codec does matter: Exploring the semantic shortcoming of codec for audio language model
Z. Ye, P. Sun, J. Lei, H. Lin, X. Tan, Z. Dai, Q. Kong, J. Chen, J. Pan, Q. Liu, et al · 2024
Later among the works it cites.
Map-neo: Highly capable and transparent bilingual large language model series
G. Zhang, S. Qu, J. Liu, C. Zhang, C. Lin, C. L. Yu, D. Pan, E. Cheng, J. Liu, Q. Lin, et al · 2024
Later among the works it cites.
Learnings from scaling visual tokenizers for reconstruction and generation
P. Hansen-Estruch, D. Yan, C.-Y. Chung, O. Zohar, J. Wang, T. Hou, T. Xu, S. Vishwanath, P. Vajda, and X. Chen · 2025
Closest in time.
Songcreator: Lyrics-based universal song generation
S. Lei, Y. Zhou, B. Tang, M. W. Lam, H. Liu, J. Wu, S. Kang, Z. Wu, H. Meng, et al · 2025
Closest in time.
Songgen: A single stage auto-regressive transformer for text-to-song generation
Z. Liu, S. Ding, Z. Zhang, X. Dong, P. Zhang, Y. Zang, Y. Cao, D. Lin, and J. Wang · 2025
Closest in time.
Audio-CoT: Exploring chain-of-thought reasoning in large audio language model
Z. Ma, Z. Chen, Y. Wang, E. S. Chng, and X. Chen · 2025
Closest in time.
Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound
A. Tjandra, Y.-C. Wu, B. Guo, J. Hoffman, B. Ellis, A. Vyas, B. Shi, S. Chen, M. Le, N. Zacharov, et al · 2025
Closest in time.
Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens
X. Wang, M. Jiang, Z. Ma, Z. Zhang, S. Liu, L. Li, Z. Liang, Q. Zheng, R. Wang, X. Feng, et al · 2025
Closest in time.
S. Wu, Z. Guo, R. Yuan, J. Jiang, S. Doh, G. Xia, J. Nam, X. Li, F. Yu, and M. Sun · 2025
Closest in time.
Muq: Self-supervised music representation learning with mel residual vector quantization
H. Zhu, Y. Zhou, H. Chen, J. Yu, Z. Ma, R. Gu, W. Tan, and X. Chen · 2025
Closest in time.