Fetching the paper…
Reading the bibliography…
Text-to-song (TTSong) is a music generation task that synthesizes accompanied singing voices.
Cognitive foundations of musical pitch
C. L. Krumhansl · 2001
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Musegan: Multi-track sequential generative adversarial networks for symbolic music generation and accompaniment
H.-W. Dong, W.-Y. Hsiao, L.-C. Yang, and Y.-H. Yang · 2018
Earlier work this paper cites.
Deep-rhythm for global tempo estimation in music
H. F. Aarabi and G. Peeters · 2019
Earlier work this paper cites.
Resemblyzer
resemble ai · 2019
Earlier work this paper cites.
Ultimate vocal remover
Anjok07 and aufr33 · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
J. Kong, J. Kim, and J. Bae · 2020
Earlier work this paper cites.
Multi-singer: Fast multi-singer singing voice vocoder with a large-scale corpus
R. Huang, F. Chen, Y. Ren, J. Liu, C. Cui, and Z. Zhao · 2021
Earlier work this paper cites.
Songmass: Automatic song writing with pre-training and alignment constraint
Z. Sheng, K. Song, X. Tan, Y. Ren, W. Ye, S. Zhang, and T. Qin · 2021
Earlier work this paper cites.
Silero vad: pre-trained enterprise-grade voice activity detector (vad), number detector and language classifier
S. Team · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le · 2021
Earlier work this paper cites.
Soundstream: An end-to-end neural audio codec, 2021
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi · 2021
Earlier work this paper cites.
Mulan: A joint embedding of music audio and natural language
Q. Huang, A. Jansen, J. Lee, R. Ganti, J. Y. Li, and D. P. Ellis · 2022
Earlier work this paper cites.
Diffsinger: Singing voice synthesis via shallow diffusion mechanism
J. Liu, C. Li, Y. Ren, F. Chen, and Z. Zhao · 2022
Earlier work this paper cites.
Introducing chatgpt
OpenAI · 2022
Cited alongside, same era.
Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech
Y. Ren, M. Lei, Z. Huang, S. Zhang, Q. Chen, Z. Yan, and Z. Zhao · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis
Y. Wang, X. Wang, P. Zhu, J. Wu, H. Li, H. Xue, Y. Zhang, L. Xie, and M. Bi · 2022
Cited alongside, same era.
Museformer: Transformer with fine-and coarse-grained attention for music generation
B. Yu, P. Lu, R. Wang, W. Hu, X. Tan, W. Ye, S. Zhang, T. Qin, and T.-Y. Liu · 2022
Cited alongside, same era.
Audioldm: Text-to-audio generation with latent diffusion models
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley · 2023
Later among the works it cites.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Later among the works it cites.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
K. Shen, Z. Ju, X. Tan, Y. Liu, Y. Leng, L. He, T. Qin, S. Zhao, and J. Bian · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Neural codec language models are zero-shot text to speech synthesizers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Zhang, R. Li, S. Wang, L. Deng, J. Liu, Y. Ren, J. He, R. Huang, J. Zhu, X. Chen, et al · 2022
Cited alongside, same era.
Musiclm: Generating music from text
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, et al · 2023
Cited alongside, same era.
Whisperx: Time-accurate speech transcription of long-form audio
M. Bain, J. Huh, T. Han, and A. Zisserman · 2023
Cited alongside, same era.
Audiolm: a language modeling approach to audio generation
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov, O. Pietquin, M. Sharifi, D. Roblek, O. Teboul, D. Grangier, M. Tagliasacchi, et al · 2023
Cited alongside, same era.
Lp-musiccaps: Llm-based pseudo music captioning
S. Doh, K. Choi, J. Lee, and J. Nam · 2023
Cited alongside, same era.
Singsong: Generating musical accompaniments from singing
C. Donahue, A. Caillon, A. Roberts, E. Manilow, P. Esling, A. Agostinelli, M. Verzetti, I. Simon, O. Pietquin, N. Zeghidour, et al · 2023
Cited alongside, same era.
Clap learning audio concepts from natural language supervision
B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang · 2023
Cited alongside, same era.
C. Wang, S. Chen, Y. Wu, Z. Zhang, L. Zhou, S. Liu, Z. Chen, Y. Liu, H. Wang, J. Li, et al · 2023
Later among the works it cites.
Uniaudio: An audio foundation model toward universal audio generation
D. Yang, J. Tian, X. Tan, R. Huang, S. Liu, X. Chang, J. Shi, S. Zhao, J. Bian, X. Wu, et al · 2023
Later among the works it cites.
Midi-voice: Expressive zero-shot singing voice synthesis via midi-driven priors
D.-M. Byun, S.-H. Lee, J.-S. Hwang, and S.-W. Lee · 2024
Closest in time.
Musicldm: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies
K. Chen, Y. Wu, H. Liu, M. Nezhurina, T. Berg-Kirkpatrick, and S. Dubnov · 2024
Closest in time.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, et al · 2024
Closest in time.
Simple and controllable music generation
J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez · 2024
Closest in time.
Songcomposer: A large language model for lyric and melody composition in song generation
S. Ding, Z. Liu, X. Dong, P. Zhang, R. Qian, C. He, D. Lin, and J. Wang · 2024
Closest in time.
Efficient neural music generation
M. W. Lam, Q. Tian, T. Li, Z. Yin, S. Feng, M. Tu, Y. Ji, R. Xia, M. Ma, X. Song, et al · 2024
Closest in time.
Robust singing voice transcription serves synthesis, 2024
R. Li, Y. Zhang, Y. Wang, Z. Hong, R. Huang, and Z. Zhao · 2024
Closest in time.
Prompt-singer: Controllable singing-voice-synthesis with natural language prompt
Y. Wang, R. Hu, R. Huang, Z. Hong, R. Li, W. Liu, F. You, T. Jin, and Z. Zhao · 2024
Closest in time.
Stylesinger: Style transfer for out-of-domain singing voice synthesis
Y. Zhang, R. Huang, R. Li, J. He, Y. Xia, F. Chen, X. Duan, B. Huai, and Z. Zhao · 2024
Closest in time.
Text-to-song: Towards controllable music generation incorporating vocals and accompaniment
H. Zhiqing, H. Rongjie, C. Xize, W. Yongqi, L. Ruiqi, Y. Fuming, Z. Zhou, and Z. Zhimeng · 2024
Closest in time.