Fetching the paper…
Reading the bibliography…
Rap, a prominent genre of vocal performance, remains underexplored in vocal generation.
The NUS sung and spoken lyrics corpus: A quantitative comparison of singing and speech
Duan, Z.; Fang, H.; Li, B.; Sim, K. C.; and Wang, Y. 2013 · 2013
Earlier work this paper cites.
MuseGAN: Multi-track Sequential Generative Adversarial Networks for Symbolic Music Generation and Accompaniment
Dong, H.; Hsiao, W.; Yang, L.; and Yang, Y. 2018 · 2018
Earlier work this paper cites.
Phonemizer: Text to Phones Transcription for Multiple Languages in Python
Bernard, M.; and Titeux, H. 2021 · 2021
Earlier work this paper cites.
Unsupervised Cross-Lingual Representation Learning for Speech Recognition
Conneau, A.; Baevski, A.; Collobert, R.; Mohamed, A.; and Auli, M. 2021 · 2021
Earlier work this paper cites.
Layer-Wise Analysis of a Self-Supervised Speech Representation Model
Pasad, A.; Chou, J.; and Livescu, K. 2021 · 2021
Earlier work this paper cites.
Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech
Popov, V.; Vovk, I.; Gogoryan, V.; Sadekova, T.; and Kudinov, M. A. 2021 · 2021
Earlier work this paper cites.
Dnsmos: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors
Reddy, C. K. A.; Gopal, V.; and Cutler, R. 2021 · 2021
Earlier work this paper cites.
SongMASS: Automatic Song Writing with Pre-training and Alignment Constraint
Sheng, Z.; Song, K.; Tan, X.; Ren, Y.; Ye, W.; Zhang, S.; and Qin, T. 2021 · 2021
Earlier work this paper cites.
DeepRapper: Neural Rap Generation with Rhyme and Rhythm Modeling
Xue, L.; Song, K.; Wu, D.; Tan, X.; Zhang, N. L.; Qin, T.; Zhang, W.; and Liu, T. 2021 · 2021
Earlier work this paper cites.
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
Chen, S.; Wang, C.; Chen, Z.; Wu, Y.; Liu, S.; Chen, Z.; Li, J.; Kanda, N.; Yoshioka, T.; Xiao, X.; Wu, J.; Zhou, L.; Ren, S.; Qian, Y.; Qian, Y.; Wu, J.; Zeng, M.; Yu, X.; and Wei, F. 2022 · 2022
Earlier work this paper cites.
MuLan: A Joint Embedding of Music Audio and Natural Language
Huang, Q.; Jansen, A.; Lee, J.; Ganti, R.; Li, J. Y.; and Ellis, D. P. W. 2022 · 2022
Earlier work this paper cites.
DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism
Liu, J.; Li, C.; Ren, Y.; Chen, F.; and Zhao, Z. 2022 · 2022
Earlier work this paper cites.
High-Resolution Image Synthesis with Latent Diffusion Models
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Ommer, B. 2022 · 2022
Earlier work this paper cites.
Opencpop: A High-Quality Open Source Chinese Popular Song Corpus for Singing Voice Synthesis
Wang, Y.; Wang, X.; Zhu, P.; Wu, J.; Li, H.; Xue, H.; Zhang, Y.; Xie, L.; and Bi, M. 2022 · 2022
Cited alongside, same era.
M4Singer: A Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing Corpus
Zhang, L.; Li, R.; Wang, S.; Deng, L.; Liu, J.; Ren, Y.; He, J.; Huang, R.; Zhu, J.; Chen, X.; and Zhao, Z. 2022 · 2022
Cited alongside, same era.
MusicLM: Generating Music From Text
Agostinelli, A.; Denk, T. I.; Borsos, Z.; Engel, J. H.; Verzetti, M.; Caillon, A.; Huang, Q.; Jansen, A.; Roberts, A.; Tagliasacchi, M.; Sharifi, M.; Zeghidour, N.; and Frank, C. H. 2023 · 2023
Cited alongside, same era.
AudioLM: A Language Modeling Approach to Audio Generation
Borsos, Z.; Marinier, R.; Vincent, D.; Kharitonov, E.; Pietquin, O.; Sharifi, M.; Roblek, D.; Teboul, O.; Grangier, D.; Tagliasacchi, M.; and Zeghidour, N. 2023 · 2023
Cited alongside, same era.
RMSSinger: Realistic-Music-Score based Singing Voice Synthesis
He, J.; Liu, J.; Ye, Z.; Huang, R.; Cui, C.; Liu, H.; and Zhao, Z. 2023 · 2023
Musicldm: Enhancing novelty in text-to-music generation using beat-synchronous mixup strategies
Chen, K.; Wu, Y.; Liu, H.; Nezhurina, M.; Berg-Kirkpatrick, T.; and Dubnov, S. 2024 · 2024
Closest in time.
SongComposer: A Large Language Model for Lyric and Melody Composition in Song Generation
Ding, S.; Liu, Z.; Dong, X.; Zhang, P.; Qian, R.; He, C.; Lin, D.; and Wang, J. 2024 · 2024
Closest in time.
Multi-Scale Sub-Band Constant-Q Transform Discriminator for High-Fidelity Vocoder
Gu, Y.; Zhang, X.; Xue, L.; and Wu, Z. 2024 · 2024
Closest in time.
Text-to-Song: Towards Controllable Music Generation Incorporating Vocals and Accompaniment
Hong, Z.; Huang, R.; Cheng, X.; Wang, Y.; Li, R.; You, F.; Zhao, Z.; and Zhang, Z. 2024 · 2024
Closest in time.
VoiceTuner: Self-Supervised Pre-training and Efficient Fine-tuning For Voice Generation
Huang, R.; Wang, Y.; Hu, R.; Xu, X.; Hong, Z.; Yang, D.; Cheng, X.; Wang, Z.; Jiang, Z.; Ye, Z.; et al. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
Kharitonov, E.; Vincent, D.; Borsos, Z.; Marinier, R.; Girgin, S.; Pietquin, O.; Sharifi, M.; Tagliasacchi, M.; and Zeghidour, N. 2023 · 2023
Cited alongside, same era.
High-Fidelity Audio Compression with Improved RVQGAN
Kumar, R.; Seetharaman, P.; Luebs, A.; Kumar, I.; and Kumar, K. 2023 · 2023
Cited alongside, same era.
Efficient Neural Music Generation
Lam, M. W. Y.; Tian, Q.; Li, T.; Yin, Z.; Feng, S.; Tu, M.; Ji, Y.; Xia, R.; Ma, M.; Song, X.; Chen, J.; Wang, Y.; and Wang, Y. 2023 · 2023
Cited alongside, same era.
BigVGAN: A Universal Neural Vocoder with Large-Scale Training
Lee, S.; Ping, W.; Ginsburg, B.; Catanzaro, B.; and Yoon, S. 2023 · 2023
Cited alongside, same era.
Robust Speech Recognition via Large-Scale Weak Supervision
Radford, A.; Kim, J. W.; Xu, T.; Brockman, G.; McLeavey, C.; and Sutskever, I. 2023 · 2023
Cited alongside, same era.
LLaMA: Open and Efficient Foundation Language Models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lample, G. 2023 · 2023
Cited alongside, same era.
Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
Wu, Y.; Chen, K.; Zhang, T.; Hui, Y.; Berg-Kirkpatrick, T.; and Dubnov, S. 2023 · 2023
Cited alongside, same era.
Closest in time.
Stack-and-delay: a new codebook pattern for music generation
Le Lan, G.; Nagaraja, V.; Chang, E.; Kant, D.; Ni, Z.; Shi, Y.; Iandola, F.; and Chandra, V. 2024 · 2024
Closest in time.
Music Source Separation With Band-Split Rope Transformer
Lu, W. T.; Wang, J.; Kong, Q.; and Hung, Y. 2024 · 2024
Closest in time.
WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
Ma, L.; Guo, D.; Song, K.; Jiang, Y.; Wang, S.; Xue, L.; Xu, W.; Zhao, H.; Zhang, B.; and Xie, L. 2024 · 2024
Closest in time.
Matcha-TTS: A fast TTS architecture with conditional flow matching
Mehta, S.; Tu, R.; Beskow, J.; Székely, É.; and Henter, G. E. 2024 · 2024
Closest in time.
NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Shen, K.; Ju, Z.; Tan, X.; Liu, E.; Leng, Y.; He, L.; Qin, T.; Zhao, S.; and Bian, J. 2024 · 2024
Closest in time.
Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt
Wang, Y.; Hu, R.; Huang, R.; Hong, Z.; Li, R.; Liu, W.; You, F.; Jin, T.; and Zhao, Z. 2024 · 2024
Closest in time.
Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts
Yao, J.; Yang, Y.; Lei, Y.; Ning, Z.; Hu, Y.; Pan, Y.; Yin, J.; Zhou, H.; Lu, H.; and Xie, L. 2024 · 2024
Closest in time.
VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
Yu, Y.; Shi, J.; Wu, Y.; and Watanabe, S. 2024 · 2024
Closest in time.