Fetching the paper…
Reading the bibliography…
Recent singing-voice-synthesis (SVS) methods have achieved remarkable audio quality and naturalness, yet they lack the capability to control the style attributes of the synthesized singing explicitly.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
On the Sensations of Tone as a Physiological Basis for the Theory of Music
Hermann Von Helmholtz. 1912 · 1912
Earlier work this paper cites.
Hifisinger: Towards high-fidelity neural singing voice synthesis
Jiawei Chen, Xu Tan, Jian Luan, Tao Qin, and Tie-Yan Liu. 2020 · 2009
Earlier work this paper cites.
Aishell-3: A multi-speaker mandarin tts corpus and the baselines
Yao Shi, Hui Bu, Xin Xu, Shaoji Zhang, and Ming Li. 2020 · 2010
Earlier work this paper cites.
Thchs-30 : A free chinese speech corpus
Zhiyong Zhang Dong Wang, Xuewei Zhang. 2015 · 2015
Earlier work this paper cites.
Montreal forced aligner: Trainable text-speech alignment using kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger. 2017 · 2017
Earlier work this paper cites.
Harvest: A high-performance fundamental frequency estimator from speech signals
Masanori Morise et al. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Didispeech: A large scale mandarin speech corpus
Tingwei Guo, Cheng Wen, Dongwei Jiang, Ne Luo, Ruixiong Zhang, Shuaijiang Zhao, Wubo Li, Cheng Gong, Wei Zou, Kun Han, et al. 2021 · 2021
Earlier work this paper cites.
Multi-singer: Fast multi-singer singing voice vocoder with a large-scale corpus
Rongjie Huang, Feiyang Chen, Yi Ren, Jinglin Liu, Chenye Cui, and Zhou Zhao. 2021 · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi. 2021 · 2021
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. 2022 · 2022
Cited alongside, same era.
Singgan: Generative adversarial network for high-fidelity singing voice generation
Rongjie Huang, Chenye Cui, Feiyang Chen, Yi Ren, Jinglin Liu, Zhou Zhao, Baoxing Huai, and Zhefeng Wang. 2022 · 2022
Cited alongside, same era.
Audiogen: Textually guided audio generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. 2022 · 2022
Clap learning audio concepts from natural language supervision
Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang. 2023a · 2023
Later among the works it cites.
Prompttts: Controllable text-to-speech with text descriptions
Zhifang Guo, Yichong Leng, Yihan Wu, Sheng Zhao, and Xu Tan. 2023 · 2023
Later among the works it cites.
Unisinger: Unified end-to-end singing voice synthesis with cross-modality information matching
Zhiqing Hong, Chenye Cui, Rongjie Huang, Lichao Zhang, Jinglin Liu, Jinzheng He, and Zhou Zhao. 2023 · 2023
Later among the works it cites.
Prompttts 2: Describing and generating voices with text prompt
Yichong Leng, Zhifang Guo, Kai Shen, Xu Tan, Zeqian Ju, Yanqing Liu, Yufei Liu, Dongchao Yang, Leying Zhang, Kaitao Song, et al. 2023 · 2023
Later among the works it cites.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bigvgan: A universal neural vocoder with large-scale training
Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon. 2022 · 2022
Cited alongside, same era.
Diffsinger: Singing voice synthesis via shallow diffusion mechanism
Jinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen, and Zhou Zhao. 2022 · 2022
Cited alongside, same era.
Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis
Yu Wang, Xinsheng Wang, Pengcheng Zhu, Jie Wu, Hanzhao Li, Heyang Xue, Yongmao Zhang, Lei Xie, and Mengxiao Bi. 2022 · 2022
Cited alongside, same era.
Visinger: Variational inference with adversarial learning for end-to-end singing voice synthesis
Yongmao Zhang, Jian Cong, Heyang Xue, Lei Xie, Pengcheng Zhu, and Mengxiao Bi. 2022b · 2022
Cited alongside, same era.
Audiolm: a language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al. 2023 · 2023
Cited alongside, same era.
Natural language supervision for general-purpose audio representations
Benjamin Elizalde, Soham Deshmukh, and Huaming Wang. 2023b
Cited in the paper.
Make-an-audio 2: Temporal-enhanced text-to-audio generation
Jiawei Huang, Yi Ren, Rongjie Huang, Dongchao Yang, Zhenhui Ye, Chen Zhang, Jinglin Liu, Xiang Yin, Zejun Ma, and Zhou Zhao. 2023a
Cited in the paper.
Later among the works it cites.
Neural codec language models are zero-shot text to speech synthesizers
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, et al. 2023 · 2023
Later among the works it cites.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2023 · 2023
Later among the works it cites.
Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts
Jixun Yao, Yuguang Yang, Yi Lei, Ziqian Ning, Yanni Hu, Yu Pan, Jingjing Yin, Hongbin Zhou, Heng Lu, and Lei Xie. 2023 · 2023
Later among the works it cites.
Wesinger 2: fully parallel singing voice synthesis via multi-singer conditional adversarial training
Zewang Zhang, Yibin Zheng, Xinhui Li, and Li Lu. 2023b · 2023
Later among the works it cites.
Megabyte: Predicting million-byte sequences with multiscale transformers
Lili Yu, Dániel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis. 2024 · 2024
Closest in time.