Fetching the paper…
Reading the bibliography…
The lack of a publicly-available large-scale and diverse dataset has long been a significant bottleneck for singing voice applications like Singing Voice Synthesis (SVS) and Singing Voice Conversion (SVC).
Mel-cepstral distance measure for objective speech quality assessment
Robert Kubichek · 1993
Earlier work this paper cites.
STRAIGHT, exploitation of the other aspect of VOCODER: Perceptually isomorphic decomposition of speech sounds
Hideki Kawahara · 2006
Earlier work this paper cites.
On the Improvement of Singing Voice Separation for Monaural Recordings Using the MIR-1K Dataset
Chao-Ling Hsu and Jyh-Shing Roger Jang · 2010
Earlier work this paper cites.
The NUS sung and spoken lyrics corpus: A quantitative comparison of singing and speech
Zhiyan Duan, Haotian Fang, Bo Li, Khe Chai Sim, and Ye Wang · 2013
Earlier work this paper cites.
Automatic identification of emotional cues in Chinese opera singing
Dawn AA Black, Ma Li, and Mi Tian · 2014
Earlier work this paper cites.
Librivox: Free public domain audiobooks
Jodi Kearns · 2014
Earlier work this paper cites.
librosa: Audio and Music Signal Analysis in Python
Brian McFee, Colin Raffel, Dawen Liang, Daniel P. W. Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto · 2015
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
WORLD: a vocoder-based high-quality speech synthesis system for real-time applications
Masanori Morise, Fumiya Yokomori, and Kenji Ozawa · 2016
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew W. Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Introducing Parselmouth: A Python interface to Praat
Y. Jadoul, Bill Thompson, and Bart de Boer · 2018
Earlier work this paper cites.
Efficient Neural Audio Synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Earlier work this paper cites.
VocalSet: A Singing Voice Dataset
Julia Wilkins, Prem Seetharaman, Alison Wahl, and Bryan Pardo · 2018
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Waveglow: A Flow-based Generative Network for Speech Synthesis
Ryan Prenger, Rafael Valle, and Bryan Catanzaro · 2019
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew · 2019
Earlier work this paper cites.
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Earlier work this paper cites.
Children’s song dataset for singing voice research
Soonbeom Choi, Wonil Kim, Saebyul Park, Sangeon Yong, and Juhan Nam · 2020
Earlier work this paper cites.
Multiple F0 Estimation in Vocal Ensembles using Convolutional Neural networks
Helena Cuesta, Brian McFee, and Emilia Gómez · 2020
Earlier work this paper cites.
PJS: phoneme-balanced Japanese singing-voice corpus
Junya Koguchi, Shinnosuke Takamichi, and Masanori Morise · 2020
Earlier work this paper cites.
Deepsinger: Singing voice synthesis with data mined from the web
Yi Ren, Xu Tan, Tao Qin, Jian Luan, Zhou Zhao, and Tie-Yan Liu · 2020
Earlier work this paper cites.
HiFi-GAN: High-Fidelity Denoising and Dereverberation Based on Speech Deep Features in Adversarial Networks
Jiaqi Su, Zeyu Jin, and Adam Finkelstein · 2020
Cited alongside, same era.
w2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training
Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu · 2021
Cited alongside, same era.
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Cited alongside, same era.
Multi-Singer: Fast Multi-Singer Singing Voice Vocoder With A Large-Scale Corpus
Rongjie Huang, Feiyang Chen, Yi Ren, Jinglin Liu, Chenye Cui, and Zhou Zhao · 2021
Cited alongside, same era.
DiffWave: A Versatile Diffusion Model for Audio Synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2021
Cited alongside, same era.
BigVGAN: A Universal Neural Vocoder with Large-Scale Training
Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, and Sungroh Yoon · 2023
Later among the works it cites.
Robust Speech Recognition via Large-Scale Weak Supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Later among the works it cites.
Benchmarks and leaderboards for sound demixing tasks, 2023
Roman Solovyev, Alexander Stempkovskiy, and Tatiana Habruseva · 2023
Later among the works it cites.
Leveraging Content-based Features from Multiple Acoustic Models for Singing Voice Conversion
Xueyao Zhang, Yicheng Gu, Haopeng Chen, Zihao Fang, Lexiao Zou, Liumeng Xue, and Zhizheng Wu · 2023
Later among the works it cites.
LyricWhiz: Robust Multilingual Zero-Shot Lyrics Transcription by Whispering to ChatGPT
Le Zhuo, Ruibin Yuan, Jiahao Pan, Yinghao Ma, Yizhi Li, Ge Zhang, Si Liu, Roger B. Dannenberg, Jie Fu, Chenghua Lin, Emmanouil Benetos, Wenhu Chen, Wei Xue, and Yike Guo · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
DiffSVC: A Diffusion Probabilistic Model for Singing Voice Conversion
Songxiang Liu, Yuewen Cao, Dan Su, and Helen Meng · 2021
Cited alongside, same era.
Tohoku kiritan singing database: A singing database for statistical parametric singing synthesis using japanese pop songs
Itsuki Ogawa and Masanori Morise · 2021
Cited alongside, same era.
NHSS: A speech and singing parallel database
Bidisha Sharma, Xiaoxue Gao, Karthika Vijayan, Xiaohai Tian, and Haizhou Li · 2021
Cited alongside, same era.
WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei · 2022
Cited alongside, same era.
Generalization Ability of MOS Prediction Networks
Erica Cooper, Wen-Chin Huang, Tomoki Toda, and Junichi Yamagishi · 2022
Cited alongside, same era.
SingAug: Data Augmentation for Singing Voice Synthesis with Cycle-consistent Training Strategy
Shuai Guo, Jiatong Shi, Tao Qian, Shinji Watanabe, and Qin Jin · 2022
Cited alongside, same era.
Audio Slicer, 2022
Openvpi · 2022
Cited alongside, same era.
Shihao Chen, Yu Gu, Jie Zhang, Na Li, Rilin Chen, Liping Chen, and Lirong Dai · 2024
Later among the works it cites.
The sound demixing challenge 2023-music demixing track
G. Fabbro, S. Uhlich, C.-H. Lai, W. Choi, M. Martínez-Ramírez, W. Liao, Gadelha I., G. Ramos, E. Hsu, H. Rodrigues, F.-R. Stöter, A. Défossez, Y. Luo, J. Yu, D. Chakraborty, S. Mohanty, R. Solovyev, A. Stempkovskiy, T. Habruseva, N. Goswami, T. Harada, M. Kim, J. H. Lee, Y. Dong, X. Zhang, J. Liu, and Y Mitsufuji · 2024
Later among the works it cites.
Multi-scale sub-band constant-q transform discriminator for high-fidelity vocoder
Yicheng Gu, Xueyao Zhang, Liumeng Xue, and Zhizheng Wu · 2024
Later among the works it cites.
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
Haorui He, Zengqiang Shang, Chaoren Wang, Xuyuan Li, Yicheng Gu, Hua Hua, Liwei Liu, Chen Yang, Jiaqi Li, Peiyang Shi, Yuancheng Wang, Kai Chen, Pengyuan Zhang, and Zhizheng Wu · 2024
Later among the works it cites.
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
Ziyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He, Zhenhui Ye, Shengpeng Ji, Qian Yang, Chen Zhang, Pengfei Wei, Chunfeng Wang, Xiang Yin, Zejun Ma, and Zhou Zhao · 2024
Later among the works it cites.
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Eric Liu, Yichong Leng, Kaitao Song, Siliang Tang, Zhizheng Wu, Tao Qin, Xiangyang Li, Wei Ye, Shikun Zhang, Jiang Bian, Lei He, Jinyu Li, and Sheng Zhao · 2024
Later among the works it cites.
MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
Yizhi Li, Ruibin Yuan, Ge Zhang, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin, Anton Ragni, Emmanouil Benetos, Norbert Gyenge, Roger B. Dannenberg, Ruibo Liu, Wenhu Chen, Gus Xia, Yemin Shi, Wenhao Huang, Zili Wang, Yike Guo, and Jie Fu · 2024
Later among the works it cites.
DiffSinger Community Vocoders, 2024
Openvpi · 2024
Later among the works it cites.
NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Kai Shen, Zeqian Ju, Xu Tan, Eric Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian · 2024
Later among the works it cites.
Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and KiSing-v2
Jiatong Shi, Yueqian Lin, Xinyi Bai, Keyi Zhang, Yuning Wu, Yuxun Tang, Yifeng Yu, Qin Jin, and Shinji Watanabe · 2024
Later among the works it cites.
SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction
Yuxun Tang, Jiatong Shi, Yuning Wu, and Qin Jin · 2024
Later among the works it cites.
SaMoye: Zero-shot Singing Voice Conversion Based on Feature Disentanglement and Synthesis
Zihao Wang, Le Ma, Yan Liu, and Kejun Zhang · 2024
Later among the works it cites.
Amphion: An Open-Source Audio, Music and Speech Generation Toolkit
Xueyao Zhang, Liumeng Xue, Yicheng Gu, Yuancheng Wang, Jiaqi Li, Haorui He, Chaoren Wang, Ting Song, Xi Chen, Zihao Fang, Haopeng Chen, Junan Zhang, Tze Ying Tang, Lexiao Zou, Mingxuan Wang, Jun Han, Kai Chen, Haizhou Li, and Zhizheng Wu · 2024
Later among the works it cites.
SinTechSVS: A Singing Technique Controllable Singing Voice Synthesis System
Junchuan Zhao, Low Qi Hong Chetwin, and Ye Wang · 2024
Later among the works it cites.
FT-GAN: Fine-Grained Tune Modeling for Chinese Opera Synthesis
Meizhen Zheng, Peng Bai, Xiaodong Shi, Xun Zhou, and Yiting Yan · 2024
Later among the works it cites.