Fetching the paper…
Reading the bibliography…
Singing Voice Synthesis (SVS) has witnessed significant advancements with the advent of deep learning techniques.
“Singing voice synthesis combining excitation plus resonance and sinusoidal plus residual models,”
Jordi Bonada, Òscar Celma Herrada, Àlex Loscos, Jaume Ortolà, Xavier Serra, Yasuo Yoshioka, Hiraku Kayama, Yuji Hisaminato, and Hideki Kenmochi, · 2001
Earlier work this paper cites.
“Sample-based singing voice synthesizer by spectral concatenation,”
Jordi Bonada, Alex Loscos, and H Kenmochi, · 2003
Earlier work this paper cites.
“Vocaloid - commercial singing synthesizer based on sample concatenation.,”
Hideki Kenmochi and Hayato Ohshita, · 2007
Earlier work this paper cites.
“XiaoiceSing: A high-quality and integrated singing voice synthesis system,”
Peiling Lu, Jie Wu, Jian Luan, et al., · 2020
Earlier work this paper cites.
“Hifisinger: Towards high-fidelity neural singing voice synthesis,”
Jiawei Chen, Xu Tan, Jian Luan, et al., · 2020
Earlier work this paper cites.
“HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,”
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae, · 2020
Earlier work this paper cites.
“Improved RawNet with Feature Map Scaling for Text-Independent Speaker Verification Using Raw Waveforms,”
Jee weon Jung, Seung bin Kim, Hye jin Shim, Ju ho Kim, and Ha-Jin Yu, · 2020
Earlier work this paper cites.
“Sequence-to-sequence singing voice synthesis with perceptual entropy loss,”
Jiatong Shi, Shuai Guo, Nan Huo, Yuekai Zhang, and Qin Jin, · 2021
Earlier work this paper cites.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Earlier work this paper cites.
“SUPERB: Speech Processing Universal PERformance Benchmark,”
Shu-Wen Yang, Po-Han Chi, Yung-Sung Chuang, et al., · 2021
Earlier work this paper cites.
“Layer-wise analysis of a self-supervised speech representation model,”
Ankita Pasad, Ju-Chieh Chou, and Karen Livescu, · 2021
Earlier work this paper cites.
“UnivNet: A Neural Vocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation,”
Won Jang, Dan Lim, Jaesam Yoon, Bongwan Kim, and Juntae Kim, · 2021
Earlier work this paper cites.
“VISinger: Variational inference with adversarial learning for end-to-end singing voice synthesis,”
Yongmao Zhang, Jian Cong, Heyang Xue, et al., · 2022
Earlier work this paper cites.
“DiffSinger: Singing voice synthesis via shallow diffusion mechanism,”
Jinglin Liu, Chengxi Li, Yi Ren, et al., · 2022
Cited alongside, same era.
“Singaug: Data augmentation for singing voice synthesis with cycle-consistent training strategy,”
Shuai Guo, Jiatong Shi, Tao Qian, Shinji Watanabe, and Qin Jin, · 2022
Cited alongside, same era.
“Data2vec: A general framework for self-supervised learning in speech, vision and language,”
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli, · 2022
Cited alongside, same era.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2022
Cited alongside, same era.
“Opencpop: A High-Quality Open Source Chinese Popular Song Corpus for Singing Voice Synthesis,”
Yu Wang, Xinsheng Wang, Pengcheng Zhu, Jie Wu, Hanzhao Li, Heyang Xue, Yongmao Zhang, Lei Xie, and Mengxiao Bi, · 2022
Cited alongside, same era.
“SALTTS: Leveraging Self-Supervised Speech Representations for improved Text-to-Speech Synthesis,”
Ramanan Sivaguru, Vasista Sai Lodagala, and S Umesh, · 2023
Later among the works it cites.
Cheng Gong, Xin Wang, Erica Cooper, Dan Wells, Longbiao Wang, Jianwu Dang, Korin Richmond, and Junichi Yamagishi, · 2023
Later among the works it cites.
“The singing voice conversion challenge 2023,”
Wen-Chin Huang, Lester Phillip Violeta, Songxiang Liu, Jiatong Shi, and Tomoki Toda, · 2023
Later among the works it cites.
“Singing voice data scaling-up: An introduction to ACE-Opencpop and KiSing-v2,”
Jiatong Shi, Yueqian Lin, Xinyi Bai, Keyi Zhang, Yuning Wu, Yuxun Tang, Yifeng Yu, Qin Jin, and Shinji Watanabe, · 2024
Closest in time.
“Multi-resolution HuBERT: Multi-resolution speech self-supervised learning with masked unit prediction,”
Jiatong Shi, Hirofumi Inaguma, Xutai Ma, Ilia Kulikov, and Anna Sun, · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“DB Production: Futon P,” https://sites.google.com/view/oftn-utagoedb/%E3%83%9B%E3%83%BC%E3%83%A0,
P Futon, · 2022
Cited alongside, same era.
“Muskits: an end-to-end music processing toolkit for singing voice synthesis,”
Jiatong Shi, Shuai Guo, Tao Qian, et al., · 2022
Cited alongside, same era.
“VISinger2: High-Fidelity End-to-End Singing Voice Synthesis Enhanced by Digital Signal Processing Synthesizer,”
Yongmao Zhang, Heyang Xue, Hanzhao Li, Lei Xie, Tingwei Guo, Ruixiong Zhang, and Caixia Gong, · 2023
Cited alongside, same era.
“Xiaoicesing 2: A High-Fidelity Singing Voice Synthesizer Based on Generative Adversarial Network,”
Wang Chunhui, Chang Zeng, and Xing He, · 2023
Cited alongside, same era.
“A systematic exploration of joint-training for singing voice synthesis,”
Yuning Wu, Yifeng Yu, Jiatong Shi, Tao Qian, and Qin Jin, · 2023
Cited alongside, same era.
“MERT: Acoustic music understanding model with large-scale self-supervised training,”
LI Yizhi, Ruibin Yuan, Ge Zhang, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin, Anton Ragni, Emmanouil Benetos, et al., · 2023
Cited alongside, same era.
“ML-SUPERB: Multilingual Speech Universal PERformance Benchmark,”
Jiatong Shi, Dan Berrebbi, William Chen, et al., · 2023
Cited alongside, same era.
Closest in time.
“MARBLE: Music audio representation benchmark for universal evaluation,”
Ruibin Yuan, Yinghao Ma, Yizhi Li, Ge Zhang, Xingran Chen, Hanzhi Yin, Yiqi Liu, Jiawen Huang, Zeyue Tian, Binyue Deng, et al., · 2024
Closest in time.
“Low-resource cross-domain singing voice synthesis via reduced self-supervised speech representations,”
Panos Kakoulidis, Nikolaos Ellinas, Georgios Vamvoukakis, Myrsini Christidou, Alexandra Vioni, Georgia Maniati, Junkwang Oh, Gunu Jho, Inchul Hwang, Pirros Tsiakoulis, et al., · 2024
Closest in time.
“Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters,”
Kenichi Fujita, Hiroshi Sato, Takanori Ashihara, Hiroki Kanagawa, Marc Delcroix, Takafumi Moriya, and Yusuke Ijima, · 2024
Closest in time.
“ParrotTTS: Text-to-speech synthesis exploiting disentangled self-supervised representations,”
Neil Shah, Saiteja Kosgi, Vishal Tambrahalli, Neha Sahipjohn, Anil Kumar Nelakanti, and Vineet Gandhi, · 2024
Closest in time.
“Towards universal speech discrete tokens: A case study for asr and tts,”
Yifan Yang, Feiyu Shen, Chenpeng Du, Ziyang Ma, Kai Yu, Daniel Povey, and Xie Chen, · 2024
Closest in time.
“Muskits-espnet: A comprehensive toolkit for singing voice synthesis in new paradigm,”
Yuning Wu, Jiatong Shi, Yifeng Yu, Yuxun Tang, Tao Qian, Yueqian Lin, Jionghao Han, Xinyi Bai, Shinji Watanabe, and Qin Jin, · 2024
Closest in time.
“Singmos: An extensive open-source singing voice dataset for mos prediction,” 2024
Yuxun Tang, Jiatong Shi, Yuning Wu, and Qin Jin, · 2024
Closest in time.
Jee-weon Jung, Wangyou Zhang, Jiatong Shi, Zakaria Aldeneh, Takuya Higuchi, Barry-John Theobald, Ahmed Hussen Abdelaziz, and Shinji Watanabe, · 2024
Closest in time.