Fetching the paper…
Reading the bibliography…
In speech generation tasks, human subjective ratings, usually referred to as the opinion score, are considered the "gold standard" for speech quality evaluation, with the mean opinion score (MOS) serving as the primary evaluation metric.
“Mel-cepstral distance measure for objective speech quality assessment”
Robert Kubichek · 1993
Earlier work this paper cites.
“The reliability of the ITU-P. 85 standard for the evaluation of text-to-speech systems”
Yolanda Vazquez-Alvarez and Mark Huckvale · 2002
Earlier work this paper cites.
“Tacotron: Towards End-to-End Speech Synthesis”
Yuxuan Wang et al · 2017
Earlier work this paper cites.
“MOSNet: Deep Learning-Based Objective Assessment for Voice Conversion”
Chen-Chou Lo et al · 2019
Earlier work this paper cites.
“FastSpeech 2: Fast and High-Quality End-to-End Text to Speech”
Yi Ren et al · 2020
Earlier work this paper cites.
“wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations”
Alexei Baevski et al · 2020
Earlier work this paper cites.
“JVS-MuSiC: Japanese multispeaker singing-voice corpus”
Hiroki Tamaru et al · 2020
Earlier work this paper cites.
“XiaoiceSing: A High-Quality and Integrated Singing Voice Synthesis System”
Peiling Lu, Jie Wu and Jian Luan · 2020
Earlier work this paper cites.
“HuBERT: How much can a bad teacher benefit ASR pre-training”
Wei-Ning Hsu et al · 2020
Earlier work this paper cites.
“HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis”
Jungil Kong, Jaehyeon Kim and Jaekyoung Bae · 2020
Earlier work this paper cites.
“MBNet: MOS prediction for synthesized speech with mean-bias network”
Yichong Leng et al · 2021
Earlier work this paper cites.
“Utilizing Self-supervised Representations for MOS Prediction”
Wei-Cheng Tseng et al · 2021
Earlier work this paper cites.
“Tohoku Kiritan singing database: A singing database for statistical parametric singing synthesis using Japanese pop songs”
Itsuki Ogawa and Masanori Morise · 2021
Earlier work this paper cites.
“Sequence-To-Sequence Singing Voice Synthesis With Perceptual Entropy Loss”
Jiatong Shi et al · 2021
Earlier work this paper cites.
“LDNet: Unified listener dependent modeling in MOS prediction for synthetic speech”
Wen-Chin Huang et al · 2022
Earlier work this paper cites.
“Generalization ability of MOS prediction networks”
Erica Cooper et al · 2022
Earlier work this paper cites.
“UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022”
Takaaki Saeki et al · 2022
Cited alongside, same era.
“DDOS: A MOS Prediction Framework utilizing Domain Adaptive Pre-training and Distribution of Opinion Scores”
Wei-Cheng Tseng, Wei-Tsung Kao and Hung-yi Lee · 2022
Cited alongside, same era.
“Fusion of Self-supervised Learned Models for MOS Prediction”
Zhengdong Yang et al · 2022
Cited alongside, same era.
“The VoiceMOS challenge 2022”
Wen-Chin Huang et al · 2022
Cited alongside, same era.
“Opencpop: A High-Quality Open Source Chinese Popular Song Corpus for Singing Voice Synthesis”
Yu Wang et al · 2022
Cited alongside, same era.
“M4Singer: A Multi-Style, Multi-Singer and Musical Score Provided Mandarin Singing Corpus”
Lichao Zhang et al · 2022
Cited alongside, same era.
“The singing voice conversion challenge 2023”
Wen-Chin Huang et al · 2023
Later among the works it cites.
“VISinger2: High-Fidelity End-to-End Singing Voice Synthesis Enhanced by Digital Signal Processing Synthesizer”
Yongmao Zhang et al · 2023
Later among the works it cites.
“A Comparative Study of Voice Conversion Models with Large-Scale Speech and Singing Data: The T13 Systems for the Singing Voice Conversion Challenge 2023”
Ryuichi Yamamoto et al · 2023
Later among the works it cites.
“On the utility of self-supervised models for prosody-related tasks”
Guan-Ting Lin et al · 2023
Later among the works it cites.
“Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models”
Zeqian Ju et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Muskits: an End-to-end Music Processing Toolkit for Singing Voice Synthesis”
Jiatong Shi et al · 2022
Cited alongside, same era.
“Diffsinger: Singing voice synthesis via shallow diffusion mechanism”
Jinglin Liu et al · 2022
Cited alongside, same era.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing”
Sanyuan Chen et al · 2022
Cited alongside, same era.
“Contentvec: An improved self-supervised speech representation by disentangling speakers”
Kaizhi Qian et al · 2022
Cited alongside, same era.
“XLS-R: Self-supervised cross-lingual speech representation learning at scale”
Arun Babu et al · 2022
Cited alongside, same era.
“Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis”
Hubert Siuzdak · 2023
Cited alongside, same era.
“The Interspeech 2024 Challenge on Speech Processing Using Discrete Units”
Xuankai Chang et al · 2024
Closest in time.
“ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models”
Jee-weon Jung et al · 2024
Closest in time.
“USAT: A Universal Speaker-Adaptive Text-to-Speech Approach”
Wenbin Wang, Yang Song and Sanjay Jha · 2024
Closest in time.
“CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection”
Yongyi Zang et al · 2024
Closest in time.
“Singing Voice Data Scaling-up: An Introduction to ACE-Opencpop and ACE-KiSing”
Jiatong Shi et al · 2024
Closest in time.
“NNSVS: A Neural Network-Based Singing Voice Synthesis Toolkit”
Ryuichi Yamamoto, Reo Yoneyama and Tomoki Toda · 2024
Closest in time.
“The Interspeech 2024 Challenge on Speech Processing Using Discrete Units”
Xuankai Chang et al · 2024
Closest in time.
“TokSing: Singing Voice Synthesis based on Discrete Tokens”
Yuning Wu et al · 2024
Closest in time.
“SingOMD: Singing Oriented Multi-resolution Discrete Representation Construction from Speech Models”
Yuxun Tang et al · 2024
Closest in time.
“Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction”
Jiatong Shi et al · 2024
Closest in time.