Fetching the paper…
Reading the bibliography…
We present our system (denoted as T05) for the VoiceMOS Challenge (VMC) 2024.
“Catastrophic forgetting in connectionist networks,”
R. M. French, · 1999
Earlier work this paper cites.
“The Blizzard Challenge 2008,”
Vasilis Karaiskos, Simon King, Robert AJ Clark, and Catherine Mayo, · 2008
Earlier work this paper cites.
“ImageNet: A large-scale hierarchical image database,”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, · 2009
Earlier work this paper cites.
“The Blizzard Challenge 2009,”
Alan W Black, Simon King, and Keiichi Tokuda, · 2009
Earlier work this paper cites.
“The Blizzard Challenge 2010,”
Alan W Black, Simon King, and Keiichi Tokuda, · 2010
Earlier work this paper cites.
“The Blizzard Challenge 2011,”
Simon King and Vasilis Karaiskos, · 2011
Earlier work this paper cites.
“LibriSpeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Sentiment analysis using image-based deep spectrum features,”
Shahin Amiriparian, Nicholas Cummins, Sandra Ottl, Maurice Gerczuk, and Björn Schuller, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“SGDR: Stochastic gradient descent with warm restarts,”
Ilya Loshchilov and Frank Hutter, · 2017
Earlier work this paper cites.
“VoxCeleb2: Deep Speaker Recognition,”
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, · 2018
Earlier work this paper cites.
“mixup: Beyond empirical risk minimization,”
Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz, · 2018
Cited alongside, same era.
“MOSNet: Deep Learning-Based Objective Assessment for Voice Conversion,”
Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang, · 2019
Cited alongside, same era.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2019
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“EfficientNetV2: Smaller models and faster training,”
Mingxing Tan and Quoc V. Le, · 2021
Cited alongside, same era.
“How do voices from past speech synthesis challenges compare today?,”
Erica Cooper and Junichi Yamagishi, · 2021
“The VoiceMOS Challenge 2022,”
Wen-Chin Huang, Erica Cooper, Yu Tsao, Hsin-Min Wang, Tomoki Toda, and Junichi Yamagishi, · 2022
Later among the works it cites.
“SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis,”
Georgia Maniati, Alexandra Vioni, Nikolaos Ellinas, Karolos Nikitaras, Konstantinos Klapsas, June Sig Sung, Gunu Jho, Aimilios Chalamandaris, and Pirros Tsiakoulis, · 2022
Later among the works it cites.
“Investigating range-equalizing bias in mean opinion score ratings of synthesized speech,”
Erica Cooper and Junichi Yamagishi, · 2023
Later among the works it cites.
“The VoiceMOS Challenge 2023: Zero-shot subjective speech quality prediction for multiple domains,”
Erica Cooper, Wen-Chin Huang, Yu Tsao, Hsin-Min Wang, Tomoki Toda, and Junichi Yamagishi, · 2023
Later among the works it cites.
“LE-SSL-MOS: Self-supervised learning MOS prediction with listener enhancement,”
Zili Qi, Xinhui Hu, Wangjin Zhou, Sheng Li, Hao Wu, Jian Lu, and Xinkang Xu, · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“EfficientNetV2: Smaller models and faster training,”
Mingxing Tan and Quoc V. Le, · 2021
Cited alongside, same era.
“Better aggregation in test-time augmentation,”
Divya Shanmugam, Davis W. Blalock, Guha Balakrishnan, and John V. Guttag, · 2021
Cited alongside, same era.
“Generalization ability of mos prediction networks,”
Erica Cooper, Wen-Chin Huang, Tomoki Toda, and Junichi Yamagishi, · 2022
Cited alongside, same era.
“UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022,”
Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari, · 2022
Cited alongside, same era.
“MOSPC: MOS prediction based on pairwise comparison,”
Kexin Wang, Yunlong Zhao, Qianqian Dong, Tom Ko, and Mingxuan Wang, · 2023
Later among the works it cites.
“The Interspeech 2024 Challenge on Speech Processing Using Discrete Units,”
Xuankai Chang, Jiatong Shi, Jinchuan Tian, Yuning Wu, Yuxun Tang, Yihan Wu, Shinji Watanabe, Yossi Adi, Xie Chen, and Qin Jin, · 2024
Closest in time.
“NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,”
Zeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan, Detai Xin, Dongchao Yang, Yanqing Liu, Yichong Leng, Kaitao Song, Siliang Tang, Zhizheng Wu, Tao Qin, Xiang-Yang Li, Wei Ye, Shikun Zhang, Jiang Bian, Lei He, Jinyu Li, and Sheng Zhao, · 2024
Closest in time.
“The VoiceMOS Challenge 2024: Beyond speech quality prediction,”
Wen-Chin Huang, Szu-Wei Fu, Erica Cooper, Ryandhimas E. Zezario, Tomoki Toda, Hsin-Min Wang, Junichi Yamagishi, and Yu Tsao, · 2024
Closest in time.
“Paralinguistic classification of mask wearing by image classifiers and fusion,”
Jeno Szep and Salim Hariri, · 2091
Closest in time.