Fetching the paper…
Reading the bibliography…
We present the first edition of the VoiceMOS Challenge, a scientific event that aims to promote the study of automatic prediction of the mean opinion score (MOS) of synthetic speech.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (PESQ) - a new method for speech quality assessment of telephone networks and codecs,” pp. 749–752, 2001
2001
Earlier work this paper cites.
ITUT Recommendation, “Vocabulary for performance and quality of service,”
2006
Earlier work this paper cites.
V. Karaiskos, S. King, R. A. Clark, and C. Mayo, “The Blizzard Challenge 2008,” in
2008
Earlier work this paper cites.
A. W. Black, S. King, and K. Tokuda, “The Blizzard Challenge 2009,” in
2009
Earlier work this paper cites.
S. Möller, F. Hinterleitner, T. H. Falk, and T. Polzehl, “Comparison of Approaches for Instrumentally Predicting the Quality of Text-to-speech Systems,” in
2010
Earlier work this paper cites.
S. King and V. Karaiskos, “The Blizzard Challenge 2010,” 2010
2010
Earlier work this paper cites.
——, “The Blizzard Challenge 2011,” 2011
2011
Earlier work this paper cites.
C. R. Norrenbrock, F. Hinterleitner, U. Heute, and S. Möller, “Towards perceptual quality modeling of synthesized audiobooks-Blizzard Challenge 2012,” in
2012
Earlier work this paper cites.
——, “The Blizzard Challenge 2013,” 2013
2013
Earlier work this paper cites.
——, “Quality Prediction of Synthesized Speech based on Perceptual Quality Dimensions,”
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Earlier work this paper cites.
T. Yoshimura, G. E. Henter, O. Watts, M. Wester, J. Yamagishi, and K. Tokuda, “A Hierarchical Predictor of Synthetic Speech Naturalness Using Neural Networks,” in
2016
Cited alongside, same era.
B. Patton, Y. Agiomyrgiannakis, M. Terry, K. Wilson, R. A. Saurous, and D. Sculley, “AutoMOS: Learning a Non-intrusive Assessor of Naturalness-of-speech,” in
2016
Cited alongside, same era.
——, “The Blizzard Challenge 2016,” 2016
2016
Cited alongside, same era.
T. Toda, L.-H. Chen, D. Saito, F. Villavicencio, M. Wester, Z. Wu, and J. Yamagishi, “The Voice Conversion Challenge 2016,” in
2016
Cited alongside, same era.
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, and Z. Ling, “The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods,” in
2018
Cited alongside, same era.
Y. Zhao, W.-C. Huang, X. Tian, J. Yamagishi, R. K. Das, T. Kinnunen, Z. Ling, and T. Toda, “Voice Conversion Challenge 2020 - Intra-lingual semi-parallel and cross-lingual voice conversion -,” in
2020
Later among the works it cites.
T. Hayashi, R. Yamamoto, K. Inoue, T. Yoshimura, S. Watanabe, T. Toda, K. Takeda, Y. Zhang, and X. Tan, “ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit,” 2020
2020
Later among the works it cites.
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in
2020
Later among the works it cites.
S. Gururangan, A. Marasović, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith, “Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks,” in
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C.-C. Lo, S.-W. Fu, W.-C. Huang, X. Wang, J. Yamagishi, Y. Tsao, and H.-M. Wang, “MOSNet: Deep Learning-Based Objective Assessment for Voice Conversion,” in
2019
Cited alongside, same era.
Z. Wu, Z. Xie, and S. King, “The Blizzard Challenge 2019,” in
2019
Cited alongside, same era.
G. Mittag and S. Möller, “Deep Learning Based Assessment of Synthetic Speech Naturalness,” in
2020
Cited alongside, same era.
Y. Choi, Y. Jung, and H. Kim, “Deep MOS Predictor for Synthetic Speech using Cluster-Based Modeling,” in
2020
Cited alongside, same era.
J. Williams, J. Rownicka, P. Oplustil, and S. King, “Comparison of speech representations for automatic quality estimation in multi-speaker text-to-speech synthesis,”
2020
Cited alongside, same era.
2021
Later among the works it cites.
W.-C. Tseng, C.-Y. Huang, W.-T. Kao, Y. Y. Lin, and H.-Y. Lee, “Utilizing Self-supervised Representations for MOS Prediction,” in
2021
Later among the works it cites.
E. Cooper and J. Yamagishi, “How do voices from past speech synthesis challenges compare today?” in
2021
Later among the works it cites.
2021
Later among the works it cites.
E. Cooper, W.-C. Huang, T. Toda, and J. Yamagishi, “Generalization ability of MOS prediction networks,” in
2022
Closest in time.
W.-C. Huang, E. Cooper, J. Yamagishi, and T. Toda, “LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech,” in
2022
Closest in time.