Fetching the paper…
Reading the bibliography…
In this work, we present the SOMOS dataset, the first large-scale mean opinion scores (MOS) dataset consisting of solely neural text-to-speech (TTS) samples.
K. E. A. Silverman, M. E. Beckman, J. F. Pitrelli, M. Ostendorf, C. W. Wightman, P. J. Price, J. B. Pierrehumbert, and J. Hirschberg, “TOBI: a standard for labeling English prosody,” in
1992
Earlier work this paper cites.
M. Bisani and H. Ney, “Bootstrap estimates for confidence intervals in ASR performance evaluation,” in
2004
Earlier work this paper cites.
A. W. Black and K. Tokuda, “The Blizzard Challenge - 2005: Evaluating Corpus-Based Speech Synthesis on Common Datasets,” in
2005
Earlier work this paper cites.
T. H. Falk, S. Möller, V. Karaiskos, and S. King, “Improving instrumental quality prediction performance for the Blizzard Challenge,” in
2008
Earlier work this paper cites.
R. Snow, B. O’connor, D. Jurafsky, and A. Y. Ng, “Cheap and fast–but is it good? evaluating non-expert annotations for natural language tasks,” in
2008
Earlier work this paper cites.
A. Kittur, E. H. Chi, and B. Suh, “Crowdsourcing user studies with Mechanical Turk,” in
2008
Earlier work this paper cites.
F. Hinterleitner, S. Möller, T. H. Falk, and T. Polzehl, “Comparison of Approaches for Instrumentally Predicting the Quality of Text-to-Speech Systems: Data from Blizzard Challenges 2008 and 2009,” in
2010
Earlier work this paper cites.
C. Callison-Burch and M. Dredze, “Creating speech and language data with amazon’s mechanical turk,” in
2010
Earlier work this paper cites.
S. Novotney and C. Callison-Burch, “Cheap, fast and good enough: Automatic speech recognition with non-expert transcription,” in
2010
Earlier work this paper cites.
M. Marge, S. Banerjee, and A. I. Rudnicky, “Using the Amazon Mechanical Turk for transcription of spoken language,” in
2010
Earlier work this paper cites.
C. R. Norrenbrock, F. Hinterleitner, U. Heute, and S. Möller, “Towards Perceptual Quality Modeling of Synthesized Audiobooks – Blizzard Challenge 2012,” in
2012
Earlier work this paper cites.
K. Crowston, “Amazon Mechanical Turk: A Research Tool for Organizations and Information Systems Scholars,” in
2012
Earlier work this paper cites.
F. Hinterleitner, C. Manolaina, and S. Möller, “Influence of a voice on the quality of synthesized speech,” in
2014
Earlier work this paper cites.
M. Honnibal and M. Johnson, “An Improved Non-monotonic Transition System for Dependency Parsing,” in
2015
Earlier work this paper cites.
T. Yoshimura, G. E. Henter, O. Watts, M. Wester, J. Yamagishi, and K. Tokuda, “A Hierarchical Predictor of Synthetic Speech Naturalness Using Neural Networks.” in
2016
Cited alongside, same era.
B. Patton, Y. Agiomyrgiannakis, M. Terry, K. Wilson, R. A. Saurous, and D. Sculley, “AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech,” in
2016
Cited alongside, same era.
T. Hoßfeld, P. E. Heegaard, M. Varela, and S. Möller, “QoE beyond the MOS: an in-depth look at QoE via better metrics and their relation to MOS,”
2016
Cited alongside, same era.
K. Ito and L. Johnson, “The LJ Speech Dataset,” https://keithito.com/LJ-Speech-Dataset/, 2017
2017
Cited alongside, same era.
Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengio
2017
Cited alongside, same era.
G. Mittag and S. Möller, “Deep Learning Based Assessment of Synthetic Speech Naturalness,” in
2020
Later among the works it cites.
J. Williams, J. Rownicka, P. Oplustil, and S. King, “Comparison of speech representations for automatic quality estimation in multi-speaker text-to-speech synthesis,” in
2020
Later among the works it cites.
R. Vipperla, S. Park, K. Choo, S. Ishtiaq, K. Min, S. Bhattacharya, A. Mehrotra, A. G. C. P. Ramos, and N. D. Lane, “Bunched lpcnet: Vocoder for low-cost neural text-to-speech systems,” in
2020
Later among the works it cites.
N. Ellinas, G. Vamvoukakis, K. Markopoulos, A. Chalamandaris, G. Maniati, P. Kakoulidis, S. Raptis, J. S. Sung, H. Park, and P. Tsiakoulis, “High quality streaming speech synthesis with low, sentence-length-independent latency,” in
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S.-W. Fu, Y. Tsao, H.-T. Hwang, and H.-m. Wang, “Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model Based on BLSTM,” in
2018
Cited alongside, same era.
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, and Z. Ling, “The Voice Conversion Challenge 2018: Promoting development of parallel and nonparallel methods,” in
2018
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan
2018
Cited alongside, same era.
C.-C. Lo, S.-W. Fu, W.-C. Huang, X. Wang, J. Yamagishi, Y. Tsao, and H.-M. Wang, “MOSNet: Deep Learning-Based Objective Assessment for Voice Conversion,” in
2019
Cited alongside, same era.
M. Todisco, X. Wang, V. Vestman, M. Sahidullah, H. Delgado, A. Nautsch, J. Yamagishi, N. Evans, T. Kinnunen, and K. A. Lee, “ASVspoof 2019: Future horizons in spoofed and fake audio detection,” in
2019
Cited alongside, same era.
J.-M. Valin and J. Skoglund, “LPCNet: Improving neural speech synthesis through linear prediction,” in
2019
Cited alongside, same era.
M. Domínguez, P. L. Rohrer, and J. Soler-Company, “PyToBI: A Toolkit for ToBI Labeling Under Python,” in
2019
Cited alongside, same era.
2020
Later among the works it cites.
T. Raitio, R. Rasipuram, and D. Castellani, “Controllable neural text-to-speech synthesis using intuitive prosodic features,” in
2020
Later among the works it cites.
Z.-H. Ling, X. Zhou, and S. King, “The Blizzard Challenge 2021,” in
2021
Later among the works it cites.
Y. Leng, X. Tan, S. Zhao, F. Soong, X.-Y. Li, and T. Qin, “MBNet: MOS Prediction for Synthesized Speech with Mean-Bias Network,” in
2021
Later among the works it cites.
E. Cooper and J. Yamagishi, “How do voices from past speech synthesis challenges compare today?” in
2021
Later among the works it cites.
Y. Zou, S. Liu, X. Yin, H. Lin, C. Wang, H. Zhang, and Z. Ma, “Fine-grained prosody modeling in neural speech synthesis using ToBI representation,” in
2021
Later among the works it cites.
K. Klapsas, N. Ellinas, J. S. Sung, H. Park, and S. Raptis, “Word-Level Style Control for Expressive, Non-attentive Speech Synthesis,” in
2021
Later among the works it cites.
2022
Closest in time.
W.-C. Huang, E. Cooper, J. Yamagishi, and T. Toda, “LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech,” in
2022
Closest in time.
E. Cooper, W.-C. Huang, T. Toda, and J. Yamagishi, “Generalization ability of MOS prediction networks,” in
2022
Closest in time.