Fetching the paper…
Reading the bibliography…
In this paper, we present a new objective prediction model for synthetic speech naturalness.
A. Mariniak, “A global framework for the assessment of synthetic speech without subjects,” in
1993
Earlier work this paper cites.
C.-C. Lo, S.-W. Fu, W.-C. Huang, X. Wang, J. Yamagishi, Y. Tsao, and H.-M. Wang, “MOSNet: Deep Learning-Based Objective Assessment for Voice Conversion,” in
2003
Earlier work this paper cites.
K. Seget, “Untersuchungen zur auditiven qualität von sprachsyntheseverfahren,”
2007
Earlier work this paper cites.
T. H. Falk and S. Moller, “Towards signal-based instrumental quality diagnosis for text-to-speech systems,”
2008
Earlier work this paper cites.
T. H. Falk, S. Möller, V. Karaiskos, and S. King, “Improving instrumental quality prediction performance for the blizzard challenge,” in
2008
Earlier work this paper cites.
V. Karaiskos, S. King, R. A. J. Clark, and C. Mayo, “The blizzard challenge 2008,” in
2008
Earlier work this paper cites.
S. King and V. Karaiskos, “The blizzard challenge 2009,” in
2009
Earlier work this paper cites.
S. Möller, F. Hinterleitner, T. H. Falk, and T. Polzehl, “Comparison of approaches for instrumentally predicting the quality of text-to-speech systems,” in
2010
Earlier work this paper cites.
F. Hinterleitner, S. Möller, T. H. Falk, and T. Polzehl, “Comparison of approaches for instrumentally predicting the quality of text-to-speech systems: Data from blizzard challenges 2008 and 2009,” in
2010
Earlier work this paper cites.
——, “The blizzard challenge 2010,” in
2010
Earlier work this paper cites.
——, “The blizzard challenge 2011,” in
2011
Earlier work this paper cites.
F. Hinterleitner, S. Möller, C. Norrenbrock, and U. Heute, “Perceptual quality dimensions of text-to-speech systems,” in
2011
Earlier work this paper cites.
——, “The blizzard challenge 2012,” in
2012
Earlier work this paper cites.
F. Hinterleitner, C. Norrenbrock, S. Möller, and U. Heute, “What makes this voice sound so bad? a multidimensional analysis of state-of-the-art text-to-speech systems,” in
2012
Cited alongside, same era.
F. Hinterleitner, C. Norrenbrock, S. Möller, and U. Heute, “Predicting the quality of text-to-speech systems from a large-scale feature set.” in
2013
Cited alongside, same era.
——, “The blizzard challenge 2013,” in
2013
Cited alongside, same era.
S. King, “Measuring a decade of progress in text-to-speech,”
2014
Cited alongside, same era.
D. Estival, S. Cassidy, F. Cox, and D. Burnham, “Austalk: an audio-visual corpus of australian english,” in
2014
Cited alongside, same era.
B. Patton, Y. Agiomyrgiannakis, M. Terry, K. Wilson, R. A. Saurous, and D. Sculley, “Automos: Learning a non-intrusive assessor of naturalness-of-speech,” in
2016
Later among the works it cites.
S. King and V. Karaiskos, “The blizzard challenge 2016,” in
2016
Later among the works it cites.
T. Toda, L.-H. Chen, D. Saito, F. Villavicencio, M. Wester, Z. Wu, and J. Yamagishi, “The voice conversion challenge 2016,” in
2016
Later among the works it cites.
M. Wester, Z. Wu, and J. Yamagishi, “Analysis of the voice conversion challenge 2016 evaluation results,” in
2016
Later among the works it cites.
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2014
Cited alongside, same era.
C. R. Norrenbrock, F. Hinterleitner, U. Heute, and S. Möller, “Quality prediction of synthesized speech based on perceptual quality dimensions,”
2015
Cited alongside, same era.
L. Latacz and W. Verhelst, “Double-ended prediction of the naturalness ratings of the blizzard challenge 2008-2013,” in
2015
Cited alongside, same era.
K. Prahallad, A. Vadapalli, S. K. Rallabandi, S. Kesiraju, H. Murthy, T. Nagarajan, B. Singh, S. T, K. S. Rao, S. V. Gangashetty, K. Simon, K. Tokuda, and A. W. Black, “The blizzard challenge 2015,” in
2015
Cited alongside, same era.
R. Gupta, H. J. Banville, and T. H. Falk, “PhySyQX: A database for physiological evaluation of synthesised speech quality-of-experience,” in
2015
Cited alongside, same era.
M. H. Soni and H. A. Patil, “Non-intrusive quality assessment of synthesized speech using spectral features and support vector regression.” in
2016
Cited alongside, same era.
T. Yoshimura, G. E. Henter, O. Watts, M. Wester, J. Yamagishi, and K. Tokuda, “A hierarchical predictor of synthetic speech naturalness using neural networks.” in
2016
Cited alongside, same era.
L. Fernández Gallardo and B. Weiss, “The nautilus speaker characterization corpus: Speech recordings and labels of speaker characteristics and voice descriptions,” in
2018
Later among the works it cites.
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, and Z. Ling, “The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods,” in
2018
Later among the works it cites.
G. Mittag and S. Möller, “Non-intrusive speech quality assessment for super-wideband speech communication networks,” in
2019
Later among the works it cites.
M. Tang and J. Zhu, “Text-to-speech quality evaluation based on lstm recurrent neural networks,” in
2019
Later among the works it cites.
Z. Wu, Z. Xie, and S. King, “The blizzard challenge 2019,” in
2019
Later among the works it cites.
Y. Guo and J. Zhu, “Naturalness evaluation of synthetic speech based on residual learning networks,” in
2020
Later among the works it cites.
G. Mittag and S. Möller, “Full-reference speech quality estimation with attentional siamese neural networks,” in
2020
Later among the works it cites.