Fetching the paper…
Reading the bibliography…
Existing objective evaluation metrics for voice conversion (VC) are not always correlated with human perception.
C. Spearman, “The proof and measurement of association between two things,”
1904
Earlier work this paper cites.
K. Pearson, “Notes on The History of Correlation,”
1920
Earlier work this paper cites.
R. Kubichek, “Mel-cepstral distance measure for objective speech quality assessment,” in
1993
Earlier work this paper cites.
B. Efron and R. J. Tibshirani,
1993
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,” in
2001
Earlier work this paper cites.
M. Bisani and H. Ney, “Bootstrap estimates for confidence intervals in ASR performance evaluation,” in
2004
Earlier work this paper cites.
M. Cernak and M. Rusko, “An evaluation of synthetic speech using the PESQ measure,” in
2005
Earlier work this paper cites.
T. H. Falk and W. Y. Chan, “Single-ended speech quality measurement using machine learning methods,”
2006
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted Boltzmann machines,” in
2010
Earlier work this paper cites.
D.-Y. Huang, “Prediction of perceived sound quality of synthetic speech,” in
2011
Cited alongside, same era.
U. Remes, R. Karhila, and M. Kurimo, “Objective evaluation measures for speaker-adaptive HMM-TTS systems,” in
2013
Cited alongside, same era.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,”
2014
Cited alongside, same era.
G. J. Mysore, “Can we automatically transform speech recorded on common consumer devices in real-world environments into professional production quality speech? — a dataset, insights, and challenges,”
2015
Cited alongside, same era.
T. Sainath, O. Vinyals, A. Senior, and H. Sak, “Convolutional, long short-term memory, fully connected deep neural networks,” in
2015
Cited alongside, same era.
B. Patton, Y. Agiomyrgiannakis, M. Terry, K. W. Wilson, R. A. Saurous, and D. Sculley, “AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech,” in
2016
Later among the works it cites.
S.-W. Fu, Y. Tsao, H.-T. Hwang, and H.-M. Wang, “Quality-Net: An end-to-end non-intrusive speech quality assessment model based on BLSTM,” in
2018
Later among the works it cites.
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, and Z. Ling, “The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods,” in
2018
Later among the works it cites.
H. Zhao, S. Zarar, I. Tashev, and C.-H. Lee, “Convolutional-recurrent neural networks for speech enhancement,” in
2018
Later among the works it cites.
K. Tan and D. Wang, “A convolutional recurrent neural network for real-time speech enhancement,” in
2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. N. Sainath, R. J. Weiss, A. W. Senior, K. W. Wilson, and O. Vinyals, “Learning the speech front-end with raw waveform CLDNNs,” in
2015
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,”
2015
Cited alongside, same era.
T. Yoshimura, G. Eje Henter, O. Watts, M. Wester, J. Yamagishi, and K. Tokuda, “A hierarchical predictor of synthetic speech naturalness using neural networks,” in
2016
Cited alongside, same era.
Later among the works it cites.
X. Lu, P. Shen, S. Li, Y. Tsao, and H. Kawai, “Temporal attentive pooling for acoustic event detection,” in
2018
Later among the works it cites.
D. Wang, S. Lv, X. Wang, and X. Lin, “Gated convolutional LSTM for speech commands recognition,” in
2018
Later among the works it cites.
A. R. Avila, H. Gamper, C. Reddy, R. Cutler, I. Tashev, and J. Gehrke, “Non-intrusive speech quality assessment using neural networks,” in
2019
Closest in time.