Fetching the paper…
Reading the bibliography…
Several studies have proposed deep-learning-based models to predict the mean opinion score (MOS) of synthesized speech, showing the possibility of replacing human raters.
“The proof and measurement of association between two things,”
C. Spearman, · 1904
Earlier work this paper cites.
“Notes on the history of correlation,”
K. Pearson, · 1920
Earlier work this paper cites.
“Mel-cepstral distance measure for objective speech quality assessment,”
R. Kubichek, · 1993
Earlier work this paper cites.
“Multitask learning: A knowledge-based source of inductive bias,”
R. A. Caruana, · 1993
Earlier work this paper cites.
An Introduction to the Bootstrap
B. Efron and R. J. Tibshirani, · 1994
Earlier work this paper cites.
“An objective measure for estimating MOS of synthesized speech,”
M. Chu and H. Peng, · 2001
Earlier work this paper cites.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,”
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, · 2001
Earlier work this paper cites.
“ANIQUE: an auditory model for single-ended speech quality estimation,”
D.-S. Kim, · 2005
Earlier work this paper cites.
“P. 563: The ITU-T standard for single-ended speech quality assessment,”
L. Malfait, J. Berger, and M. Kastner, · 2006
Earlier work this paper cites.
“Detection of synthetic speech for the problem of imposture,”
P. L. De Leon, I. Hernaez, I. Saratxaga, M. Pucher, and J. Yamagishi, · 2011
Earlier work this paper cites.
“WaveNet: A generative model for raw audio,”
A. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, · 2016
Earlier work this paper cites.
“A hierarchical predictor of synthetic speech naturalness using neural networks,”
T. Yoshimura, G. E. Henter, O. Watts, M. Wester, J. Yamagishi, and K. Tokuda, · 2016
Cited alongside, same era.
“AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech,”
B. Patton, Y. Agiomyrgiannakis, M. Terry, K. W. Wilson, R. A. Saurous, and D. Sculley, · 2016
Cited alongside, same era.
“The Voice Conversion Challenge 2016,”
T. Toda, L. Chen, D. Saito, F. Villavicencio, M. Wester, Z. Wu, and J. Yamagishi, · 2016
Cited alongside, same era.
“An overview of multi-task learning in deep neural networks,”
S. Ruder, · 2017
Cited alongside, same era.
“Multitask learning with low-level auxiliary tasks for encoder-decoder based speech recognition,”
S. Toshniwal, H. Tang, L. Lu, and K. Livescu, · 2017
Cited alongside, same era.
“Neural TTS voice conversion,”
Z. Kons, S. Shechtman, A. Sorin, R. Hoory, C. Rabinovitz, and E. D. S. Morais, · 2018
Later among the works it cites.
“Learning general purpose distributed sentence representations via large scale multi-task learning,”
S. Subramanian, A. Trischler, Y. Bengio, and C. J. Pal, · 2018
Later among the works it cites.
“Exploring named entity recognition as an auxiliary task for slot filling in conversational language understanding,”
S. Louvan and B. Magnini, · 2018
Later among the works it cites.
“The voice conversion challenge 2018: Promoting development of parallel and nonparallel methods,”
J. Lorenzo-Trueba, J. Yamagishi, T. Toda, D. Saito, F. Villavicencio, T. Kinnunen, and Z. Ling, · 2018
Later among the works it cites.
“Neural speech synthesis with Transformer network,”
N. Li, S. Liu, Y. Liu, S. Zhao, and M. Liu, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Bingel and A. Søggard, · 2017
Cited alongside, same era.
“Focal loss for dense object detection,”
T. Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, · 2017
Cited alongside, same era.
“Deep TEN: Texture encoding network,”
H. Zhang, J. Xue, and K. Dana, · 2017
Cited alongside, same era.
“Natural TTS synthesis by conditioning WaveNet on Mel spectrogram predictions,”
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerry-Ryan, R. A. Saurous, Y. Agiomyrgiannakis, and Y. Wu, · 2018
Cited alongside, same era.
“Deep Voice 3: Scaling text-to-speech with convolutional sequence learning,”
W. Ping, K. Peng, A. Gibiansky, S. O. Arik, A. Kannan, S. Narang, J. Raiman, and J. Miller, · 2018
Cited alongside, same era.
“CycleGAN-VC: Non-parallel voice conversion using cycle-consistent adversarial networks,”
T. Kaneko and H. Kameoka, · 2018
Cited alongside, same era.
“Speech synthesis evaluation — state-of-the-art assessment and suggestion for a novel research program,”
P. Wagner, J. Beskow, S. Betz, J. Edlund, J. Gustafson, G. E. Henter, S. L. Maguer, Z. Malisz, E. Székely, C. Tånnander, and J. Voße, · 2019
Later among the works it cites.
“MOSNet: Deep learning based objective assessment for voice conversion,”
C. Lo, S. Fu, W. Huang, X. Wang, J. Yamagishi, Y. Tsao, and H. Wang, · 2019
Later among the works it cites.
“Simultaneous detection and localization of a wake-up word using multi-task learning of the duration and endpoint,”
T. Maekaku, Y. Kida, and A. Sugiyama, · 2019
Later among the works it cites.
“Multi-task learning with high-order statistics for x-vector based text-independent speaker verification,”
L. You, W. Guo, L. Dai, and J. Du, · 2019
Later among the works it cites.
“Multi-task training of hybrid DNN-TVM model for speaker verification with noisy and far-field speech,”
A. Jati, R. Peri, M. Pal, T. J. Park, N. Kumar, R. Travadi, P. Georgiou, and S. Narayanan, · 2019
Later among the works it cites.
“Deep MOS predictor for synthetic speech using cluster-based modeling,”
Y. Choi, Y. Jung, and H. Kim, · 2020
Closest in time.