Fetching the paper…
Reading the bibliography…
We present the UTokyo-SaruLab mean opinion score (MOS) prediction system submitted to VoiceMOS Challenge 2022.
D. H. Wolpert, “Stacked generalization,” Neural Networks , vol. 5, no. 2, pp. 241–259, 1992
1992
Earlier work this paper cites.
M. Ester, H.-P. Kriegel, J. Sander, and X. Xu, “A density-based algorithm for discovering clusters in large spatial databases with noise,” in Proceedings of the Second International Conference on Knowledge Discovery and Data Mining . AAAI Press, 1996, p. 226–231
1996
Earlier work this paper cites.
L. Breiman, “Stacked regressions,” Machine learning , vol. 24, pp. 49–64, 1996
1996
Earlier work this paper cites.
A. W. Black and K. Tokuda, “The Blizzard Challenge-2005: Evaluating corpus-based speech synthesis on common datasets,” in Proc. INTERSPEECH , Lisbon, Portugal, Sep. 2005
2005
Earlier work this paper cites.
Z.-H. Zhou, Ensemble methods: foundations and algorithms . CRC press, 2012
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in Proc. ICASSP , South Brisbane, Australia, Apr. 2015, pp. 5206–5210
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T.-Y. Liu, “LightGBM: A highly efficient gradient boosting decision tree,” in Proc. NIPS , vol. 30, 2017
2017
Earlier work this paper cites.
C.-C. Lo, S.-W. Fu, W.-C. Huang, X. Wang, J. Yamagishi, Y. Tsao, and H.-M. Wang, “MOSNet: Deep learning-based objective assessment for voice conversion,” Proc. Interspeech , pp. 1541–1545, 2019
2019
Cited alongside, same era.
T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A next-generation hyperparameter optimization framework,” in Proc. KDD , 2019
2019
Cited alongside, same era.
T. Hayashi, R. Yamamoto, K. Inoue, T. Yoshimura, S. Watanabe, T. Toda, K. Takeda, Y. Zhang, and X. Tan, “ESPnet-TTS: Unified, reproducible, and integratable open source end-to-end text-to-speech toolkit,” Proc. ICASSP , pp. 7654–7658, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
J. Serrà, J. Pons, and S. Pascual, “SESQA: semi-supervised learning for speech quality assessment,” in Proc. ICASSP . IEEE, 2021, pp. 381–385
2021
Later among the works it cites.
P. Manocha, Z. Jin, R. Zhang, and A. Finkelstein, “CDPAM: Contrastive learning for perceptual audio similarity,” in Proc. ICASSP . IEEE, 2021, pp. 196–200
2021
Later among the works it cites.
P. Manocha, B. Xu, and A. Kumar, “NORESQA: A framework for speech quality assessment using non-matching references,” Proc. NeurIPS , vol. 34, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in Proc. LREC 2020 , 2020, pp. 4218–4222
2020
Cited alongside, same era.
Y. Leng, X. Tan, S. Zhao, F. K. Soong, X.-Y. Li, and T. Qin, “MBNET: MOS prediction for synthesized speech with mean-bias network,” Proc. ICASSP , pp. 391–395, 2021
2021
Cited alongside, same era.
E. Cooper and J. Yamagishi, “How do voices from past speech synthesis challenges compare today?” in Proc. SSW , 2021, pp. 183–188
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Later among the works it cites.
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, “Unsupervised Cross-Lingual Representation Learning for Speech Recognition,” in Proc. Interspeech 2021 , 2021, pp. 2426–2430
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.