Fetching the paper…
Reading the bibliography…
Automatic speech quality assessment is an important, transversal task whose progress is hampered by the scarcity of human annotations, poor generalization to unseen recording conditions, and a lack of flexibility of existing approaches.
“Perceptual evaluation of speech quality (PESQ) – A new method for speech quality assessment of telephone networks and codecs,”
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, · 2001
Earlier work this paper cites.
“ANIQUE: an auditory model for single-ended speech quality estimation,”
D.-S. Kim, · 2005
Earlier work this paper cites.
“Learning to rank using gradient descent,”
C. Burges, T. Shaked, E. Renshaw, A. Lazier, M. Deeds, N. Hamilton, and G. Hullender, · 2005
Earlier work this paper cites.
“Low-complexity, nonintrusive speech quality assessment,”
V. Grancharov, D. Y. Zhao, J. Lindblom, and W. B. Kleijn, · 2006
Earlier work this paper cites.
Semi-Supervised Learning
O. Chapelle, B. Schölkopf, and A. Zien, · 2006
Earlier work this paper cites.
A guide to NumPy
T. E. Oliphant, · 2006
Earlier work this paper cites.
“Matplotlib: A 2D graphics environment,”
J. D. Hunter, · 2007
Earlier work this paper cites.
“Perception of speech and sound,”
B. Kollmeier, T. Brand, and B. Meyer, · 2008
Earlier work this paper cites.
“Limitations of perceptual evaluation of speech quality on VoIP systems,”
T. Manjunath, · 2009
Earlier work this paper cites.
“A non-intrusive quality and intelligibility measure of reverberant and dereverberated speech,”
T. H. Falk, C. Zheng, and W.-Y. Chan, · 2010
Earlier work this paper cites.
“P.563 – The ITU-T standard for single-ended speech quality assessment,”
L. Malfait, J. Berger, and M. Kastner, · 2010
Earlier work this paper cites.
“Speech quality assessment,”
P. C. Loizou, · 2011
Earlier work this paper cites.
“An algorithm for intelligibility prediction of time-frequency weighted noisy speech,”
C. H. Taal, R. C. Hendricks, R. Heusdens, and J. Jensen, · 2011
Earlier work this paper cites.
“The NumPy array: a structure for efficient numerical computation,”
S. Van Der Walt, S. C. Colbert, and G. Varoquaux, · 2011
Earlier work this paper cites.
“Perceptual objective listening quality assessment (POLQA), the third generation ITU-T standard for end-to-end speech quality measurement.,”
J. G. Beerends, C. Schmidmer, J. Berger, M Oberman, R. Ullman, J. Pomy, and M. Keyhl, · 2013
Cited alongside, same era.
“TCD-VoIP, a research database of degraded speech for assessing quality in VoIP applications,”
N. Harte, E. Gillen, and A. Hines, · 2015
Cited alongside, same era.
“ESC: dataset for environmental sound classification,”
K. J. Piczak, · 2015
Cited alongside, same era.
“librosa: audio and music signal analysis in python,”
B. McFee, C. A. Raffel, D. Liang, D. P. W. E. Ellis, M. McVicar, E. Battenberg, and O. Nieto, · 2015
Cited alongside, same era.
“Novel deep autoencoder features for non-intrusive speech quality assessment,”
M. H. Soni and H. A. Patil, · 2016
Cited alongside, same era.
“WEnets: a convolutional framework for evaluating audio waveforms,”
A. A. Catellier and S. D. Voran, · 2019
Later among the works it cites.
“MOSNet: deep learning-based objective assessment for voice conversion,”
C.-C. Lo, S.-W. Fu, W.-C. Huang, X. Wang, J. Yamagishi, Y. Tsao, and H.-M. Wang, · 2019
Later among the works it cites.
“Exploiting unlabeled data in CNNs by self-supervised learning to rank,”
X. Liu, J. van de Weijer, and A. D. Bagdanov, · 2019
Later among the works it cites.
“Learning problem-agnostic speech representations from multiple self-supervised tasks,”
S. Pascual, M. Ravanelli, J. Serrà, A. Bonafonte, and Y. Bengio, · 2019
Later among the works it cites.
“SDR - Half-baked or well done?,”
J. L. Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Patton, Y. Agiomyrgiannakis, M. Terry, K. Wilson, R. A. Saurous, and D. Sculley, · 2016
Cited alongside, same era.
“dipIQ: blind image quality assessment by learning-to-rank discriminable image pairs,”
K. Ma, W. Liu, T. Liu, Z. Wang, and D. Tao, · 2017
Cited alongside, same era.
“Quality-Net: an end-to-end non-intrusive speech quality assessment model based on BLSTM,”
S.-W. Fu, Y. Tsao, H.-T. Hwang, and H.-M. Wang, · 2018
Cited alongside, same era.
“Towards a universal neural network encoder for time series,”
J. Serrà, S. Pascual, and A. Karatzoglou, · 2018
Cited alongside, same era.
“Averaging weights leads to wider optima and better generalization,”
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson, · 2018
Cited alongside, same era.
“The Ryerson audio-visual database of emotional speech and song (RAVDESS),”
S. R. Livingstone and F. A. Russo, · 2018
Cited alongside, same era.
“Non-intrusive speech quality assessment for super-wideband speech communication networks,”
G. Mittag and S. Möller, · 2019
Cited alongside, same era.
R. Zhang, · 2019
Later among the works it cites.
“New deep learning optimizer, Ranger: synergistic combination of RAdam + Lookahead for the best of both,” Medium: https://medium.com/@lessw/new-deep-learning-optimizer-ranger-synergistic-combination-of-radam-lookahead-for-the-best-of-2dc83f79a48d , 2019
L. Wright, · 2019
Later among the works it cites.
“CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice cloning toolkit (version 0.92),”
Y. Yamagishi, C. Veaux, and K. MacDonald, · 2019
Later among the works it cites.
“Learning sound event classifiers from web audio with noisy labels,”
E. Fonseca, M. Plakal, D. P. W. E. Ellis, F. Font, X. Favory, and X. Serra, · 2019
Later among the works it cites.
“PyTorch: an imperative style, high-performance deep learning library,”
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, · 2019
Later among the works it cites.
“ViSQOL v3: an open source production ready objective speech and audio metric,”
M. Chinen, F. S. C. Lim, J. Skoglund, N. Gureev, F. O’Gorman, and A. Hines, · 2020
Closest in time.
“A differentiable perceptual audio metric learned from just noticeable differences,”
P. Manocha, A. Finkelstein, Z. Jin, N. J. Bryan, R. Zhang, and G. J. Mysore, · 2020
Closest in time.
“A survey on semi-supervised learning,”
J. E. van Engelen and H. H. Hoos, · 2020
Closest in time.