Fetching the paper…
Reading the bibliography…
Human judgments obtained through Mean Opinion Scores (MOS) are the most reliable way to assess the quality of speech signals.
R. Caruana, “Multitask learning,” Machine learning , vol. 28, no. 1, pp. 41–75, 1997
1997
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier et al. , “Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,” in 2001 IEEE ICASSP , vol. 2. IEEE, 2001, pp. 749–752
2001
Earlier work this paper cites.
S. Coren, L. M. Ward, and J. T. Enns, Sensation and perception . John Wiley & Sons Hoboken, NJ, 2004
2004
Earlier work this paper cites.
J. Breebaart, J. Engdegård, C. Falch et al. , “Spatial audio object coding (SAOC)-the upcoming mpeg standard on parametric object based audio coding,” in AES Convention 124 . AES, 2008
2008
Earlier work this paper cites.
T. Manjunath, “Limitations of perceptual evaluation of speech quality on voip systems,” in 2009 IEEE International Symposium on Broadband Multimedia Systems and Broadcasting . IEEE, 2009, pp. 1–6
2009
Earlier work this paper cites.
T. Kastner, “Evaluating physical measures for predicting the perceived quality of blindly separated audio source signals,” in AES Convention , 2009
2009
Earlier work this paper cites.
H. Abdi and L. J. Williams, “Principal component analysis,” Wiley interdisciplinary reviews: computational statistics , vol. 2, no. 4, pp. 433–459, 2010
2010
Earlier work this paper cites.
V. Emiya, E. Vincent, N. Harlander et al. , “Subjective and objective quality assessment of audio source separation,” IEEE TASLP , 2011
2011
Earlier work this paper cites.
I. Koch, V. Lawo, J. Fels et al. , “Switching in the cocktail party: exploring intentional control of auditory selective attention.” Journal of Experimental Psychology: HPP , vol. 37, no. 4, p. 1140, 2011
2011
Earlier work this paper cites.
J. G. Beerends, C. Schmidmer, J. Berger et al. , “Perceptual objective listening quality assessment (POLQA), the third generation itu-t standard for end-to-end speech quality measurement part i—temporal alignment,” Journal of the AES , vol. 61, no. 6, pp. 366–384, 2013
2013
Earlier work this paper cites.
A. Hines, J. Skoglund, A. Kokaram et al. , “Robustness of speech quality metrics to background noise and network degradations: Comparing ViSQOL, PESQ and POLQA,” in IEEE ICASSP , 2013
2013
Earlier work this paper cites.
G. J. Mysore, “Can we automatically transform speech recorded on common consumer devices in real-world environments into professional production quality speech?—a dataset, insights, and challenges,” IEEE SPS , vol. 22, no. 8, 2014
2014
Earlier work this paper cites.
A. Hines, J. Skoglund, A. C. Kokaram et al. , “ViSQOL: an objective speech quality model,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2015, no. 1, pp. 1–18, 2015
2015
Earlier work this paper cites.
N. Harte, E. Gillen, and A. Hines, “TCD-VoIP, a research database of degraded speech for assessing quality in voip applications,” in QoMEX , 2015
2015
Earlier work this paper cites.
R. C. Streijl, S. Winkler, and D. S. Hands, “Mean opinion score (MOS) revisited: methods and applications, limitations and alternatives,” Multimedia Systems , vol. 22, no. 2, pp. 213–227, 2016
2016
Earlier work this paper cites.
E. Cano, D. FitzGerald, and K. Brandenburg, “Evaluation of quality of sound source separation algorithms: Human perception vs quantitative metrics,” in 2016 24th European Signal Processing Conference (EUSIPCO) . IEEE, 2016, pp. 1758–1762
2016
Earlier work this paper cites.
B. Patton, Y. Agiomyrgiannakis, M. Terry et al. , “AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech,” arXiv , 2016
2016
Cited alongside, same era.
Z. Jin, G. J. Mysore, S. Diverdi et al. , “VoCo: Text-based insertion and replacement in audio narration,” ACM TOG , 2017
2017
Cited alongside, same era.
S.-W. Fu, Y. Tsao, H.-T. Hwang et al. , “Quality-net: end-to-end non-intrusive speech quality assessment model on blstm,” Interspeech , 2018
2018
Cited alongside, same era.
Z. Jin, A. Finkelstein, G. J. Mysore et al. , “FFTNet: A real-time speaker-dependent neural vocoder,” in ICASSP , 2018
2018
Cited alongside, same era.
2018
A. A. Catellier and S. D. Voran, “WaweNets: A no-reference convolutional waveform-based approach to estimating narrowband and wideband speech quality,” in ICASSP , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Su, Z. Jin, and A. Finkelstein, “HiFi-GAN: High-fidelity denoising and dereverberation based on speech deep features in adversarial networks,” Interspeech , 2020
2020
Later among the works it cites.
G. Mittag, B. Naderi, A. Chehadi et al. , “NISQA: A deep CNN-self-attention model for multidimensional speech quality prediction with crowdsourced datasets,” in Interspeech , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
C.-C. Lo, S.-W. Fu, W.-C. Huang et al. , “MOSNet: Deep learning based objective assessment for voice conversion,” Interspeech , 2019
2019
Cited alongside, same era.
S.-W. Fu, C.-F. Liao, Y. Tsao et al. , “MetricGAN: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” in ICML , 2019
2019
Cited alongside, same era.
X. Dong and D. S. Williamson, “A classification-aided framework for non-intrusive speech quality assessment,” in WASPAA , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
C. Shorten and T. M. Khoshgoftaar, “A survey on image data augmentation for deep learning,” Journal of big data , vol. 6, no. 1, pp. 1–48, 2019
2019
Cited alongside, same era.
B. Feng, Z. Jin, J. Su et al. , “Learning bandwidth expansion using perceptually-motivated loss,” in ICASSP , 2019, pp. 606–610
2019
Cited alongside, same era.
J. Su, A. Finkelstein, and Z. Jin, “Perceptually-motivated environment-specific speech enhancement,” in ICASSP , 2019
2019
Cited alongside, same era.
2021
Later among the works it cites.
P. Manocha, Z. Jin, R. Zhang et al. , “CDPAM: Contrastive learning for perceptual audio similarity,” ICASSP , 2021
2021
Later among the works it cites.
M. Yu, C. Zhang, Y. Xu et al. , “MetricNet: Improved modeling for non-intrusive speech quality assessment,” Interspeech , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
Z. Zhang, P. Vyas, X. Dong et al. , “An end-to-end non-intrusive model for subjective and objective real-world speech assessment using a multi-task framework,” in ICASSP , 2021, pp. 316–320
2021
Later among the works it cites.
P. Manocha, B. Xu, and A. Kumar, “NORESQA: A framework for speech quality assessment using non-matching references,” NeurIPS , vol. 34, 2021
2021
Later among the works it cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai et al. , “Hubert: Self-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM TASLP , vol. 29, pp. 3451–3460, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Su, Z. Jin, and A. Finkelstein, “HiFi-GAN-2: Studio-quality speech enhancement via generative adversarial networks conditioned on acoustic features,” in WASPAA 2021 , Oct. 2021
2021
Later among the works it cites.
2022
Closest in time.
P. Manocha, Z. Jin, and A. Finkelstein, “SQAPP: No-reference speech quality assessment via pairwise preference,” in ICASSP , May 2022
2022
Closest in time.