Fetching the paper…
Reading the bibliography…
Human subjective evaluation is the gold standard to evaluate speech quality optimized for human perception.
Feb 1998
“ITU-T Recommendation P.800: Methods for subjective determination of transmission quality,” · 1998
Earlier work this paper cites.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,”
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, · 2001
Earlier work this paper cites.
Subjective test methodology for evaluating speech communication systems that include noise suppression algorithm
ITU-T Recommendation P.835, · 2003
Earlier work this paper cites.
“Evaluation of objective measures for speech enhancement,”
Yi Hu and Philipos C Loizou, · 2006
Earlier work this paper cites.
“Influence of loudness level on the overall quality of transmitted speech,”
Côté Nicolas, Valérie Gautier-Turbin, and Sebastian Möller, · 2007
Earlier work this paper cites.
“3QUEST: 3-fold Quality Evaluation of Speech in Telecommunications Systems,” 2008
Head Acoustics Application Note, · 2008
Earlier work this paper cites.
“Perceptual Objective Listening Quality Assessment (POLQA), The Third Generation ITU-T Standard for End-to-End Speech Quality Measurement Part II-Perceptual Model,”
John Beerends, Christian Schmidmer, Jens Berger, Matthias Obermann, Raphael Ullmann, Joachim Pomy, and Michael Keyhl, · 2013
Earlier work this paper cites.
“ViSQOL: an objective speech quality model,”
Andrew Hines, Jan Skoglund, Anil C Kokaram, and Naomi Harte, · 2015
Earlier work this paper cites.
“Speech acoustic modeling from raw multichannel waveforms,”
Y. Hoshen, R. J. Weiss, and K. W. Wilson, · 2015
Earlier work this paper cites.
Subjective evaluation of speech quality with a crowdsourcing approach
ITU-T Recommendation P.808, · 2018
Cited alongside, same era.
“Quality-net: An end-to-end non-intrusive speech quality assessment model based on blstm,”
Szu-Wei Fu, Yu Tsao, Hsin-Te Hwang, and Hsin-Min Wang, · 2018
Cited alongside, same era.
“A scalable noisy speech dataset and online subjective test framework,”
Chandan KA Reddy, Ebrahim Beyrami, Jamie Pool, Ross Cutler, Sriram Srinivasan, and Johannes Gehrke, · 2019
Cited alongside, same era.
“Non-intrusive speech quality assessment using neural networks,”
A. R. Avila, H. Gamper, C. Reddy, R. Cutler, I. Tashev, and J. Gehrke, · 2019
Cited alongside, same era.
“Intrusive and non-intrusive perceptual speech quality assessment using a convolutional neural network,”
Hannes Gamper, Chandan KA Reddy, Ross Cutler, Ivan J Tashev, and Johannes Gehrke, · 2019
Cited alongside, same era.
“Wawenets: A no-reference convolutional waveform-based approach to estimating narrowband and wideband speech quality,”
Andrew A Catellier and Stephen D Voran, · 2020
Later among the works it cites.
“A differentiable perceptual audio metric learned from just noticeable differences,”
Pranay Manocha, Adam Finkelstein, Zeyu Jin, Nicholas J Bryan, Richard Zhang, and Gautham J Mysore, · 2020
Later among the works it cites.
“Real time speech enhancement in the waveform domain,”
Alexandre Defossez, Gabriel Synnaeve, and Yossi Adi, · 2020
Later among the works it cites.
“Subjective evaluation of noise suppression algorithms in crowdsourcing,”
Babak Naderi and Ross Cutler, · 2021
Closest in time.
“INTERSPEECH 2021 Deep Noise Suppression Challenge,”
Chandan K A Reddy, Harishchandra Dubey, Kazuhito Koishida, Arun Nair, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, and Sriram Srinivasan, · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Improving deep models of speech quality prediction through voice activity detection and entropy-based measures,”
Jasper Ooster and Bernd T Meyer, · 2019
Cited alongside, same era.
“Non-intrusive speech quality prediction using modulation energies and lstm-network,”
Benjamin Cauchi, Kai Siedenburg, Joao F Santos, Tiago H Falk, Simon Doclo, and Stefan Goetze, · 2019
Cited alongside, same era.
“WaveCycleGAN2: Time-domain neural post-filter for speech waveform generation,”
Kou Tanaka, Hirokazu Kameoka, Takuhiro Kaneko, and Nobukatsu Hojo, · 2019
Cited alongside, same era.
“An attention enhanced multi-task model for objective speech assessment in real-world environments,”
Xuan Dong and Donald S Williamson, · 2020
Cited alongside, same era.
“DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,”
Chandan K A Reddy, Vishak Gopal, and Ross Cutler, · 2021
Closest in time.
“Dccrn+: Channel-wise subband dccrn with snr estimation for speech enhancement,”
Shubo Lv, Yanxin Hu, Shimin Zhang, and Lei Xie, · 2021
Closest in time.
“A simultaneous denoising and dereverberation framework with target decoupling,”
Andong Li, Wenzhe Liu, Xiaoxue Luo, Guochen Yu, Chengshi Zheng, and Xiaodong Li, · 2021
Closest in time.
“Sesqa: semi-supervised learning for speech quality assessment,”
Joan Serrà, Jordi Pons, and Santiago Pascual, · 2021
Closest in time.