Fetching the paper…
Reading the bibliography…
To obtain improved speech enhancement models, researchers often focus on increasing performance according to specific instrumental metrics.
C. A. E. Goodhart, Problems of Monetary Management: The UK Experience . London: Macmillan Education UK, 1984, pp. 91–121
1984
Earlier work this paper cites.
Y. Ephraim and D. Malah, “Speech enhancement using a minimum mean-square error log-spectral amplitude estimator,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 33, no. 2, pp. 443–445, 1985
1985
Earlier work this paper cites.
M. Strathern, “Improving ratings: audit in the british university system,” European Review , vol. 5, no. 3, p. 305–321, 1997
1997
Earlier work this paper cites.
J. H. L. Hansen and B. L. Pellom, “An effective quality evaluation protocol for speech enhancement algorithms,” in Proc. 5th International Conference on Spoken Language Processing (ICSLP 1998) , 1998, p. paper 0917
1998
Earlier work this paper cites.
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (PESQ) - a new method for speech quality assessment of telephone networks and codecs,” in IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , vol. 2, 2001, pp. 749–752 vol.2
2001
Earlier work this paper cites.
“Perceptual evaluation of speech quality (PESQ): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs,” Rec. ITU-T P.862, International Telecommunications Union, Geneva, Switzerland, Recommendation, 2001
2001
Earlier work this paper cites.
E. Vincent, R. Gribonval, and C. Fevotte, “Performance measurement in blind audio source separation,” IEEE Trans. on Audio, Speech, and Lang. Process. (TASLP) , vol. 14, no. 4, pp. 1462–1469, 2006
2006
Earlier work this paper cites.
“Application guide for objective quality measurement based on recommendations P.862, P.862.1 and P.862.2,” Rec. ITU-T P.862.3, International Telecommunications Union, Geneva, Switzerland, Recommendation (Withdrawn), Nov. 2007
2007
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time-frequency weighted noisy speech,” IEEE Trans. on Audio, Speech, and Lang. Process. (TASLP) , vol. 19, no. 7, pp. 2125–2136, 2011
2011
Earlier work this paper cites.
P. Loizou, Speech Enhancement: Theory and Practice, Second Edition . CRC Press, 2013
2013
Earlier work this paper cites.
J. G. Beerends, C. Schmidmer, J. Berger, M. Obermann, R. Ullmann, J. Pomy, and M. Keyhl, “Perceptual objective listening quality assessment (POLQA), the third generation ITU-T standard for end-to-end speech quality measurement part i - temporal alignment,” Journal of the Audio Engineering Society (AES) , vol. 61, no. 6, pp. 366–384, 2013
2013
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in Int. Conf. on Learning Representations (ICLR) , 2015
2015
Cited alongside, same era.
C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Investigating RNN-based speech enhancement methods for noise-robust text-to-speech,” in 9th ISCA Workshop on Speech Synthesis Workshop (SSW 9) , 2016
2016
Cited alongside, same era.
Y. Koizumi, K. Niwa, Y. Hioka, K. Kobayashi, and Y. Haneda, “DNN-based source enhancement to increase objective sound quality assessment score,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 10, pp. 1780–1792, 2018
2018
Cited alongside, same era.
J. M. Martin-Doñas, A. M. Gomez, J. A. Gonzalez, and A. M. Peinado, “A deep learning loss function based on the perceptual evaluation of the speech quality,” IEEE Signal Process. Lett. (SPL) , vol. 25, no. 11, pp. 1680–1684, 2018
2018
Cited alongside, same era.
S.-W. Fu, C. Yu, T.-A. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y. Tsao, “MetricGAN+: An improved version of MetricGAN for speech enhancement,” in Interspeech , 2021, pp. 201–205
2021
Later among the works it cites.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in Int. Conf. on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
Z. Xu, M. Strake, and T. Fingscheidt, “Deep noise suppression maximizing non-differentiable PESQ mediated by a non-intrusive PESQNet,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1572–1585, 2022
2022
Later among the works it cites.
X. Bie, S. Leglaive, X. Alameda-Pineda, and L. Girin, “Unsupervised speech enhancement using dynamical variational autoencoders,” IEEE Trans. on Audio, Speech, and Lang. Process. (TASLP) , vol. 30, pp. 2993–3007, 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Perceptual objective listening quality prediction,” Rec. ITU-T P.863, International Telecommunications Union, Recommendation, Mar. 2018
2018
Cited alongside, same era.
J. L. Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR - half-baked or well done?” in IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , 2019, pp. 626–630
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Cited alongside, same era.
S.-W. Fu, C.-F. Liao, and Y. Tsao, “Learning with learned loss function: Speech enhancement with quality-net to improve perceptual evaluation of speech quality,” IEEE Signal Process. Lett. (SPL) , vol. 27, pp. 26–30, 2020
2020
Cited alongside, same era.
S. Kriman, S. Beliaev, B. Ginsburg, J. Huang, O. Kuchaiev, V. Lavrukhin, R. Leary, J. Li, and Y. Zhang, “Quartznet: Deep automatic speech recognition with 1d time-channel separable convolutions,” in IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , 2020, pp. 6124–6128
2020
Cited alongside, same era.
J. Heitkaemper, D. Jakobeit, C. Boeddeker, L. Drude, and R. Haeb-Umbach, “Demystifying tasnet: A dissecting approach,” in IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , 2020, pp. 6359–6363
2020
Cited alongside, same era.
C. K. A. Reddy, V. Gopal, and R. Cutler, “DNSMOS P.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,” in IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , 2022, pp. 886–890
2022
Later among the works it cites.
R. Cao, S. Abdulatif, and B. Yang, “CMGAN: Conformer-based metric GAN for speech enhancement,” in Interspeech , 2022, pp. 936–940
2022
Later among the works it cites.
J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, “Speech enhancement and dereverberation with diffusion-based generative models,” IEEE Trans. on Audio, Speech, and Lang. Process. (TASLP) , vol. 31, pp. 2351–2364, 2023
2023
Later among the works it cites.
J.-M. Lemercier, J. Richter, S. Welker, and T. Gerkmann, “Analysing diffusion-based generative approaches versus discriminative approaches for speech restoration,” in IEEE Int. Conf. on Acoustics, Speech and Signal Process. (ICASSP) , 2023, pp. 1–5
2023
Later among the works it cites.
2024
Closest in time.
S.-W. Fu, C.-F. Liao, Y. Tsao, and S.-D. Lin, “MetricGAN: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” in Int. Conf. on Machine Learning (ICML) , ser. Proceedings of Machine Learning Research, vol. 97. PMLR, 2019, pp. 2031–2041
2041
Closest in time.