Fetching the paper…
Reading the bibliography…
Speech enhancement (SE) aims to improve speech quality and intelligibility, which are both related to a smooth transition in speech segments that may carry linguistic information, e.g.
S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,” IEEE TASSP , vol. 27, no. 2, pp. 113–120, 1979
1979
Earlier work this paper cites.
J. Lim and A. Oppenheim, “Enhancement and bandwidth compression of noisy speech,” Proceedings of the IEEE , vol. 67, no. 12, pp. 1586–1604, 1979
1979
Earlier work this paper cites.
I. Olkin and F. Pukelsheim, “The distance between two random vectors with given dispersion matrices,” Linear Algebra and its Applications , vol. 48, pp. 257–263, 1982
1982
Earlier work this paper cites.
K. Paliwal and A. Basu, “A speech enhancement method based on kalman filtering,” in Proc. ICASSP , 1987
1987
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,” in Proc. ICASSP , 2001
2001
Earlier work this paper cites.
C. Villani, Optimal transport – Old and new . Springer Science & Business Media, 2008, vol. 338, pp. xxii+973
2008
Earlier work this paper cites.
Y. Hu and P. Loizou, “Evaluation of objective quality measures for speech enhancement,” IEEE TSAP , vol. 16, pp. 229–238, 2008
2008
Earlier work this paper cites.
2010
Earlier work this paper cites.
P. Loizou, Speech enhancement: theory and practice . CRC press, 2013
2013
Earlier work this paper cites.
X. Lu, Y. Tsao, S. Matsuda, and C. Hori, “Speech enhancement based on deep denoising autoencoder.” in Proc. Interspeech , vol. 2013, 2013, pp. 436–440
2013
Earlier work this paper cites.
C. Veaux, J. Yamagishi, and S. King, “The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,” in Proc. O-COCOSDA/CASLRE , 2013
2013
Earlier work this paper cites.
J. Thiemann, N. Ito, and E. Vincent, “Demand: a collection of multi-channel recordings of acoustic noise in diverse environments,” in Proc. Meetings Acoust , 2013
2013
Earlier work this paper cites.
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “A regression approach to speech enhancement based on deep neural networks,” IEEE/ACM TASLP , vol. 23, no. 1, pp. 7–19, 2014
2014
Earlier work this paper cites.
F. Weninger, F. Eyben, and B. Schuller, “Single-channel speech separation with memory-enhanced recurrent neural networks,” in Proc. ICASSP , 2014
2014
Earlier work this paper cites.
F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. L. Roux, J. Hershey, and B. Schuller, “Speech enhancement with lstm recurrent neural networks and its application to noise-robust asr,” in Proc. LVA/ICA , 2015
2015
Earlier work this paper cites.
J. Johnson, A. Alahi, and F.-F. Li, “Perceptual losses for real-time style transfer and super-resolution,” in Proc. ECCV , 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
S. S. Pascual, A. Bonafonte, and J. Serrà, “Segan: Speech enhancement generative adversarial network,” in Proc. Interspeech , 2017
2017
Cited alongside, same era.
X. Huang and S. J. Belongie, “Arbitrary style transfer in real-time with adaptive instance normalization.” in Proc. ICCV , 2017
2017
Cited alongside, same era.
H. Zhao, S. Zarar, I. Tashev, and C.-H. Lee, “Convolutional-recurrent neural networks for speech enhancement,” in Proc. ICASSP , 2018
2018
2018
Later among the works it cites.
H.-S. Choi, J.-H. Kim, J. Huh, A. Kim, J.-W. Ha, and K. Lee, “Phase-aware speech enhancement with deep complex u-net,” in Proc. ICLR , 2018
2018
Later among the works it cites.
P. Rodríguez, M. A. Bautista, J. Gonzalez, and S. Escalera, “Beyond one-hot encoding: Lower dimensional target embedding,” Image and Vision Computing , vol. 75, pp. 21–31, 2018
2018
Later among the works it cites.
A. Pandey and D. Wang, “Tcnn: Temporal convolutional neural network for real-time speech enhancement in the time domain,” in Proc. ICASSP , 2019
2019
Later among the works it cites.
D. Baby and S. Verhulst, “Sergan: Speech enhancement using relativistic generative adversarial networks with gradient penalty,” in Proc. ICASSP , 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
S.-W. Fu, T.-W. Wang, Y. Tsao, X. Lu, and H. Kawai, “End-to-end waveform utterance enhancement for direct evaluation metrics optimization by fully convolutional neural networks,” IEEE/ACM TASLP , vol. 26, no. 9, pp. 1570–1584, 2018
2018
Cited alongside, same era.
K. Tan and D. Wang, “A convolutional recurrent neural network for real-time speech enhancement.” in Proc. Interspeech , 2018
2018
Cited alongside, same era.
M. Soni, N. Shah, and H. Patil, “Time-frequency masking-based speech enhancement using generative adversarial network,” in Proc. ICASSP , 2018
2018
Cited alongside, same era.
A. Pandey and D. Wang, “On adversarial training and loss functions for speech enhancement,” in Proc. ICASSP , 2018
2018
Cited alongside, same era.
S. Qin and T. Jiang, “Improved wasserstein conditional generative adversarial network speech enhancement,” EURASIP Journal on Wireless Communications and Networking , vol. 2018, no. 1, p. 181, 2018
2018
Cited alongside, same era.
J. Martín-Doñas, A. Gomez, J. Gonzalez, and A. Peinado, “A deep learning loss function based on the perceptual evaluation of the speech quality,” IEEE Signal processing letters , vol. 25, no. 11, pp. 1680–1684, 2018
2018
Cited alongside, same era.
M. Kolbæk, Z.-H. Tan, and J. Jensen, “Monaural speech enhancement using deep neural networks by maximizing a short-time objective intelligibility measure,” in Proc. ICASSP , 2018
2018
Cited alongside, same era.
2019
Later among the works it cites.
S.-W. Fu, C.-F. Liao, Y. Tsao, and S.-D. Lin, “Metricgan: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” in Proc. ICML , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
S.-W. Fu, C.-F. Liao, and Y. Tsao, “Learning with learned loss function: Speech enhancement with quality-net to improve perceptual evaluation of speech quality,” IEEE Signal Processing Letters , vol. 27, pp. 26–30, 2019
2019
Later among the works it cites.
F. Germain, Q. Chen, and V. Koltun, “Speech denoising with deep feature losses,” Proc. Interspeech , 2019
2019
Later among the works it cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition,” Proc. Interspeech , 2019
2019
Later among the works it cites.
J. Yao and A. Al-Dahle, “Coarse-to-fine optimization for speech enhancement.” in Proc. Interspeech , 2019
2019
Later among the works it cites.
J. Su, Z. Jin, and A. Finkelstein, “Hifi-gan: High-fidelity denoising and dereverberation based on speech deep features in adversarial networks,” Proc. Interspeech , 2020
2020
Closest in time.
2020
Closest in time.
S. Kataria, J. Villalba, and N. Dehak, “Perceptual loss based speech denoising with an ensemble of audio pattern recognition and self-supervised models,” in Proc. ICASSP , 2021
2021
Closest in time.