Fetching the paper…
Reading the bibliography…
While traditional statistical signal processing model-based methods can derive the optimal estimators relying on specific statistical assumptions, current learning-based methods further promote the performance upper bound via deep neural networks but at the expense of high encapsulation and lack adequate interpretability.
Y. Ephraim and D. Malah, “Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator,” IEEE Trans. Acoust., Speech, Signal Process. , vol. 32, no. 6, pp. 1109–1121, 1984
1984
Earlier work this paper cites.
D. B. Paul and J. Baker, “The design for the wall street journal-based csr corpus,” in Workshop on Speech and Natural Lang. , 1992, pp. 357–362
1992
Earlier work this paper cites.
A. Varga and H. J. Steeneken, “Assessment for automatic speech recognition: II. NOISEX-92: A database and an experiment to study the effect of additive noise on speech recognition systems,” Speech. Commun. , vol. 12, no. 3, pp. 247–251, 1993
1993
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,” in Proc. ICASSP , vol. 2. IEEE, 2001, pp. 749–752
2001
Earlier work this paper cites.
R. I.-T. P. ITU, “862.2: Wideband extension to recommendation p. 862 for the assessment of wideband telephone networks and speech codecs. itu-telecommunication standardization sector, 2007.”
2007
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “A short-time objective intelligibility measure for time-frequency weighted noisy speech,” in Proc. ICASSP . IEEE, 2010, pp. 4214–4217
2010
Earlier work this paper cites.
T. Gerkmann and M. Krawczyk, “MMSE-optimal spectral amplitude estimation given the STFT-phase,” IEEE Signal Process. Lett. , vol. 20, no. 2, pp. 129–132, 2012
2012
Earlier work this paper cites.
P. C. Loizou, Speech enhancement: theory and practice . CRC press, 2013
2013
Earlier work this paper cites.
T. Gerkmann, “Bayesian estimation of clean speech spectral coefficients given a priori knowledge of the phase,” IEEE Trans. Signal Process. , vol. 62, no. 16, pp. 4199–4208, 2014
2014
Earlier work this paper cites.
T. Yoshioka, N. Ito, M. Delcroix, A. Ogawa, K. Kinoshita, M. Fujimoto, C. Yu, W. J. Fabian, M. Espi, T. Higuchi et al. , “The NTT CHiME-3 system: Advances in speech enhancement and recognition for mobile multi-microphone devices,” in Proc. ASRU . IEEE, 2015, pp. 436–443
2015
Earlier work this paper cites.
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The third “CHiME” speech separation and recognition challenge: Dataset, task and baselines,” in Proc. ASRU . IEEE, 2015, pp. 504–511
2015
Earlier work this paper cites.
J. Jensen and C. H. Taal, “An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,” IEEE/ACM Trans. Audio. Speech, Lang. Process. , vol. 24, no. 11, pp. 2009–2022, 2016
2016
Earlier work this paper cites.
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier, “Language modeling with gated convolutional networks,” in Proc. ICML . PMLR, 2017, pp. 933–941
2017
Earlier work this paper cites.
D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE/ACM Trans. Audio. Speech, Lang. Process. , vol. 26, no. 10, pp. 1702–1726, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE/ACM Trans. Audio. Speech, Lang. Process. , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Cited alongside, same era.
S. Wisdom, J. R. Hershey, K. Wilson, J. Thorpe, M. Chinen, B. Patton, and R. A. Saurous, “Differentiable consistency constraints for improved deep speech enhancement,” in Proc. ICASSP . IEEE, 2019, pp. 900–904
2019
Cited alongside, same era.
Y. Hu, Y. Liu, S. Lv, M. Xing, S. Zhang, Y. Fu, J. Wu, B. Zhang, and L. Xie, “DCCRN: Deep complex convolution recurrent network for phase-aware speech enhancement,” in Proc. Interspeech , 2020, pp. 2472–2476
2020
Later among the works it cites.
A. Pandey and D. Wang, “Densely connected neural network with dilated convolutions for real-time speech enhancement in the time domain,” in Proc. ICASSP , 2020, pp. 6629–6633
2020
Later among the works it cites.
A. Defossez, G. Synnaeve, and Y. Adi, “Real time speech enhancement in the waveform domain,” in Proc. Interspeech , 2020, pp. 3291–3295
2020
Later among the works it cites.
D. Yin, C. Luo, Z. Xiong, and W. Zeng, “PHASEN: A phase-and-harmonics-aware speech enhancement network,” in Proc. AAAI , 2020, pp. 9458–9465
2020
Later among the works it cites.
A. Li, W. Liu, X. Luo, G. Yu, C. Zheng, and X. Li, “A simultaneous denoising and dereverberation framework with target decoupling,” in Proc. Interspeech , 2021, pp. 2801–2805
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “Sdr–half-baked or well done?” in Proc. ICASSP . IEEE, 2019, pp. 626–630
2019
Cited alongside, same era.
K. Tan and D. Wang, “Learning complex spectral mapping with gated convolutional recurrent networks for monaural speech enhancement,” IEEE/ACM Trans. Audio. Speech, Lang. Process. , vol. 28, pp. 380–390, 2020
2020
Cited alongside, same era.
A. Li, M. Yuan, C. Zheng, and X. Li, “Speech enhancement using progressive learning-based convolutional recurrent neural network,” Appl. Acoust. , vol. 166, p. 107347, 2020
2020
Cited alongside, same era.
X. Hao, X. Su, S. Wen, Z. Wang, Y. Pan, F. Bao, and W. Chen, “Masking and inpainting: A two-stage speech enhancement approach for low SNR and non-stationary noise,” in Proc. ICASSP . IEEE, 2020, pp. 6959–6963
2020
Cited alongside, same era.
X. Qin, Z. Zhang, C. Huang, M. Dehghan, O. R. Zaiane, and M. Jagersand, “U2-net: Going deeper with nested U-structure for salient object detection,” Pattern Recognition , vol. 106, p. 107404, 2020
2020
Cited alongside, same era.
Q. Zhang, A. Nicolson, M. Wang, K. K. Paliwal, and C. Wang, “DeepMMSE: A deep learning approach to mmse-based noise power spectral density estimation,” IEEE/ACM Trans. Audio. Speech, Lang. Process. , vol. 28, pp. 1404–1415, 2020
2020
Cited alongside, same era.
C. Reddy, V. Gopal, R. Culter, E. Beyrami, R. Cheng, H. Dubey, S. Matusevych, R. Aichner, A. Aazami, S. Braun, P. Rana, S. Srinivasan, and J. Gehrke, “The interspeech 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,” in Proc. Interspeech , 2020, pp. 2492–2496
2020
Cited alongside, same era.
N. Westhausen and B. Meyer, “Dual-signal transformation lstm network for real-time noise suppression,” in Proc. Interspeech , 2020, pp. 2477–2481
2020
Cited alongside, same era.
2021
Later among the works it cites.
Z.-Q. Wang, G. Wichern, and J. Le Roux, “On The Compensation Between Magnitude and Phase in Speech Separation,” IEEE Signal Process. Lett. , vol. 28, pp. 2018–2022, 2021
2021
Later among the works it cites.
C. Zheng, X. Peng, Y. Zhang, S. Srinivasan, and Y. Lu, “Interactive Speech and Noise Modeling for Speech Enhancement,” in Proc. AAAI , vol. 35, no. 16, 2021, pp. 14 549–14 557
2021
Later among the works it cites.
A. Li, W. Liu, C. Zheng, C. Fan, and X. Li, “Two Heads are Better Than One: A Two-Stage Complex Spectral Mapping Approach for Monaural Speech Enhancement,” IEEE/ACM Trans. Audio. Speech, Lang. Process. , vol. 29, pp. 1829–1843, 2021
2021
Later among the works it cites.
A. Li, C. Zheng, R. Peng, and X. Li, “On the importance of power compression and phase estimation in monaural speech dereverberation,” JASA Express Letters , vol. 1, no. 1, p. 014802, 2021
2021
Later among the works it cites.
X. Hao, X. Su, R. Horaud, and X. Li, “Fullsubnet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement,” in Proc. ICASSP , 2021, pp. 6633–6637
2021
Later among the works it cites.
H. Choi, S. Park, J. Lee, H. Heo, D. Jeon, and K. Lee, “Real-Time Denoising and Dereverberation wtih Tiny Recurrent U-Net,” in Proc. ICASSP , 2021, pp. 5789–5793
2021
Later among the works it cites.
A. Li, C. Zheng, L. Zhang, and X. Li, “Glance and gaze: A collaborative learning framework for single-channel speech enhancement,” Appl. Acoust. , vol. 187, p. 108499, 2022
2022
Closest in time.