Fetching the paper…
Reading the bibliography…
Curriculum learning begins to thrive in the speech enhancement area, which decouples the original spectrum estimation task into multiple easier sub-tasks to achieve better performance.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,”
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, · 2001
Earlier work this paper cites.
“Evaluation of objective quality measures for speech enhancement,”
Y. Hu and P. C. Loizou, · 2007
Earlier work this paper cites.
“A short-time objective intelligibility measure for time-frequency weighted noisy speech,”
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, · 2010
Earlier work this paper cites.
“The importance of phase in speech enhancement,”
K. Paliwal, K. Wójcicki, and B. Shannon, · 2011
Earlier work this paper cites.
Speech enhancement: theory and practice
P. C. Loizou, · 2013
Earlier work this paper cites.
“The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,”
C. Veaux, J. Yamagishi, and S. King, · 2013
Earlier work this paper cites.
“The diverse environments multi-channel acoustic noise database: A database of multichannel environmental noise recordings,”
J. Thiemann, N. Ito, and E. Vincent, · 2013
Earlier work this paper cites.
“On training targets for supervised speech separation,”
Y. Wang, A. Narayanan, and D. L. Wang, · 2014
Earlier work this paper cites.
“A regression approach to speech enhancement based on deep neural networks,”
Y. Xu, J. Du, L-R. Dai, and C-H. Lee, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. Kingma and J. Ba, · 2014
Earlier work this paper cites.
“Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network,”
W. Shi, J. Caballero, F. Huszár, J. Totz, A. P. Aitken, D. Bishop, R.and Rueckert, and Z. Wang, · 2016
Earlier work this paper cites.
“Investigating RNN-based speech enhancement methods for noise-robust text-to-speech,”
C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, · 2016
Earlier work this paper cites.
“SEGAN: Speech enhancement generative adversarial network,”
S. Pascual, A. Bonafonte, and J. Serra, · 2017
Cited alongside, same era.
“Supervised speech separation based on deep learning: An overview,”
D. L. Wang and J. Chen, · 2018
Cited alongside, same era.
“Time-frequency masking-based speech enhancement using generative adversarial network,”
M. H. Soni, N. Shah, and H. A. Patil, · 2018
Cited alongside, same era.
“Learning complex spectral mapping with gated convolutional recurrent networks for monaural speech enhancement,”
K. Tan and D. L. Wang, · 2019
Cited alongside, same era.
“Phase-aware Speech Enhancement with Deep Complex U-Net,”
H. S. Choi, J. H. Kim, J. Huh, A. Kim, J. W. Ha, and K. Lee, · 2019
Cited alongside, same era.
“Speech enhancement using self-adaptation and multi-head self-attention,”
Y. Koizumi, K. Yatabe, M. Delcroix, Y. Masuyama, and D. Takeuchi, · 2020
Later among the works it cites.
“T-gsa: Transformer with gaussian-weighted self-attention for speech enhancement,”
J. Kim, M. El-Khamy, and J. Lee, · 2020
Later among the works it cites.
“Two Heads are Better Than One: A Two-Stage Complex Spectral Mapping Approach for Monaural Speech Enhancement,”
A. Li, W. Liu, C. Zheng, C. Fan, and X. Li, · 2021
Closest in time.
“TSTNN: Two-Stage Transformer Based Neural Network for Speech Enhancement in the Time Domain,”
K. Wang, B. He, and W. P. Zhu, · 2021
Closest in time.
“A simultaneous denoising and dereverberation framework with target decoupling,”
A. Li, W. Liu, X. Luo, G. Yu, C. Zheng, and X. Li, · 2021
Closest in time.
“Glance and Gaze: A Collaborative Learning Framework for Single-channel Speech Enhancement,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Pandey and D. Wang, · 2019
Cited alongside, same era.
“Dccrn: Deep complex convolution recurrent network for phase-aware speech enhancement,”
Y. Hu, Y. Liu, S. Lv, M. Xing, and L. Xie, · 2020
Cited alongside, same era.
“Real time speech enhancement in the waveform domain,”
A. Defossez, G. Synnaeve, and Y. Adi, · 2020
Cited alongside, same era.
“Dual-path transformer network: Direct context-aware modeling for end-to-end monaural speech separation,”
J. Chen, Q. Mao, and D. Liu, · 2020
Cited alongside, same era.
“On loss functions and recurrency training for gan-based speech enhancement systems,”
Z. Zhang, C. Deng, Y. Shen, D. S. Williamson, Y. Sha, Y. Zhang, H. Song, and X. Li, · 2020
Cited alongside, same era.
“Deep residual-dense lattice network for speech enhancement,”
M. Nikzad, A. Nicolson, Y. Gao, J. Zhou, K. K. Paliwal, and F. Shang, · 2020
Cited alongside, same era.
“PHASEN: A phase-and-harmonics-aware speech enhancement network,”
D. Yin, C. Luo, Z. Xiong, and W. Zeng, · 2020
Cited alongside, same era.
A. Li, C. Zheng, L. Zhang, and X. Li, · 2021
Closest in time.
“On The Compensation Between Magnitude and Phase in Speech Separation,”
Z.-Q. Wang, G. Wichern, and J. Le Roux, · 2021
Closest in time.
“Cyclegan-based non-parallel speech enhancement with an adaptive attention-in-attention mechanism,”
G. Yu, Y. Wang, C. Zheng, H. Wang, and Q. Zhang, · 2021
Closest in time.
“On the importance of power compression and phase estimation in monaural speech dereverberation,”
A. Li, C. Zheng, R. Peng, and X. Li, · 2021
Closest in time.
“MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement,”
S.-W. Fu, C. Yu, T.-A. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y. Tsao, · 2021
Closest in time.
“Se-Conformer: Time-Domain Speech Enhancement using Conformer,”
E. Kim and H. Seo, · 2021
Closest in time.
“MetricGAN: Generative adversarial networks based black-box metric scores optimization for speech enhancement,”
S.-W. Fu, C.-F. Liao, Y. Tsao, and S.-D. Lin, · 2041
Closest in time.