Fetching the paper…
Reading the bibliography…
Phase information has a significant impact on speech perceptual quality and intelligibility.
A. Gray and J. Markel, “Distance measures for speech processing,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 24, no. 5, pp. 380–391, 1976
1976
Earlier work this paper cites.
J. Lim and A. Oppenheim, “All-pole modeling of degraded speech,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 26, no. 3, pp. 197–210, 1978
1978
Earlier work this paper cites.
M. Berouti, R. Schwartz, and J. Makhoul, “Enhancement of speech corrupted by acoustic noise,” in Proc. ICASSP , vol. 4, 1979, pp. 208–211
1979
Earlier work this paper cites.
D. Wang and J. Lim, “The unimportance of phase in speech enhancement,” IEEE/ACM Transactions on Acoustics, Speech, and Signal Processing , vol. 30, no. 4, pp. 679–681, 1982
1982
Earlier work this paper cites.
M. Dendrinos, S. Bakamidis, and G. Carayannis, “Speech enhancement from noise: A regenerative approach,” Speech Communication , vol. 10, no. 1, pp. 45–57, 1991
1991
Earlier work this paper cites.
Y. Ephraim, “Statistical-model-based speech enhancement systems,” Proceedings of the IEEE , vol. 80, no. 10, pp. 1526–1555, 1992
1992
Earlier work this paper cites.
Y. Ephraim and H. L. Van Trees, “A signal subspace approach for speech enhancement,” IEEE Transactions on Speech and Audio Processing , vol. 3, no. 4, pp. 251–266, 1995
1995
Earlier work this paper cites.
T. Robinson, J. Fransen, D. Pye, J. Foote, and S. Renals, “Wsjcam0: a british english speech corpus for large vocabulary continuous speech recognition,” in Proc. ICASSP , vol. 1, 1995, pp. 81–84
1995
Earlier work this paper cites.
G. Hu and D. Wang, “Speech segregation based on pitch tracking and amplitude modulation,” in Proc. WASPAA , 2001, pp. 79–82
2001
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,” in Proc. ICASSP , vol. 2, 2001, pp. 749–752
2001
Earlier work this paper cites.
M. Lincoln, I. McCowan, J. Vepa, and H. K. Maganti, “The multi-channel wall street journal audio visual corpus (MC-WSJ-AV): Specification and initial experiments,” in Proc. ASRU , 2005, pp. 357–362
2005
Earlier work this paper cites.
S. Srinivasan, N. Roman, and D. Wang, “Binary and ratio time-frequency masks for robust speech recognition,” Speech Communication , vol. 48, no. 11, pp. 1486–1501, 2006
2006
Earlier work this paper cites.
J. Benesty, S. Makino, and J. Chen, Speech enhancement . Springer Science & Business Media, 2006
2006
Earlier work this paper cites.
Y. Hu and P. C. Loizou, “Evaluation of objective quality measures for speech enhancement,” IEEE Transactions on audio, speech, and language processing , vol. 16, no. 1, pp. 229–238, 2007
2007
Earlier work this paper cites.
J. Le Roux, N. Ono, and S. Sagayama, “Explicit consistency constraints for STFT spectrograms and their application to phase reconstruction.” in Proc. SAPA , 2008, pp. 23–28
2008
Earlier work this paper cites.
T. Nakatani, T. Yoshioka, K. Kinoshita, M. Miyoshi, and B.-H. Juang, “Speech dereverberation based on variance-normalized delayed linear prediction,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 18, no. 7, pp. 1717–1731, 2010
2010
Earlier work this paper cites.
J. Le Roux, H. Kameoka, N. Ono, and S. Sagayama, “Fast signal reconstruction from magnitude STFT spectrogram based on spectrogram consistency,” in Proc. DAFx , vol. 10, 2010, pp. 397–403
2010
Earlier work this paper cites.
K. Paliwal, K. Wójcicki, and B. Shannon, “The importance of phase in speech enhancement,” Speech Communication , vol. 53, no. 4, pp. 465–494, 2011
2011
Earlier work this paper cites.
N. Roman and J. Woodruff, “Intelligibility of reverberant noisy speech with ideal binary masking,” The Journal of the Acoustical Society of America , vol. 130, no. 4, pp. 2153–2161, 2011
2011
Earlier work this paper cites.
X. Glorot, A. Bordes, and Y. Bengio, “Deep sparse rectifier neural networks,” in Proc. AISTATS , 2011, pp. 315–323
2011
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time-frequency weighted noisy speech,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 19, no. 7, pp. 2125–2136, 2011
2011
Earlier work this paper cites.
A. Narayanan and D. Wang, “Ideal ratio mask estimation using deep neural networks for robust speech recognition,” in Proc. ICASSP , 2013, pp. 7092–7096
2013
Earlier work this paper cites.
C. Veaux, J. Yamagishi, and S. King, “The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,” in Proc. O-COCOSDA/CASLRE , 2013, pp. 1–4
2013
Earlier work this paper cites.
J. Thiemann, N. Ito, and E. Vincent, “The diverse environments multi-channel acoustic noise database (DEMAND): A database of multichannel environmental noise recordings,” in Proc. ICA , vol. 19, no. 1, 2013, p. 035081
2013
Earlier work this paper cites.
J. L. Desjardins and K. A. Doherty, “The effect of hearing aid noise reduction on listening effort in hearing-impaired adults,” Ear and hearing , vol. 35, no. 6, pp. 600–610, 2014
2014
Earlier work this paper cites.
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “A regression approach to speech enhancement based on deep neural networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 1, pp. 7–19, 2014
2014
Earlier work this paper cites.
Y. Wang, A. Narayanan, and D. Wang, “On training targets for supervised speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 22, no. 12, pp. 1849–1858, 2014
2014
Earlier work this paper cites.
F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. Le Roux, J. R. Hershey, and B. Schuller, “Speech enhancement with LSTM recurrent neural networks and its application to noise-robust ASR,” in International Conference on Latent Variable Analysis and Signal Separation (LVA/ICA) , 2015, pp. 91–99
2015
Earlier work this paper cites.
H. Erdogan, J. R. Hershey, S. Watanabe, and J. Le Roux, “Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks,” in Proc. ICASSP , 2015, pp. 708–712
2015
Cited alongside, same era.
D. S. Williamson, Y. Wang, and D. Wang, “Complex ratio masking for monaural speech separation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 24, no. 3, pp. 483–492, 2015
2015
Cited alongside, same era.
Z.-H. Ling, S.-Y. Kang, H. Zen, A. Senior, M. Schuster, X.-J. Qian, H. M. Meng, and L. Deng, “Deep learning for acoustic modeling in parametric speech generation: A systematic review of existing techniques and future trends,” IEEE Signal Processing Magazine , vol. 32, no. 3, pp. 35–52, 2015
2015
Cited alongside, same era.
K. Li and C.-H. Lee, “A deep neural network approach to speech bandwidth expansion,” in Proc. ICASSP , 2015, pp. 4395–4399
2015
Cited alongside, same era.
D. Yin, C. Luo, Z. Xiong, and W. Zeng, “PHASEN: A phase-and-harmonics-aware speech enhancement network,” in Proc. AAAI , vol. 34, no. 05, 2020, pp. 9458–9465
2020
Later among the works it cites.
A. Pandey and D. Wang, “Densely connected neural network with dilated convolutions for real-time speech enhancement in the time domain,” in Proc. ICASSP , 2020, pp. 6629–6633
2020
Later among the works it cites.
C. K. Reddy, V. Gopal, R. Cutler, E. Beyrami, R. Cheng, H. Dubey, S. Matusevych, R. Aichner, A. Aazami, S. Braun et al. , “The INTERSPEECH 2020 deep noise suppression challenge: Datasets, subjective testing framework, and challenge results,” in Proc. Interspeech , 2020, pp. 2492–2496
2020
Later among the works it cites.
A. Gritsenko, T. Salimans, R. van den Berg, J. Snoek, and N. Kalchbrenner, “A spectral energy distance for parallel speech synthesis,” Proc. NeurIPS , vol. 33, pp. 13 062–13 072, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in Proc. ICCV , 2015, pp. 1026–1034
2015
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Cited alongside, same era.
W. Shi, J. Caballero, F. Huszár, J. Totz, A. P. Aitken, R. Bishop, D. Rueckert, and Z. Wang, “Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network,” in Proc. CVPR , 2016, pp. 1874–1883
2016
Cited alongside, same era.
C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Investigating RNN-based speech enhancement methods for noise-robust text-to-speech.” in Proc. SSW , 2016, pp. 146–152
2016
Cited alongside, same era.
K. Kinoshita, M. Delcroix, S. Gannot, E. A. P. Habets, R. Haeb-Umbach, W. Kellermann, V. Leutnant, R. Maas, T. Nakatani, B. Raj et al. , “A summary of the REVERB challenge: state-of-the-art and remaining challenges in reverberant speech processing research,” EURASIP Journal on Advances in Signal Processing , vol. 2016, pp. 1–19, 2016
2016
Cited alongside, same era.
S. Pascual, A. Bonafonte, and J. Serrà, “SEGAN: Speech enhancement generative adversarial network,” in Proc. Interspeech , 2017, pp. 3642–3646
2017
Cited alongside, same era.
V. Kuleshov, S. Z. Enam, and S. Ermon, “Audio super-resolution using neural nets,” in Proc. ICLR (Workshop Track) , 2017
2017
Cited alongside, same era.
M. Chinen, F. S. Lim, J. Skoglund, N. Gureev, F. O’Gorman, and A. Hines, “ViSQOL v3: An open source production ready objective speech and audio metric,” in Proc. QoMEX , 2020, pp. 1–6
2020
Later among the works it cites.
E. Kim and H. Seo, “SE-Conformer: Time-domain speech enhancement using conformer.” in Proc. Interspeech , 2021, pp. 2736–2740
2021
Later among the works it cites.
H. Wang and D. Wang, “Towards robust speech super-resolution,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 2058–2066, 2021
2021
Later among the works it cites.
X. Hao, X. Su, R. Horaud, and X. Li, “FullSubNet: A full-band and sub-band fusion model for real-time single-channel speech enhancement,” in Proc. ICASSP , 2021, pp. 6633–6637
2021
Later among the works it cites.
A. Li, W. Liu, C. Zheng, C. Fan, and X. Li, “Two heads are better than one: A two-stage complex spectral mapping approach for monaural speech enhancement,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 1829–1843, 2021
2021
Later among the works it cites.
Z.-Q. Wang, G. Wichern, and J. Le Roux, “On the compensation between magnitude and phase in speech separation,” IEEE Signal Processing Letters , vol. 28, pp. 2018–2022, 2021
2021
Later among the works it cites.
S.-W. Fu, C. Yu, T.-A. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y. Tsao, “MetricGAN+: An improved version of MetricGAN for speech enhancement,” in Proc. Interspeech , 2021, pp. 201–205
2021
Later among the works it cites.
K. Wang, B. He, and W.-P. Zhu, “TSTNN: Two-stage transformer based neural network for speech enhancement in the time domain,” in Proc. ICASSP , 2021, pp. 7098–7102
2021
Later among the works it cites.
V. Kothapally and J. H. Hansen, “SkipConvGAN: Monaural speech dereverberation using generative adversarial networks via complex time-frequency masking,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1600–1613, 2022
2022
Later among the works it cites.
S. Zhao, B. Ma, K. N. Watcharasupat, and W.-S. Gan, “FRCRN: Boosting feature representation using frequency recurrence for monaural speech enhancement,” in Proc. ICASSP , 2022, pp. 9281–9285
2022
Later among the works it cites.
F. Dang, H. Chen, and P. Zhang, “DPT-FSNet: Dual-path transformer based full-band and sub-band fusion network for speech enhancement,” in Proc. ICASSP , 2022, pp. 6857–6861
2022
Later among the works it cites.
A. Li, S. You, G. Yu, C. Zheng, and X. Li, “Taylor, can you hear me now? a taylor-unfolding framework for monaural speech enhancement,” in Proc. IJCAI , 2022, pp. 4193–4200
2022
Later among the works it cites.
G. Yu, A. Li, C. Zheng, Y. Guo, Y. Wang, and H. Wang, “Dual-branch attention-in-attention transformer for single-channel speech enhancement,” in Proc. ICASSP , 2022, pp. 7847–7851
2022
Later among the works it cites.
2022
Later among the works it cites.
H. Liu, W. Choi, X. Liu, Q. Kong, Q. Tian, and D. Wang, “Neural vocoder is all you need for speech super-resolution,” in Proc. Interspeech , 2022, pp. 4227–4231
2022
Later among the works it cites.
Y. Fu, Y. Liu, J. Li, D. Luo, S. Lv, Y. Jv, and L. Xie, “Uformer: A unet based dilated complex & real dual-path conformer network for simultaneous speech enhancement and dereverberation,” in Proc. ICASSP , 2022, pp. 7417–7421
2022
Later among the works it cites.
Y.-X. Lu, Y. Ai, and Z.-H. Ling, “MP-SENet: A speech enhancement model with parallel denoising of magnitude and phase spectra,” in Proc. Interspeech , 2023, pp. 3834–3838
2023
Closest in time.
D. Yin, Z. Zhao, C. Tang, Z. Xiong, and C. Luo, “TridentSE: Guiding speech enhancement with 32 global tokens,” in Proc. Interspeech , 2023, pp. 3839–3843
2023
Closest in time.
Z.-Q. Wang, S. Cornell, S. Choi, Y. Lee, B.-Y. Kim, and S. Watanabe, “TF-GridNet: Making time-frequency domain models great again for monaural speaker separation,” in Proc. ICASSP , 2023, pp. 1–5
2023
Closest in time.
Y. Ai and Z.-H. Ling, “Neural speech phase prediction based on parallel estimation architecture and anti-wrapping losses,” in Proc. ICASSP , 2023
2023
Closest in time.
L. Liu, H. Guan, J. Ma, W. Dai, G. Wang, and S. Ding, “A mask free neural network for monaural speech enhancement,” in Proc. Interspeech , 2023, pp. 2468–2472
2023
Closest in time.
V. Kothapally and J. H. Hansen, “Monaural speech dereverberation using deformable convolutional networks,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024
2024
Closest in time.
S.-W. Fu, C.-F. Liao, Y. Tsao, and S.-D. Lin, “MetricGAN: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” in Proc. ICML , 2019, pp. 2031–2041
2041
Closest in time.