Fetching the paper…
Reading the bibliography…
Recently, convolution-augmented transformer (Conformer) has achieved promising performance in automatic speech recognition (ASR) and time-domain speech enhancement (SE), as it can capture both local and global dependencies in the speech signal.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,”
A. Rix, J. Beerends, M. Hollier and A. Hekstra, · 2001
Earlier work this paper cites.
“Evaluation of objective quality measures for speech enhancement,”
Y. Hu and P. C. Loizou, · 2008
Earlier work this paper cites.
“A short-time objective intelligibility measure for time-frequency weighted noisy speech,”
C. H. Taal, R. C. Hendriks, R. Heusdens and J. Jensen, · 2010
Earlier work this paper cites.
Speech Enhancement: Theory and Practice
P. C. Loizou, · 2013
Earlier work this paper cites.
C. Veaux, J. Yamagishi and S. King, “The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,” in Oriental International Conference on Speech Database and Assessments COCOSDA , 2013, pp. 1–4
2013
Earlier work this paper cites.
J. Thiemann, N. Ito and E. Vincent, “The diverse environments multi-channel acoustic noise database (DEMAND): A database of multichannel environmental noise recordings,” in Proceedings of Meetings on Acoustics , vol. 19, no. 1, Acoustical Society of America, 2013, pp. 035081
2013
Earlier work this paper cites.
J. Barker, R. Marxer, E. Vincent and S. Watanabe, “The third ‘CHiME’speech separation and recognition challenge: Dataset, task and baselines,” in IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) . 2015, pp. 504–511
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in IEEE International Conference on Computer Vision (ICCV) , 2015, pp. 1026–1034
2015
Earlier work this paper cites.
C. Valentini-Botinhao, X. Wang, S. Takaki and J. Yamagishi, “Investigating RNN-based speech enhancement methods for noise-robust text-to-speech.” in SSW , 2016, pp. 146–152
2016
Earlier work this paper cites.
K. Kinoshita et al. , “A summary of the reverb challenge: State-of-the-art and remaining challenges in reverberant speech processing research,” Journal on Advances in Signal Processing , vol. 2016, no. 01, pp. 1–19, 2016
2016
Earlier work this paper cites.
“Complex ratio masking for monaural speech separation,”
D. S. Williamson, Y. Wang and P. Wang, · 2016
Earlier work this paper cites.
“SEGAN: Speech enhancement generative adversarial network,”
S. Pascual, A. Bonafonte and J. Serra, · 2017
Earlier work this paper cites.
A. Vaswani et al. , “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Cited alongside, same era.
X. Mao et al. , “Least squares generative adversarial networks,” in IEEE International Conference on Computer Vision (ICCV) , 2017, pp. 2813–2821
2017
Cited alongside, same era.
P. Isola et al. , “Image-to-Image translation with conditional adversarial networks,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 5967–5976
2017
Cited alongside, same era.
D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE/ACM Transactions on Audio, Speech and Language Processing , vol. 26, no. 10, pp. 1702–1726, 2018
2018
Cited alongside, same era.
A. Pandey and D. Wang, “Densely connected neural network with dilated convolutions for real-time speech enhancement in the time domain,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 6629–6633
2020
Later among the works it cites.
“Dual-branch attention-in-attention transformer for single-channel speech enhancement,”
G. Yu et al. , · 2021
Later among the works it cites.
K. Wang, B. He and W.-P. Zhu, “TSTNN: Two-stage transformer based neural network for speech enhancement in the time domain,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 7098–7102
2021
Later among the works it cites.
E. Kim and H. Seo, “SE-Conformer: Time-Domain Speech Enhancement Using Conformer,” in Proc. Interspeech , 2021, pp. 2736–2740
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Macartney and T. Weyde, · 2018
Cited alongside, same era.
J. Lee, J. Skoglund, T. Shabestary and H.-G. Kang, “Phase-sensitive joint learning algorithms for deep learning-based speech enhancement,” IEEE Signal Processing Letters , vol. 25, no. 8, pp. 1276–1280, 2018
2018
Cited alongside, same era.
K. Wilson et al. , “Exploring tradeoffs in models for low-latency speech enhancement,” in 16th International Workshop on Acoustic Signal Enhancement (IWAENC) , 2018, pp. 366–370
2018
Cited alongside, same era.
D. Yin, C. Luo, Z. Xiong and W. Zeng, “Phasen: A phase-and-harmonics-aware speech enhancement network,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 05, 2020, pp. 9458–9465
2020
Cited alongside, same era.
“Real time speech enhancement in the waveform domain,”
A. Defossez, G. Synnaeve and Y. Adi, · 2020
Cited alongside, same era.
S. Abdulatif et al. , “AeGAN: Time-frequency speech denoising via generative adversarial networks,” in 28th European Signal Processing Conference (EUSIPCO) , 2020, pp. 451–455
2020
Cited alongside, same era.
A. Gulati et al. , “Conformer: Convolution-augmented transformer for speech recognition,” in Proc. Interspeech , 2020, pp. 5036–5040
2020
Cited alongside, same era.
J. Chen, Q. Mao and D. Liu, “Dual-path transformer network: Direct context-aware modeling for end-to-end monaural speech separation,” in Proc. Interspeech , 2020, pp. 2642-2646
2020
Cited alongside, same era.
S. Abdulatif et al. , “Investigating cross-domain losses for speech enhancement,” in 29th European Signal Processing Conference (EUSIPCO) , 2021, pp. 411–415
2021
Later among the works it cites.
Z.-Q. Wang, G. Wichern and J. Le Roux, “On the compensation between magnitude and phase in speech separation,” IEEE Signal Processing Letters , vol. 28, pp. 2018–2022, 2021
2021
Later among the works it cites.
S. Chen et al. , “Continuous speech separation with conformer,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 5749–5753
2021
Later among the works it cites.
S. Braun and I. Tashev, “A consolidated view of loss functions for supervised deep learning-based speech enhancement,” in 44th International Conference on Telecommunications and Signal Processing (TSP) , 2021, pp. 72–76
2021
Later among the works it cites.
S.-W. Fu et al. , “MetricGAN+: An improved version of metricGAN for speech enhancement,” in Proc. Interspeech , 2021, pp. 201–205
2021
Later among the works it cites.
H. Dubey et al. , “ICASSP 2022 Deep noise suppression challenge,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022
2022
Closest in time.
A. Li, C. Zheng, L. Zhang and X. Li, “Glance and gaze: A collaborative learning framework for single-channel speech enhancement,” Applied Acoustics , vol. 187, p. 108499, 2022
2022
Closest in time.
S.-W. Fu, C.-F. Liao, Y. Tsao and S.-D. Lin, “MetricGAN: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” in International Conference on Machine Learning . PMLR, 2019, pp. 2031–2041
2041
Closest in time.