Fetching the paper…
Reading the bibliography…
This work aims to investigate the use of a recently proposed, attention-free, scalable state-space model (SSM), Mamba, for the speech enhancement (SE) task.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,”
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, · 2001
Earlier work this paper cites.
“Speech processing in vocoder-centric cochlear implants,”
P. C. Loizou, · 2006
Earlier work this paper cites.
“An algorithm for intelligibility prediction of time–frequency weighted noisy speech,”
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, · 2011
Earlier work this paper cites.
Speech Enhancement: Theory and Practice
P. C. Loizou, · 2013
Earlier work this paper cites.
“Speech enhancement based on deep denoising autoencoder,”
X. Lu, Y. Tsao, S. Matsuda, and C. Hori, · 2013
Earlier work this paper cites.
“The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,”
C. Veaux, J. Yamagishi, and S. King, · 2013
Earlier work this paper cites.
“The diverse environments multi-channel acoustic noise database (DEMAND): A database of multichannel environmental noise recordings,”
J. Thiemann, N. Ito, and E. Vincent, · 2013
Earlier work this paper cites.
“An experimental study on speech enhancement based on deep neural networks,”
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, · 2014
Earlier work this paper cites.
“Experiments on deep learning for speech denoising,”
D. Liu, P. Smaragdis, and M. Kim, · 2014
Earlier work this paper cites.
“Speech enhancement and recognition using multi-task learning of long short-term memory recurrent neural networks,”
Z. Chen, S. Watanabe, H. Erdogan, and J. R. Hershey, · 2015
Earlier work this paper cites.
“Investigating RNN-based speech enhancement methods for noise-robust text-to-speech.,”
C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, · 2016
Earlier work this paper cites.
“Deep learning reinvents the hearing aid,”
D. Wang, · 2017
Earlier work this paper cites.
“Conditional generative adversarial networks for speech enhancement and noise-robust speaker verification,”
D. Michelsanti and Z.-H. Tan, · 2017
Earlier work this paper cites.
“Speech enhancement generative adversarial network,”
S. Pascual, A. Bonafonte, and J. Serrà, · 2017
Earlier work this paper cites.
“Supervised speech separation based on deep learning: An overview,”
D. Wang and J. Chen, · 2018
Earlier work this paper cites.
“End-to-end waveform utterance enhancement for direct evaluation metrics optimization by fully convolutional neural networks,”
S.-W. Fu, T.-W. Wang, Y. Tsao, X. Lu, and H. Kawai, · 2018
Cited alongside, same era.
“A convolutional recurrent neural network for real-time speech enhancement.,”
K. Tan and D. Wang, · 2018
Cited alongside, same era.
“SDR-half-baked or well done?,”
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, · 2019
Cited alongside, same era.
“MetricGAN: Generative adversarial networks based black-box metric scores optimization for speech enhancement,”
S.-W. Fu, C.-F. Liao, Y. Tsao, and S.-D. Lin, · 2019
Cited alongside, same era.
“Robust speaker recognition based on single-channel and multi-channel speech enhancement,”
H. Taherian, Z.-Q. Wang, J. Chang, and D. Wang, · 2020
Cited alongside, same era.
V. Zadorozhnyy, Q. Ye, and K. Koishida, · 2022
Later among the works it cites.
“DPT-FSNet: Dual-path transformer based full-band and sub-band fusion network for speech enhancement,”
F. Dang, H. Chen, and P. Zhang, · 2022
Later among the works it cites.
“D4AM: A general denoising framework for downstream acoustic models,”
C.-C. Lee, Y. Tsao, H.-M. Wang, and C.-S. Chen, · 2023
Later among the works it cites.
“MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra,”
Y.-X. Lu, Y. Ai, and Z.-H. Ling, · 2023
Later among the works it cites.
Y.-X. Lu, Y. Ai, and Z.-H. Ling, · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Li, M. Yuan, C. Zheng, and X. Li, · 2020
Cited alongside, same era.
“Speech enhancement using self-adaptation and multi-head self-attention,”
Y. Koizumi, K. Yatabe, M. Delcroix, Y. Masuyama, and D. Takeuchi, · 2020
Cited alongside, same era.
“On mean absolute error for deep neural network based vector-to-vector regression,”
J. Qi, J. Du, S. M. Siniscalchi, X. Ma, and C.-H. Lee, · 2020
Cited alongside, same era.
“Real time speech enhancement in the waveform domain,”
A. Défossez, G. Synnaeve, and Y. Adi, · 2020
Cited alongside, same era.
“MetricGAN+: An improved version of metricgan for speech enhancement,”
S.-W. Fu, C. Yu, T.-A. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y. Tsao, · 2020
Cited alongside, same era.
“Boosting objective scores of a speech enhancement model by metricgan post-processing,”
S.-W. Fu, C.-F. Liao, T.-A. Hsieh, et al., · 2020
Cited alongside, same era.
“Efficiently modeling long sequences with structured state spaces,”
A Gu, K Goel, and C Re, · 2021
Cited alongside, same era.
Later among the works it cites.
“Attention-based speech enhancement using human quality perception modelling,”
K. M. Nayem and D. Williamson, · 2023
Later among the works it cites.
“Mamba: Linear-time sequence modeling with selective state spaces,”
A. Gu and T. Dao, · 2023
Later among the works it cites.
“A multi-dimensional deep structured state space approach to speech enhancement using small-footprint models,”
P.-J. Ku, C.-H. Huck Yang, S. M. Siniscalchi, and C.-H. Lee, · 2023
Later among the works it cites.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever, · 2023
Later among the works it cites.
T. Ochiai, K. Iwamoto, M. Delcroix, R. Ikeshita, H. Sato, S. Araki, and S. Katagiri, · 2024
Closest in time.
“CMGAN: Conformer-based metric-gan for monaural speech enhancement,”
S. Abdulatif, R. Cao, and B. Yang, · 2024
Closest in time.
X. Jiang, C. Han, and N. Mesgarani, · 2024
Closest in time.
“Spmamba: State-space model is all you need in speech separation,”
K. Li and G. Chen, · 2024
Closest in time.
“Dual-branch modeling based on state-space model for speech enhancement,”
L. Sun, S. Yuan, A. Gong, L. Ye, and E. S. Chng, · 2024
Closest in time.
“Spiking structured state space model for monaural speech enhancement,”
Y. Du, X. Liu, and Y. Chua, · 2024
Closest in time.