Fetching the paper…
Reading the bibliography…
We propose a multi-dimensional structured state space (S4) approach to speech enhancement.
A. Tustin, “A method of analysing the behaviour of linear systems in terms of time series,”
1947
Earlier work this paper cites.
S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,”
1979
Earlier work this paper cites.
J. Lim and A. Oppenheim, “Enhancement and bandwidth compression of noisy speech,”
1979
Earlier work this paper cites.
M. Berouti, R. Schwartz, and J. Makhoul, “Enhancement of speech corrupted by acoustic noise,” in
1979
Earlier work this paper cites.
J. R. Bunch, “A note on the stable decompostion of skew-symmetric matrices,”
1982
Earlier work this paper cites.
K. Paliwal and A. Basu, “A speech enhancement method based on kalman filtering,” in
1987
Earlier work this paper cites.
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in
2001
Earlier work this paper cites.
V. Krishnan, S. Siniscalchi, D. Anderson, and M. Clements, “Noise robust aurora-2 speech recognition employing a codebook-constrained kalman filter preprocessor,” in
2006
Earlier work this paper cites.
D. Wang, “Time-frequency masking for speech separation and its potential for hearing aid design,”
2008
Earlier work this paper cites.
I. Markovsky, “Structured low-rank approximation and its applications,”
2008
Earlier work this paper cites.
R. Pascanu, T. Mikolov, and Y. Bengio, “On the difficulty of training recurrent neural networks,” in
2013
Earlier work this paper cites.
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “An experimental study on speech enhancement based on deep neural networks,”
2014
Earlier work this paper cites.
E. M. Grais, M. U. Sen, and H. Erdogan, “Deep neural networks for single channel source separation,” in
2014
Earlier work this paper cites.
Y. Wang, A. Narayanan, and D. Wang, “On training targets for supervised speech separation,”
2014
Earlier work this paper cites.
P.-S. Huang, M. Kim, M. Hasegawa-Johnson, and P. Smaragdis, “Deep learning for monaural speech separation,” in
2014
Earlier work this paper cites.
J. H. Hansen and T. Hasan, “Speaker recognition by machines and humans: A tutorial review,”
2015
Cited alongside, same era.
——, “A regression approach to speech enhancement based on deep neural networks,”
2015
Cited alongside, same era.
F. Weninger, H. Erdogan, S. Watanabe, E. Vincent, J. L. Roux, J. R. Hershey, and B. Schuller, “Speech enhancement with LSTM recurrent neural networks and its application to noise-robust ASR,” in
2015
Cited alongside, same era.
A. Kessy, A. Lewin, and K. Strimmer, “Optimal whitening and decorrelation,”
2015
Cited alongside, same era.
B. Wu, K. Li, F. Ge, Z. Huang, M. Yang, S. M. Siniscalchi, and C.-H. Lee, “An end-to-end deep learning approach to simultaneous speech dereverberation and acoustic modeling for robust speech recognition,”
2017
Cited alongside, same era.
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in
2020
Later among the works it cites.
Y. Hu, Y. Liu, S. Lv, M. Xing, S. Zhang, Y. Fu, J. Wu, B. Zhang, and L. Xie, “DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement,” in
2020
Later among the works it cites.
J. Qi, J. Du, S. M. Siniscalchi, X. Ma, and C.-H. Lee, “On mean absolute error for deep neural network based vector-to-vector regression,”
2020
Later among the works it cites.
S.-W. Fu, C. Yu, T.-A. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y. Tsao, “MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement,” in
2021
Later among the works it cites.
H. Fang, G. Carbajal, S. Wermter, and T. Gerkmann, “Variational autoencoder for speech enhancement with a noise-aware encoder,” in
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Pascual, A. Bonafonte, and J. Serrà, “Segan: Speech enhancement generative adversarial network,” in
2017
Cited alongside, same era.
C. Valentini-Botinhao, “Noisy speech database for training speech enhancement algorithms and tts models,” 2017
2017
Cited alongside, same era.
T. Gao, J. Du, L.-R. Dai, and C.-H. Lee, “Densely connected progressive learning for lstm-based speech enhancement,” in
2018
Cited alongside, same era.
C. Macartney and T. Weyde, “Improved speech enhancement with the wave-u-net,”
2018
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,”
2019
Cited alongside, same era.
S.-W. Fu, C.-F. Liao, Y. Tsao, and S.-D. Lin, “MetricGAN: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” in
2019
Cited alongside, same era.
R. Giri, U. Isik, and A. Krishnaswamy, “Attention wave-u-net for speech enhancement,” in
2019
Cited alongside, same era.
2021
Later among the works it cites.
A. Gu, I. Johnson, K. Goel, K. K. Saab, T. Dao, A. Rudra, and C. Re, “Combining recurrent, convolutional, and continuous-time models with linear state space layers,” in
2021
Later among the works it cites.
Y.-J. Lu, Z.-Q. Wang, S. Watanabe, A. Richard, C. Yu, and Y. Tsao, “Conditional diffusion probabilistic model for speech enhancement,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
H. Yen, F. G. Germain, G. Wichern, and J. L. Roux, “Cold diffusion for speech enhancement,”
2022
Later among the works it cites.
A. Gu, K. Goel, and C. Re, “Efficiently modeling long sequences with structured state spaces,” in
2022
Later among the works it cites.
K. Goel, A. Gu, C. Donahue, and C. Re, “It’s raw! Audio generation with state-space models,” in
2022
Later among the works it cites.
2022
Later among the works it cites.
E. Nguyen, K. Goel, A. Gu, G. Downs, P. Shah, T. Dao, S. Baccus, and C. Ré, “S4ND: Modeling images and videos as multidimensional signals with state spaces,” in
2022
Later among the works it cites.
S. Lv, Y. Fu, M. Xing, J. Sun, L. Xie, J. Huang, Y. Wang, and T. Yu, “S-dccrn: Super wide band dccrn with learnable complex feature for speech enhancement,” in
2022
Later among the works it cites.