Fetching the paper…
Reading the bibliography…
Diffusion model, as a new generative model which is very popular in image generation and audio synthesis, is rarely used in speech enhancement.
J. Garofolo, D. Graff, D. Paul, and D. Pallett, “Csr-i (wsj0) complete ldc93s6a,” in Web Download. Philadelphia: Linguistic Data Consortium , vol. 83, 1993
1993
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” vol. 2, 2001, pp. 749–752
2001
Earlier work this paper cites.
Y. Hu and P. C. Loizou, “Evaluation of objective quality measures for speech enhancement,” IEEE Transactions on audio, speech, and language processing , vol. 16, no. 1, pp. 229–238, 2007
2007
Earlier work this paper cites.
M. Abd El-Fattah, M. I. Dessouky, S. Diab, and F. Abd El-Samie, “Speech enhancement using an adaptive wiener filtering approach,” Progress In Electromagnetics Research M , vol. 4, pp. 167–184, 2008
2008
Earlier work this paper cites.
K. Paliwal, K. Wójcicki, and B. Shannon, “The importance of phase in speech enhancement,” Speech Communication , vol. 53, no. 4, pp. 465–494, 2011
2011
Earlier work this paper cites.
C. Veaux, J. Yamagishi, and S. King, “The voice bank corpus: Design, collection and data analysis of a large regional accent speech database,” in Proc. CASLRE 2013
2013
Earlier work this paper cites.
J. Thiemann, N. Ito, and E. Vincent, “The diverse environments multi-channel acoustic noise database (demand): A database of multichannel environmental noise recordings,” in Proceedings of Meetings on Acoustics , vol. 19, no. 1, 2013, p. 035081
2013
Earlier work this paper cites.
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. H. Soni, N. Shah, and H. A. Patil, “Time-frequency masking-based speech enhancement using generative adversarial network,” in Proc. ICASSP , 2018, pp. 5039–5043
2018
Earlier work this paper cites.
C. Donahue, B. Li, and R. Prabhavalkar, “Exploring speech enhancement with generative adversarial networks for robust speech recognition,” in Proc. ICASSP , 2018, pp. 5024–5028
2018
Earlier work this paper cites.
S. Leglaive, L. Girin, and R. Horaud, “A variance modeling framework based on variational autoencoders for speech enhancement,” in Proc. MLSP , 2018, pp. 1–6
2018
Cited alongside, same era.
Y. Bando, M. Mimura, K. Itoyama, K. Yoshii, and T. Kawahara, “Statistical speech enhancement based on probabilistic integration of variational autoencoder and non-negative matrix factorization,” in Proc. ICASSP , 2018, pp. 716–720
2018
Cited alongside, same era.
D. Baby and S. Verhulst, “Sergan: Speech enhancement using relativistic generative adversarial networks with gradient penalty,” in Proc. ICASSP , 2019, pp. 106–110
2019
Cited alongside, same era.
H. Zhang, I. Goodfellow, D. Metaxas, and A. Odena, “Self-attention generative adversarial networks,” in Proceedings of the 36th International Conference on Machine Learning , vol. 97, 2019, pp. 7354–7363
2019
Cited alongside, same era.
2020
Later among the works it cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems , vol. 33, 2020, pp. 6840–6851
2020
Later among the works it cites.
2020
Later among the works it cites.
H. Phan, I. V. McLoughlin, L. Pham, O. Y. Chén, P. Koch, M. De Vos, and A. Mertins, “Improving gans for speech enhancement,” IEEE Signal Processing Letters , vol. 27, pp. 1700–1704, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Leglaive, L. Girin, and R. Horaud, “Semi-supervised multichannel speech enhancement with variational autoencoders and non-negative matrix factorization,” in Proc. ICASSP , 2019, pp. 101–105
2019
Cited alongside, same era.
S. Leglaive, U. Şimşekli, A. Liutkus, L. Girin, and R. Horaud, “Speech enhancement with variational autoencoders and alpha-stable distributions,” in Proc. ICASSP , 2019, pp. 541–545
2019
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE/ACM Trans. Audio, speech, Lang. Process. , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Cited alongside, same era.
G. Liu, K. Gong, X. Liang, and Z. Chen, “Cp-gan: Context pyramid generative adversarial network for speech enhancement,” in Proc. ICASSP , 2020, pp. 6624–6628
2020
Cited alongside, same era.
A. A. Nugraha, K. Sekiguchi, and K. Yoshii, “A flow-based deep latent variable model for speech spectrogram modeling and enhancement,” IEEE/ACM Trans. Audio, speech, Lang. Process. , vol. 28, pp. 1104–1117, 2020
2020
Cited alongside, same era.
S. Leglaive, X. Alameda-Pineda, L. Girin, and R. Horaud, “A recurrent variational autoencoder for speech enhancement,” in Proc. ICASSP , 2020, pp. 371–375
2020
Cited alongside, same era.
2020
Cited alongside, same era.
T.-A. Hsieh, H.-M. Wang, X. Lu, and Y. Tsao, “Wavecrn: An efficient convolutional recurrent neural network for end-to-end speech enhancement,” IEEE Signal Processing Letters , vol. 27, pp. 2149–2153, 2020
2020
Later among the works it cites.
M. Strauss and B. Edler, “A flow-based neural network for time domain speech enhancement,” in ICASSP , 2021, pp. 5754–5758
2021
Later among the works it cites.
M. Sadeghi and X. Alameda-Pineda, “Mixture of inference networks for vae-based audio-visual speech enhancement,” IEEE Transactions on Signal Processing , vol. 69, pp. 1899–1909, 2021
2021
Later among the works it cites.
Y.-J. Lu, Y. Tsao, and S. Watanabe, “A study on speech enhancement based on diffusion probabilistic model,” in Proc. APSIPA ASC , 2021, pp. 659–666
2021
Later among the works it cites.
Y.-J. Lu, Z.-Q. Wang, S. Watanabe, A. Richard, C. Yu, and Y. Tsao, “Conditional Diffusion Probabilistic Model for Speech Enhancement,” in Proc. ICASSP , 2022
2022
Closest in time.
J. Whang, M. Delbracio, H. Talebi, C. Saharia, A. G. Dimakis, and P. Milanfar, “Deblurring via stochastic refinement,” in Proc. CVPR , June 2022, pp. 16 293–16 303
2022
Closest in time.
S.-W. Fu, C.-F. Liao, Y. Tsao, and S.-D. Lin, “MetricGAN: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” in Proceedings of the 36th International Conference on Machine Learning , vol. 97, 2019, pp. 2031–2041
2041
Closest in time.