Fetching the paper…
Reading the bibliography…
Diffusion models have recently shown promising results for difficult enhancement tasks such as the conditional and unconditional restoration of natural images and audio signals.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,” in
2001
Earlier work this paper cites.
Y. Hu and P. C. Loizou, “Evaluation of objective quality measures for speech enhancement,”
2008
Earlier work this paper cites.
X. Lu, Y. Tsao, S. Matsuda, and C. Hori, “Speech enhancement based on deep denoising autoencoder,” in
2013
Earlier work this paper cites.
J. Li, L. Deng, Y. Gong, and R. Haeb-Umbach, “An overview of noise-robust automatic speech recognition,”
2014
Earlier work this paper cites.
Y. Xu, J. Du, L.-R. Dai, and C.-H. Lee, “An experimental study on speech enhancement based on deep neural networks,”
2014
Earlier work this paper cites.
Y. Wang, A. Narayanan, and D. Wang, “On training targets for supervised speech separation,”
2014
Earlier work this paper cites.
H. Erdogan, J. R. Hershey, S. Watanabe, and J. Le Roux, “Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks,” in
2015
Earlier work this paper cites.
Z. Chen, S. Watanabe, H. Erdogen, and J. R. Hershey, “Speech enhancement and recognition using multi-task learning of long short-term memory recurrent neural networks,” in
2015
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification,” in
2015
Earlier work this paper cites.
A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves
2016
Earlier work this paper cites.
D. S. Williamson, Y. Wang, and D. Wang, “Complex ratio masking for monaural speech separation,”
2016
Earlier work this paper cites.
D. Wang, “Deep learning reinvents the hearing aid,”
2017
Earlier work this paper cites.
H. Erdogan, J. R. Hershey, S. Watanabe, and J. Le Roux, “Deep recurrent networks for separation and recognition of single-channel speech in nonstationary background audio,” in
2017
Earlier work this paper cites.
S. Pascual, A. Bonafonte, and J. Serrà, “SEGAN: Speech enhancement generative adversarial network,” in
2017
Cited alongside, same era.
C. Valentini-Botinhao, “Noisy speech database for training speech enhancement algorithms and TTS models,” 2017,
2017
Cited alongside, same era.
K. Tan and D. Wang, “A convolutional recurrent neural network for real-time speech enhancement,” in
2018
Cited alongside, same era.
C. Trabelsi, O. Bilaniuk, Y. Zhang, D. Serdyuk, S. Subramanian, J. F. Santos
2018
Cited alongside, same era.
M. L. Vik, “Speech enhancement with a generative adversarial network,” Master’s thesis, NTNU, Trondheim, Norway, 2019
2019
Cited alongside, same era.
G. Carbajal, J. Richter, and T. Gerkmann, “Disentanglement learning for variational autoencoders applied to audio-visual speech enhancement,” in
2021
Later among the works it cites.
H. Fang, G. Carbajal, S. Wermter, and T. Gerkmann, “Variational autoencoder for speech enhancement with a noise-aware encoder,” in
2021
Later among the works it cites.
M. Strauss and B. Edler, “A flow-based neural network for time domain speech enhancement,” in
2021
Later among the works it cites.
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in
2021
Later among the works it cites.
Y.-J. Lu, Y. Tsao, and S. Watanabe, “A study on speech enhancement based on diffusion probabilistic model,” in
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
F. G. Germain, Q. Chen, and V. Koltun, “Speech denoising with deep feature losses,” in
2019
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time-frequency magnitude masking for speech separation,”
2019
Cited alongside, same era.
S.-W. Fu, C.-F. Liao, Y. Tsao, and S.-D. Lin, “MetricGAN: Generative adversarial networks based black-box metric scores optimization for speech enhancement,” in
2019
Cited alongside, same era.
H. Taherian, Z.-Q. Wang, J. Chang, and D. Wang, “Robust speaker recognition based on single-channel and multi-channel speech enhancement,”
2020
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in
2020
Cited alongside, same era.
Y. Hu, Y. Liu, S. Lv, M. Xing, S. Zhang, Y. Fu
2020
Cited alongside, same era.
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “DiffWave: A versatile diffusion model for audio synthesis,” in
2021
Later among the works it cites.
A. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” in
2021
Later among the works it cites.
X. Bie, S. Leglaive, X. Alameda-Pineda, and L. Girin, “Unsupervised speech enhancement using dynamical variational autoencoders,”
2022
Closest in time.
Y.-J. Lu, Z.-Q. Wang, S. Watanabe, A. Richard, C. Yu, and Y. Tsao, “Conditional diffusion probabilistic model for speech enhancement,” in
2022
Closest in time.
S. Welker, J. Richter, and T. Gerkmann, “Speech enhancement with score-based generative models in the complex STFT domain,” in
2022
Closest in time.
2022
Closest in time.
A. Bansal, E. Borgnia, H.-M. Chu, J. S. Li, H. Kazemi, F. Huang
2022
Closest in time.
S. Lv, Y. Fu, M. Xing, J. Sun, L. Xie, J. Huang
2022
Closest in time.