Fetching the paper…
Reading the bibliography…
In this work, we build upon our previous publication and use diffusion-based generative models for speech enhancement.
J. R. Dormand and P. J. Prince, “A family of embedded Runge-Kutta formulae,” Journal of Computational and Applied Mathematics , vol. 6, pp. 19–26, 1980
1980
Earlier work this paper cites.
B. D. Anderson, “Reverse-time diffusion equation models,” Stochastic Processes and their Applications , vol. 12, no. 3, pp. 313–326, 1982
1982
Earlier work this paper cites.
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (PESQ) - a new method for speech quality assessment of telephone networks and codecs,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , vol. 2, pp. 749–752, 2001
2001
Earlier work this paper cites.
ITU-T Rec. P.835, “Subjective test methodology for evaluating speech communication systems that include noise suppression algorithm,” Int. Telecom. Union (ITU) , 2003. [Online]. Available: https://www.itu.int/rec/T-REC-P.835-200311-I/en
2003
Earlier work this paper cites.
A. Hyvärinen and P. Dayan, “Estimation of non-normalized statistical models by score matching.” Journal of Machine Learning Research , vol. 6, no. 4, 2005
2005
Earlier work this paper cites.
C. H. You, S. N. Koh, and S. Rahardja, “/spl beta/-order MMSE spectral amplitude estimation for speech enhancement,” IEEE Trans. on Audio, Speech, and Language Proc. (TASLP) , vol. 13, no. 4, pp. 475–486, 2005
2005
Earlier work this paper cites.
M. Lincoln, I. McCowan, J. Vepa, and H. K. Maganti, “The multi-channel wall street journal audio visual corpus (MC-WSJ-AV): Specification and initial experiments,” IEEE Workshop on Automatic Speech Recognition and Understanding , pp. 357–362, 2005
2005
Earlier work this paper cites.
ITU-T Rec. P.862.3, “Application guide for objective quality measurement based on Recommendations P.862, P.862.1 and P.862.2,” Int. Telecom. Union (ITU) , 2007. [Online]. Available: https://www.itu.int/rec/T-REC-P.862.3/en
2007
Earlier work this paper cites.
T. Gerkmann and R. Martin, “Empirical distributions of DFT-domain speech coefficients based on estimated speech variances,” Int. Workshop on Acoustic Echo and Noise Control , 2010
2010
Earlier work this paper cites.
C. Breithaupt and R. Martin, “Analysis of the decision-directed snr estimator for speech enhancement with respect to low-snr and transient conditions,” IEEE Trans. on Audio, Speech, and Language Proc. (TASLP) , vol. 19, no. 2, pp. 277–289, 2010
2010
Earlier work this paper cites.
P. Vincent, “A connection between score matching and denoising autoencoders,” Neural Computation , vol. 23, no. 7, pp. 1661–1674, 2011
2011
Earlier work this paper cites.
R. C. Hendriks, T. Gerkmann, and J. Jensen, DFT-domain based single-microphone noise reduction for speech enhancement: A survey of the state-of-the-art . Morgan & Claypool, 2013
2013
Earlier work this paper cites.
J. Thiemann, N. Ito, and E. Vincent, “The Diverse Environments Multi-channel Acoustic Noise Database (DEMAND): A database of multichannel environmental noise recordings,” The Journal of the Acoustical Society of America , vol. 133, no. 5, pp. 3591–3591, 2013
2013
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational Bayes,” Int. Conf. on Learning Representations (ICLR) , 2014
2014
Earlier work this paper cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Advances in Neural Inf. Proc. Systems (NeurIPS) , vol. 27, 2014
2014
Earlier work this paper cites.
ITU-R Rec. BS.1534-3, “Method for the subjective assessment of intermediate quality level of audio systems,” Int. Telecom. Union (ITU) , 2014. [Online]. Available: https://www.itu.int/rec/R-REC-BS.1534
2014
Earlier work this paper cites.
D. S. Williamson, Y. Wang, and D. Wang, “Complex ratio masking for monaural speech separation,” IEEE Trans. on Audio, Speech, and Language Proc. (TASLP) , vol. 24, no. 3, pp. 483–492, 2015
2015
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” Int. Conf. on Machine Learning (ICML) , pp. 2256–2265, 2015
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” Int. Conf. on Medical image computing and computer-assisted intervention , pp. 234–241, 2015
2015
Earlier work this paper cites.
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The third ‘CHiME’ speech separation and recognition challenge: Dataset, task and baselines,” IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , pp. 504–511, 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” Int. Conf. on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Investigating RNN-based speech enhancement methods for noise-robust text-to-speech,” ISCA Speech Synthesis Workshop (SSW) , pp. 146–152, 2016
2016
Earlier work this paper cites.
J. Jensen and C. H. Taal, “An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,” IEEE Trans. on Audio, Speech, and Language Proc. (TASLP) , vol. 24, no. 11, pp. 2009–2022, 2016
2016
Earlier work this paper cites.
S.-W. Fu, T.-Y. Hu, Y. Tsao, and X. Lu, “Complex spectrogram enhancement by convolutional neural network with multi-metrics learning,” IEEE Int. Workshop on Machine Learning for Signal Proc. (MLSP) , pp. 1–6, 2017
2017
Earlier work this paper cites.
S.-W. Fu, Y. Tsao, X. Lu, and H. Kawai, “Raw waveform-based speech enhancement by fully convolutional networks,” IEEE Asia-Pacific Signal and Inf. Proc. Assoc. Annual Summit and Conf. (APSIPA ASC) , 2017
2017
Earlier work this paper cites.
S. Pascual, A. Bonafonte, and J. Serrà, “SEGAN: Speech enhancement generative adversarial network,” ISCA Interspeech , pp. 3642–3646, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in Neural Inf. Proc. Systems (NeurIPS) , vol. 30, 2017
2017
Earlier work this paper cites.
T. Gerkmann and E. Vincent, “Spectral masking and filtering,” in Audio Source Separation and Speech Enhancement , E. Vincent, T. Virtanen, and S. Gannot, Eds. John Wiley & Sons, 2018
2018
Earlier work this paper cites.
D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE Trans. on Audio, Speech, and Language Proc. (TASLP) , vol. 26, no. 10, pp. 1702–1726, 2018
2018
Cited alongside, same era.
Y. Bando, M. Mimura, K. Itoyama, K. Yoshii, and T. Kawahara, “Statistical speech enhancement based on probabilistic integration of variational autoencoder and non-negative matrix factorization,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , pp. 716–720, 2018
2018
Cited alongside, same era.
S. Leglaive, L. Girin, and R. Horaud, “A variance modeling framework based on variational autoencoders for speech enhancement,” IEEE Int. Workshop on Machine Learning for Signal Proc. (MLSP) , pp. 1–6, 2018
2018
Cited alongside, same era.
A. Brock, J. Donahue, and K. Simonyan, “Large scale GAN training for high fidelity natural image synthesis,” Int. Conf. on Learning Representations (ICLR) , 2018
2018
Cited alongside, same era.
——, “Disentanglement learning for variational autoencoders applied to audio-visual speech enhancement,” IEEE Workshop on Applications of Signal Proc. to Audio and Acoustics (WASPAA) , pp. 126–130, 2021
2021
Later among the works it cites.
H. Fang, G. Carbajal, S. Wermter, and T. Gerkmann, “Variational autoencoder for speech enhancement with a noise-aware encoder,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , pp. 676–680, 2021
2021
Later among the works it cites.
Y.-J. Lu, Y. Tsao, and S. Watanabe, “A study on speech enhancement based on diffusion probabilistic model,” IEEE Asia-Pacific Signal and Inf. Proc. Assoc. Annual Summit and Conf. (APSIPA ASC) , pp. 659–666, 2021
2021
Later among the works it cites.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” Int. Conf. on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Wu and K. He, “Group normalization,” Proc. of the European conference on computer vision (ECCV) , pp. 3–19, 2018
2018
Cited alongside, same era.
R. Scheibler, E. Bezzam, and I. Dokmanic, “Pyroomacoustics: A python package for audio room simulation and array processing algorithms,” in IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2018
2018
Cited alongside, same era.
ITU-T Rec. P.863, “Perceptual objective listening quality prediction,” Int. Telecom. Union (ITU) , 2018. [Online]. Available: https://www.itu.int/rec/T-REC-P.863-201803-I/en
2018
Cited alongside, same era.
M. Schoeffler, S. Bartoschek, F.-R. Stöter, M. Roess, S. Westphal, B. Edler, and J. Herre, “webmushra—a comprehensive framework for web-based listening tests,” Journal of Open Research Software , vol. 6, no. 1, 2018
2018
Cited alongside, same era.
E. Aksan and O. Hilliges, “Stcn: Stochastic temporal convolutional networks,” Int. Conf. on Learning Representations (ICLR) , 2018
2018
Cited alongside, same era.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” Advances in Neural Inf. Proc. Systems (NeurIPS) , vol. 32, 2019
2019
Cited alongside, same era.
S. Särkkä and A. Solin, Applied Stochastic Differential Equations . Cambridge University Press, 2019
2019
Cited alongside, same era.
R. Zhang, “Making convolutional networks shift-invariant again,” Int. Conf. on Machine Learning (ICML) , pp. 7324–7334, 2019
2019
Cited alongside, same era.
S. Braun and I. Tashev, “A consolidated view of loss functions for supervised deep learning-based speech enhancement,” Int. Conf. on Telecom. and Signal Proc. (TSP) , pp. 72–76, 2021
2021
Later among the works it cites.
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, “WaveGrad: Estimating gradients for waveform generation,” Int. Conf. on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
P. Dhariwal and A. Nichol, “Diffusion models beat GANs on image synthesis,” Advances in Neural Inf. Proc. Systems (NeurIPS) , vol. 34, 2021
2021
Later among the works it cites.
C.-W. Huang, J. H. Lim, and A. C. Courville, “A variational perspective on diffusion-based generative models and score matching,” Advances in Neural Inf. Proc. Systems (NeurIPS) , vol. 34, 2021
2021
Later among the works it cites.
C. K. Reddy, V. Gopal, and R. Cutler, “DNSMOS: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , pp. 6493–6497, 2021
2021
Later among the works it cites.
ITU-T Rec. P.808, “Subjective evaluation of speech quality with a crowdsourcing approach,” Int. Telecom. Union (ITU) , 2021. [Online]. Available: https://www.itu.int/rec/T-REC-P.808-202106-I/en
2021
Later among the works it cites.
L. Girin, S. Leglaive, X. Bie, J. Diard, T. Hueber, and X. Alameda-Pineda, “Dynamical variational autoencoders: A comprehensive review,” Foundations and Trends in Machine Learning , vol. 15, no. 1-2, pp. 1–175, 2021
2021
Later among the works it cites.
S.-W. Fu, C. Yu, T.-A. Hsieh, P. Plantinga, M. Ravanelli, X. Lu, and Y. Tsao, “MetricGAN+: An improved version of MetricGAN for speech enhancement,” ISCA Interspeech , 2021
2021
Later among the works it cites.
D. Watson, W. Chan, J. Ho, and M. Norouzi, “Learning fast samplers for diffusion models by differentiating through sample quality,” Int. Conf. on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
X. Bie, S. Leglaive, X. Alameda-Pineda, and L. Girin, “Unsupervised speech enhancement using dynamical variational autoencoders,” IEEE Trans. on Audio, Speech, and Language Proc. (TASLP) , vol. 30, pp. 2993–3007, 2022
2022
Closest in time.
Y.-J. Lu, Z.-Q. Wang, S. Watanabe, A. Richard, C. Yu, and Y. Tsao, “Conditional diffusion probabilistic model for speech enhancement,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , pp. 7402–7406, 2022
2022
Closest in time.
S. Welker, J. Richter, and T. Gerkmann, “Speech enhancement with score-based generative models in the complex STFT domain,” ISCA Interspeech , pp. 2928–2932, 2022
2022
Closest in time.
2022
Closest in time.
Y. Koizumi, H. Zen, K. Yatabe, N. Chen, and M. Bacchiani, “SpecGrad: Diffusion probabilistic model based neural vocoder with adaptive noise spectral shaping,” ISCA Interspeech , pp. 803–807, 2022
2022
Closest in time.
T. Karras, M. Aittala, T. Aila, and S. Laine, “Elucidating the design space of diffusion-based generative models,” Advances in Neural Inf. Proc. Systems (NeurIPS) , vol. 35, 2022
2022
Closest in time.
C. K. Reddy, V. Gopal, and R. Cutler, “DNSMOS P.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2022
2022
Closest in time.
A. Li, C. Zheng, L. Zhang, and X. Li, “Glance and gaze: A collaborative learning framework for single-channel speech enhancement,” Applied Acoustics , vol. 187, p. 108499, 2022
2022
Closest in time.
G. Yu, A. Li, H. Wang, Y. Wang, Y. Ke, and C. Zheng, “DBT-Net: Dual-branch federative magnitude and phase estimation with attention-in-attention transformer for monaural speech enhancement,” IEEE Trans. on Audio, Speech, and Language Proc. (TASLP) , vol. 30, pp. 2629–2644, 2022
2022
Closest in time.
R. Cao, S. Abdulatif, and B. Yang, “CMGAN: Conformer-based metric GAN for speech enhancement,” in ISCA Interspeech , 2022, pp. 936–940
2022
Closest in time.
S.-W. Fu, C. Yu, K.-H. Hung, M. Ravanelli, and Y. Tsao, “MetricGAN-U: Unsupervised speech enhancement/dereverberation based only on noisy/reverberated speech,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , pp. 7412–7416, 2022
2022
Closest in time.
T. Peer and T. Gerkmann, “Phase-aware deep speech enhancement: It’s all about the frame length,” JASA Express Letters , vol. 2, no. 10, p. 104802, 2022
2022
Closest in time.
J.-M. Lemercier, J. Richter, S. Welker, and T. Gerkmann, “Analysing diffusion-based generative approaches versus discriminative approaches for speech restoration,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Closest in time.
P. Andreev, A. Alanov, O. Ivanov, and D. Vetrov, “HiFi++: a unified framework for bandwidth extension and speech enhancement,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2023
2023
Closest in time.