Fetching the paper…
Reading the bibliography…
Recently, score-based generative models have been successfully employed for the task of speech enhancement.
C. M. Bender and S. A. Orszag, Advanced Mathematical Methods for Scientists and Engineers . McGraw-Hill, 1978
1978
Earlier work this paper cites.
B. D. Anderson, “Reverse-time diffusion equation models,” Stochastic Processes and their Applications , vol. 12, no. 3, pp. 313–326, 1982
1982
Earlier work this paper cites.
U. G. Haussmann and E. Pardoux, “Time reversal of diffusions,” The Annals of Probability , pp. 1188–1205, 1986
1986
Earlier work this paper cites.
W. Rudin, Real and Complex Analysis , 3rd ed. McGraw-Hill, Inc., 1987
1987
Earlier work this paper cites.
I. Karatzas and S. E. Shreve, Brownian Motion and Stochastic Calculus , 2nd ed. Springer, 1996
1996
Earlier work this paper cites.
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, “Perceptual evaluation of speech quality (PESQ) - a new method for speech quality assessment of telephone networks and codecs,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , vol. 2, pp. 749–752, 2001
2001
Earlier work this paper cites.
T. Gerkmann and R. Martin, “Empirical distributions of DFT-domain speech coefficients based on estimated speech variances,” Int. Workshop on Acoustic Echo and Noise Control , 2010
2010
Earlier work this paper cites.
R. C. Hendriks, T. Gerkmann, and J. Jensen, DFT-domain based single-microphone noise reduction for speech enhancement: A survey of the state-of-the-art . Morgan & Claypool, 2013
2013
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” Int. Conf. on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The third ‘CHiME’ speech separation and recognition challenge: Dataset, task and baselines,” IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , pp. 504–511, 2015
2015
Cited alongside, same era.
J. Jensen and C. H. Taal, “An algorithm for predicting the intelligibility of speech masked by modulated noise maskers,” IEEE Trans. on Audio, Speech, and Language Proc. (TASLP) , vol. 24, no. 11, pp. 2009–2022, 2016
2016
Cited alongside, same era.
T. Gerkmann and E. Vincent, “Spectral masking and filtering,” in Audio Source Separation and Speech Enhancement , E. Vincent, T. Virtanen, and S. Gannot, Eds. John Wiley & Sons, 2018
2018
Cited alongside, same era.
D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE Trans. on Audio, Speech, and Language Proc. (TASLP) , vol. 26, no. 10, pp. 1702–1726, 2018
2018
Cited alongside, same era.
Y.-J. Lu, Y. Tsao, and S. Watanabe, “A study on speech enhancement based on diffusion probabilistic model,” IEEE Asia-Pacific Signal and Inf. Proc. Assoc. Annual Summit and Conf. (APSIPA ASC) , pp. 659–666, 2021
2021
Later among the works it cites.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” Int. Conf. on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, and M. Kudinov, “Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech,” Int. Conf. on Machine Learning (ICML) , 2021
2021
Later among the works it cites.
Y.-J. Lu, Z.-Q. Wang, S. Watanabe, A. Richard, C. Yu, and Y. Tsao, “Conditional diffusion probabilistic model for speech enhancement,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ITU-T Rec. P.863, “Perceptual objective listening quality prediction,” Int. Telecom. Union (ITU) , 2018. [Online]. Available: https://www.itu.int/rec/T-REC-P.863-201803-I/en
2018
Cited alongside, same era.
Y. Luo and N. Mesgarani, “Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,” IEEE Trans. on Audio, Speech, and Language Proc. (TASLP) , vol. 27, no. 8, pp. 1256–1266, 2019
2019
Cited alongside, same era.
S. Särkkä and A. Solin, Applied Stochastic Differential Equations , ser. Institute of Mathematical Statistics Textbooks. Cambridge University Press, 2019, no. 10
2019
Cited alongside, same era.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR–half-baked or well done?” in IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2019, pp. 626–630
2019
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in Neural Inf. Proc. Systems (NeurIPS) , vol. 33, pp. 6840–6851, 2020
2020
Cited alongside, same era.
J. S. Garofolo, D. Graff, D. Paul, and D. Pallett, “CSR-I (WSJ0) Complete.” [Online]. Available: https://catalog.ldc.upenn.edu/LDC93S6A
Cited in the paper.
S. Welker, J. Richter, and T. Gerkmann, “Speech enhancement with score-based generative models in the complex STFT domain,” Interspeech , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
V. Popov, I. Vovk, V. Gogoryan, T. Sadekova, M. S. Kudinov, and J. Wei, “Diffusion-based voice conversion with fast maximum likelihood sampling scheme,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.