Fetching the paper…
Reading the bibliography…
Score-based generative models (SGMs) have recently shown impressive results for difficult generative tasks such as the unconditional and conditional generation of natural images and audio signals.
G. E. Uhlenbeck and L. S. Ornstein, “On the theory of the Brownian motion,” Physical review , vol. 36, no. 5, p. 823, 1930
1930
Earlier work this paper cites.
G. Parisi, “Correlation functions and computer simulations,” Nuclear Physics B , vol. 180, no. 3, pp. 378–384, 1981
1981
Earlier work this paper cites.
B. D. Anderson, “Reverse-time diffusion equation models,” Stochastic Processes and their Applications , vol. 12, no. 3, pp. 313–326, 1982
1982
Earlier work this paper cites.
ITU-T Rec. P.862.3, “Application guide for objective quality measurement based on Recommendations P.862, P.862.1 and P.862.2,” Int. Telecom. Union (ITU) , 2007
2007
Earlier work this paper cites.
T. Gerkmann and R. Martin, “Empirical distributions of DFT-domain speech coefficients based on estimated speech variances,” Int. Workshop on Acoustic Echo and Noise Control , 2010
2010
Earlier work this paper cites.
T. Chakraborty and M. Kearns, “Market making and mean reversion,” Proceedings of the 12th ACM conference on Electronic commerce , pp. 307–314, 2011
2011
Earlier work this paper cites.
P. Vincent, “A connection between score matching and denoising autoencoders,” Neural Computation , vol. 23, no. 7, pp. 1661–1674, 2011
2011
Earlier work this paper cites.
R. C. Hendriks, T. Gerkmann, and J. Jensen, DFT-Domain Based Single-Microphone Noise Reduction for Speech Enhancement: A Survey of the State-of-the-Art . Morgan & Claypool, 2013
2013
Earlier work this paper cites.
J. Thiemann, N. Ito, and E. Vincent, “The Diverse Environments Multi-channel Acoustic Noise Database (DEMAND): A database of multichannel environmental noise recordings,” The Journal of the Acoustical Society of America , vol. 133, no. 5, pp. 3591–3591, 2013
2013
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” Int. Conf. on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
A. Odena, V. Dumoulin, and C. Olah, “Deconvolution and checkerboard artifacts,” Distill , vol. 1, no. 10, p. e3, 2016
2016
Earlier work this paper cites.
C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Investigating RNN-based speech enhancement methods for noise-robust Text-to-Speech.” ISCA Speech Synthesis Workshop (SSW) , pp. 146–152, 2016
2016
Earlier work this paper cites.
S. Pascual, A. Bonafonte, and J. Serrà, “SEGAN: Speech enhancement generative adversarial network,” ISCA Interspeech , pp. 3642–3646, 2017
2017
Earlier work this paper cites.
T. Gerkmann and E. Vincent, “Spectral masking and filtering,” in Audio Source Separation and Speech Enhancement , E. Vincent, T. Virtanen, and S. Gannot, Eds. John Wiley & Sons, 2018
2018
Cited alongside, same era.
Y. Bando, M. Mimura, K. Itoyama, K. Yoshii, and T. Kawahara, “Statistical speech enhancement based on probabilistic integration of variational autoencoder and non-negative matrix factorization,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , pp. 716–720, 2018
2018
Cited alongside, same era.
S. Leglaive, L. Girin, and R. Horaud, “A variance modeling framework based on variational autoencoders for speech enhancement,” IEEE Int. Workshop on Machine Learning for Signal Proc. (MLSP) , pp. 1–6, 2018
2018
Cited alongside, same era.
D. Baby and S. Verhulst, “SERGAN: Speech enhancement using relativistic generative adversarial networks with gradient penalty,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , pp. 106–110, 2019
2019
M. Tancik, P. Srinivasan, B. Mildenhall, S. Fridovich-Keil, N. Raghavan, U. Singhal, R. Ramamoorthi, J. Barron, and R. Ng, “Fourier features let networks learn high frequency functions in low dimensional domains,” Advances in Neural Inf. Proc. Systems (NeurIPS) , vol. 33, pp. 7537–7547, 2020
2020
Later among the works it cites.
Y. Song and S. Ermon, “Improved techniques for training score-based generative models,” Advances in Neural Inf. Proc. Systems (NeurIPS) , vol. 33, pp. 12 438–12 448, 2020
2020
Later among the works it cites.
G. Carbajal, J. Richter, and T. Gerkmann, “Disentanglement learning for variational autoencoders applied to audio-visual speech enhancement,” IEEE Workshop on Applications of Signal Proc. to Audio and Acoustics (WASPAA) , pp. 126–130, 2021
2021
Later among the works it cites.
H. Fang, G. Carbajal, S. Wermter, and T. Gerkmann, “Variational autoencoder for speech enhancement with a noise-aware encoder,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , pp. 676–680, 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Y. Song and S. Ermon, “Generative modeling by estimating gradients of the data distribution,” Advances in Neural Inf. Proc. Systems (NeurIPS) , vol. 32, 2019
2019
Cited alongside, same era.
S. Särkkä and A. Solin, Applied Stochastic Differential Equations . Cambridge University Press, 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
J. Le Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, “SDR–half-baked or well done?” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , pp. 626–630, 2019
2019
Cited alongside, same era.
P. Wang, K. Tan, and D. L. Wang, “Bridging the gap between monaural speech enhancement and recognition with distortion-independent acoustic modeling,” IEEE Trans. on Audio, Speech, and Language Proc. (TASLP) , vol. 28, pp. 39–48, 2020
2020
Cited alongside, same era.
J. Richter, G. Carbajal, and T. Gerkmann, “Speech enhancement with stochastic temporal convolutional networks.” ISCA Interspeech , pp. 4516–4520, 2020
2020
Cited alongside, same era.
A. A. Nugraha, K. Sekiguchi, and K. Yoshii, “A flow-based deep latent variable model for speech spectrogram modeling and enhancement,” IEEE Trans. on Audio, Speech, and Language Proc. (TASLP) , vol. 28, pp. 1104–1117, 2020
2020
Cited alongside, same era.
Y. Bando, K. Sekiguchi, and K. Yoshii, “Adaptive neural speech enhancement with a denoising variational autoencoder.” ISCA Interspeech , pp. 2437–2441, 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
Y.-J. Lu, Y. Tsao, and S. Watanabe, “A study on speech enhancement based on diffusion probabilistic model,” Asia-Pacific Signal and Inf. Proc. Association Annual’Summit and Conf. (APSIPA ASC) , pp. 659–666, 2021
2021
Later among the works it cites.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” Int. Conf. on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
P. Dhariwal and A. Nichol, “Diffusion models beat GANs on image synthesis,” Advances in Neural Inf. Proc. Systems (NeurIPS) , vol. 34, 2021
2021
Later among the works it cites.
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “DiffWave: A versatile diffusion model for audio synthesis,” Int. Conf. on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, “WaveGrad: Estimating gradients for waveform generation,” Int. Conf. on Learning Representations (ICLR) , 2021
2021
Later among the works it cites.
S. Braun and I. Tashev, “A consolidated view of loss functions for supervised deep learning-based speech enhancement,” Int. Conf. on Telecom. and Signal Proc. (TSP) , pp. 72–76, 2021
2021
Later among the works it cites.
Y.-J. Lu, Z.-Q. Wang, S. Watanabe, A. Richard, C. Yu, and Y. Tsao, “Conditional Diffusion Probabilistic Model for Speech Enhancement,” IEEE Int. Conf. on Acoustics, Speech and Signal Proc. (ICASSP) , 2022
2022
Closest in time.
J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, “Speech enhancement and dereverberation with diffusion-based generative models,” (in preparation) , 2022
2022
Closest in time.