Fetching the paper…
Reading the bibliography…
Generative diffusion models have emerged as leading models in speech and image generation.
Generating diverse high-fidelity images with vq-vae-2
Razavi, A.; Oord, A. v. d.; and Vinyals, O. 2019 · 1906
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Song, Y.; and Ermon, S. 2019 · 1907
Earlier work this paper cites.
High fidelity speech synthesis with adversarial networks
Bińkowski, M.; Donahue, J.; Dieleman, S.; Clark, A.; Elsen, E.; Casagrande, N.; Cobo, L. C.; and Simonyan, K. 2019 · 1909
Earlier work this paper cites.
Mel-cepstral distance measure for objective speech quality assessment
Kubichek, R. 1993 · 1993
Earlier work this paper cites.
Long short-term memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs
Rix, A. W.; Beerends, J. G.; Hollier, M. P.; and Hekstra, A. P. 2001 · 2001
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Hyvärinen, A.; and Dayan, P. 2005 · 2005
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2006
Earlier work this paper cites.
WaveGrad: Estimating gradients for waveform generation
Chen, N.; Zhang, Y.; Zen, H.; Weiss, R. J.; Norouzi, M.; and Chan, W. 2020 · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and Fei-Fei, L. 2009 · 2009
Earlier work this paper cites.
Diffwave: A versatile diffusion model for audio synthesis
Kong, Z.; Ping, W.; Huang, J.; Zhao, K.; and Catanzaro, B. 2020 · 2009
Earlier work this paper cites.
Score-Based Generative Modeling through Stochastic Differential Equations
Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2020 · 2011
Earlier work this paper cites.
An Algorithm for Intelligibility Prediction of Time–Frequency Weighted Noisy Speech
Taal, C. H.; Hendriks, R. C.; Heusdens, R.; and Jensen, J. 2011 · 2011
Cited alongside, same era.
Learning Energy-Based Models by Diffusion Recovery Likelihood
Gao, R.; Song, Y.; Poole, B.; Wu, Y. N.; and Kingma, D. P. 2020 · 2012
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, D.; and Ba, J. 2014 · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K.; and Zisserman, A. 2014 · 2014
Cited alongside, same era.
Deep Learning Face Attributes in the Wild
Liu, Z.; Luo, P.; Wang, X.; and Tang, X. 2015 · 2015
Cited alongside, same era.
Donahue, C.; McAuley, J.; and Puckette, M. 2018 · 2018
Later among the works it cites.
Efficient neural audio synthesis
Kalchbrenner, N.; Elsen, E.; Simonyan, K.; Noury, S.; Casagrande, N.; Lockhart, E.; Stimberg, F.; Oord, A.; Dieleman, S.; and Kavukcuoglu, K. 2018 · 2018
Later among the works it cites.
Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation
Luo, Y.; and Mesgarani, N. 2019 · 2019
Later among the works it cites.
Denoising Diffusion Implicit Models
Jiaming Song, C. M.; and Ermon, S. 2020 · 2020
Later among the works it cites.
Denoising Diffusion Probabilistic Models
Jonathan Ho, P. A., Ajay Jain. 2020 · 2020
Later among the works it cites.
Analyzing and improving the image quality of stylegan
Karras, T.; Laine, S.; Aittala, M.; Hellsten, J.; Lehtinen, J.; and Aila, T. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; and Ganguli, S. 2015 · 2015
Cited alongside, same era.
LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop
Yu, F.; Zhang, Y.; Song, S.; Seff, A.; and Xiao, J. 2015 · 2015
Cited alongside, same era.
A kernelized Stein discrepancy for goodness-of-fit tests
Liu, Q.; Lee, J.; and Jordan, M. 2016 · 2016
Cited alongside, same era.
Wavenet: A generative model for raw audio
Oord, A. v. d.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A.; and Kavukcuoglu, K. 2016 · 2016
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017 · 2017
Cited alongside, same era.
The LJ Speech Dataset
Ito, K.; and Johnson, L. 2017 · 2017
Cited alongside, same era.
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
High-fidelity performance metrics for generative models in PyTorch
Obukhov, A.; Seitzer, M.; Wu, P.-W.; Zhydenko, S.; Kyl, J.; and Lin, E. Y.-J. 2020 · 2020
Later among the works it cites.
WaveGrad
Vovk, I. 2020 · 2020
Later among the works it cites.
Argmax Flows and Multinomial Diffusion: Towards Non-Autoregressive Language Models
Hoogeboom, E.; Nielsen, D.; Jaini, P.; Forré, P.; and Welling, M. 2021 · 2021
Closest in time.
Knowledge Distillation in Iterative Generative Models for Improved Sampling Speed
Luhman, E.; and Luhman, T. 2021 · 2021
Closest in time.
Autoregressive Denoising Diffusion Models for Multivariate Probabilistic Time Series Forecasting
Rasul, K.; Seward, C.; Schuster, I.; and Vollgraf, R. 2021 · 2021
Closest in time.
Denoising Diffusion Implicit Models
Song, J.; Meng, C.; and Ermon, S. 2021 · 2021
Closest in time.