Fetching the paper…
Reading the bibliography…
Audio diffusion models can synthesize a wide variety of sounds.
Stochastic differential equations
P. E. Kloeden, E. Platen, P. E. Kloeden, and E. Platen · 1992
Earlier work this paper cites.
Analysis of the decision-directed snr estimator for speech enhancement with respect to low-snr and transient conditions
C. Breithaupt and R. Martin · 2010
Earlier work this paper cites.
FaceNet: A unified embedding for face recognition and clustering
F. Schroff, D. Kalenichenko, and J. Philbin · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
WaveNet: A generative model for raw audio
A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
SampleRNN: An unconditional end-to-end neural audio generation model
S. Mehri, K. Kumar, I. Gulrajani, R. Kumar, S. Jain, J. Sotelo, A. Courville, and Y. Bengio · 2017
Earlier work this paper cites.
Fréchet audio distance: A metric for evaluating music enhancement algorithms
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi · 2018
Earlier work this paper cites.
Speech commands: A dataset for limited-vocabulary speech recognition
P. Warden · 2018
Earlier work this paper cites.
Adversarial audio synthesis
C. Donahue, J. McAuley, and M. Puckette · 2019
Earlier work this paper cites.
Low bit-rate speech coding with vq-vae and a wavenet decoder
C. Gârbacea, A. van den Oord, Y. Li, F. S. Lim, A. Luebs, O. Vinyals, and T. C. Walters · 2019
Earlier work this paper cites.
Sound effect synthesis
D. Moffat, R. Selfridge, and J. D. Reiss · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
PANNs: Large-scale pretrained audio neural networks for audio pattern recognition
Q. Kong, Y. Cao, T. Iqbal, Y. Wang, W. Wang, and M. D. Plumbley · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2020
Earlier work this paper cites.
DiffWave: A versatile diffusion model for audio synthesis
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro · 2021
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2021
Cited alongside, same era.
It’s raw! audio generation with state-space models
K. Goel, A. Gu, C. Donahue, and C. Ré · 2022
Cited alongside, same era.
Multi-instrument music synthesis with spectrogram diffusion
C. Hawthorne, I. Simon, A. Roberts, N. Zeghidour, J. Gardner, E. Manilow, and J. Engel · 2022
Cited alongside, same era.
Classifier-free diffusion guidance, 2022
J. Ho and T. Salimans · 2022
Cited alongside, same era.
Masked autoencoders that listen
P.-Y. Huang, H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, and C. Feichtenhofer · 2022
Cited alongside, same era.
Elucidating the design space of diffusion-based generative models
T. Karras, M. Aittala, T. Aila, and S. Laine · 2022
Cited alongside, same era.
AudioLDM: Text-to-audio generation with latent diffusion models
H. Liu, Z. Chen, Y. Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley · 2023
Closest in time.
Simple pooling front-ends for efficient audio classification
X. Liu, H. Liu, Q. Kong, X. Mei, M. D. Plumbley, and W. Wang · 2023
Closest in time.
A demand-driven perspective on generative audio ai
S. Oh, M. Kang, H. Moon, K. Choi, and B. S. Chon · 2023
Closest in time.
Full-band general audio synthesis with score-based diffusion
S. Pascual, G. Bhattacharya, C. Yeh, J. Pons, and J. Serrà · 2023
Closest in time.
Speech enhancement and dereverberation with diffusion-based generative models
J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann · 2023
Closest in time.
Class-conditioned latent diffusion model for dcase 2023 foley sound synthesis challenge
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
DPM-Solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps
C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu · 2022
Cited alongside, same era.
DPM-Solver++: Fast solver for guided sampling of diffusion probabilistic models
C. Lu, Y. Zhou, F. Bao, J. Chen, C. Li, and J. Zhu · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Cited alongside, same era.
GAN you hear me? reclaiming unconditional speech synthesis from diffusion models
M. Baas and H. Kamper · 2023
Cited alongside, same era.
Foley sound synthesis at the dcase 2023 challenge
K. Choi, J. Im, L. Heller, B. McFee, K. Imoto, Y. Okamoto, M. Lagrange, and S. Takamichi · 2023
Cited alongside, same era.
Foley sound synthesis based on gan using contrastive learning without label information
H. C. Chung, Y. Lee, and J. H. Jung · 2023
Cited alongside, same era.
R. Scheibler, T. Hasumi, Y. Fujita, T. Komatsu, R. Yamamoto, and K. Tachibana · 2023
Closest in time.
Diffusion art or digital forgery? investigating data replication in diffusion models
G. Somepalli, V. Singla, M. Goldblum, J. Geiping, and T. Goldstein · 2023
Closest in time.
Understanding and mitigating copying in diffusion models
G. Somepalli, V. Singla, M. Goldblum, J. Geiping, and T. Goldstein · 2023
Closest in time.
AUDIT: Audio editing by following instructions with latent diffusion models
Y. Wang, Z. Ju, X. Tan, L. He, Z. Wu, J. Bian, and S. Zhao · 2023
Closest in time.
Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov · 2023
Closest in time.
Diffsound: Discrete diffusion model for text-to-sound generation
D. Yang, J. Yu, H. Wang, W. Wang, C. Weng, Y. Zou, and D. Yu · 2023
Closest in time.
Latent diffusion model based foley sound generation system for dcase challenge 2023 task 7
Y. Yuan, H. Liu, X. Liu, X. Kang, M. D. Plumbley, and W. Wang · 2023
Closest in time.
Fast sampling of diffusion models with exponential integrator
Q. Zhang and Y. Chen · 2023
Closest in time.
UniPC: A unified predictor-corrector framework for fast sampling of diffusion models
W. Zhao, L. Bai, Y. Rao, J. Zhou, and J. Lu · 2023
Closest in time.