Fetching the paper…
Reading the bibliography…
In this paper, we propose a novel score-base generative model for unconditional raw audio synthesis.
Reverse-time diffusion equation models, 1982
Brian D O Anderson · 1982
Earlier work this paper cites.
A connection between score matching and denoising autoencoders, 2011
Pascal Vincent · 2011
Earlier work this paper cites.
Conditional generative adversarial nets, 2014
Mehdi Mirza and Simon Osindero · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics, 2015
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation, 2015
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Wavenet: A generative model for raw audio, 2016
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Samplernn: An unconditional end-to-end neural audio generation model, 2017
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio · 2017
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer, 2017
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron Courville · 2017
Earlier work this paper cites.
Waveglow: A flow-based generative network for speech synthesis, 2018
Ryan Prenger, Rafael Valle, and Bryan Catanzaro · 2018
Earlier work this paper cites.
Speech commands: A dataset for limited-vocabulary speech recognition, 2018
Pete Warden · 2018
Earlier work this paper cites.
Melnet: A generative model for audio in the frequency domain, 2019
Sean Vasquez and Mike Lewis · 2019
Cited alongside, same era.
Gansynth: Adversarial neural audio synthesis, 2019
Jesse Engel, Kumar Krishna Agrawal, Shuo Chen, Ishaan Gulrajani, Chris Donahue, and Adam Roberts · 2019
Cited alongside, same era.
Adversarial audio synthesis, 2019
Chris Donahue, Julian McAuley, and Miller Puckette · 2019
Cited alongside, same era.
Neural drum machine: An interactive system for real-time synthesis of drum sounds
Cyran Aouameur, Philippe Esling, and Gaëtan Hadjeres · 2019
Cited alongside, same era.
Generating diverse high-fidelity images with vq-vae-2
Ali Razavi, Aaron van den Oord, and Oriol Vinyals · 2019
Cited alongside, same era.
Generative modeling by estimating gradients of the data distribution
Denoising diffusion probabilistic models, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Later among the works it cites.
Wavegrad: Estimating gradients for waveform generation, 2020
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J. Weiss, Mohammad Norouzi, and William Chan · 2020
Later among the works it cites.
Ganspace: Discovering interpretable gan controls, 2020
Erik Härkönen, Aaron Hertzmann, Jaakko Lehtinen, and Sylvain Paris · 2020
Later among the works it cites.
Fourier features let networks learn high frequency functions in low dimensional domains, 2020
Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng · 2020
Later among the works it cites.
Improved techniques for training score-based generative models, 2020
Yang Song and Stefano Ermon · 2020
Later among the works it cites.
Spectrogram inpainting for interactive generation of instrument sounds
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang Song and Stefano Ermon · 2019
Cited alongside, same era.
Conditioned-u-net: Introducing a control mechanism in the u-net for multiple source separations
Gabriel Meseguer-Brocal and G. Peeters · 2019
Cited alongside, same era.
Fréchet audio distance: A metric for evaluating music enhancement algorithms, 2019
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi · 2019
Cited alongside, same era.
A spectral energy distance for parallel speech synthesis, 2020
Alexey A. Gritsenko, Tim Salimans, Rianne van den Berg, Jasper Snoek, and Nal Kalchbrenner · 2020
Cited alongside, same era.
Drumgan: Synthesis of drum sounds with timbral feature conditioning using generative adversarial networks, 2020
J. Nistal, S. Lattner, and G. Richard · 2020
Cited alongside, same era.
Score-based generative modeling through stochastic differential equations, 2021a
Yang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole
Cited in the paper.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon
Cited in the paper.
Théis Bazin, Gaëtan Hadjeres, Philippe Esling, and Mikhail Malt · 2021
Closest in time.
Diffwave: A versatile diffusion model for audio synthesis, 2021
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2021
Closest in time.
Designing an encoder for stylegan image manipulation
Omer Tov, Yuval Alaluf, Yotam Nitzan, Or Patashnik, and Daniel Cohen-Or · 2021
Closest in time.
On maximum likelihood training of score-based generative models, 2021
Conor Durkan and Yang Song · 2021
Closest in time.
Improved denoising diffusion probabilistic models, 2021
Alex Nichol and Prafulla Dhariwal · 2021
Closest in time.