Fetching the paper…
Reading the bibliography…
In this work, we define a diffusion-based generative model capable of both music synthesis and source separation by learning the score of the joint probability density of sources sharing a context.
An Introduction to Harmonic Analysis
Yitzhak Katznelson · 2004
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Aapo Hyvärinen · 2005
Earlier work this paper cites.
Regularized estimation of image statistics by score matching
Durk P Kingma and Yann LeCun · 2010
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Pascal Vincent · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Generative adversarial source separation
Y Cem Subakan and Paris Smaragdis · 2018
Earlier work this paper cites.
Mmdenselstm: An efficient combination of convolutional and recurrent neural networks for audio source separation
Naoya Takahashi, Nabarun Goswami, and Yuki Mitsufuji · 2018
Earlier work this paper cites.
Music source separation in the waveform domain
Alexandre Défossez, Nicolas Usunier, Léon Bottou, and Francis Bach · 2019
Earlier work this paper cites.
Adversarial audio synthesis
Chris Donahue, Julian McAuley, and Miller Puckette · 2019
Earlier work this paper cites.
Universal sound separation
Ilya Kavalerov, Scott Wisdom, Hakan Erdogan, Brian Patton, Kevin Wilson, Jonathan Le Roux, and John R Hershey · 2019
Earlier work this paper cites.
Single-channel signal separation and deconvolution with generative adversarial networks
Qiuqiang Kong, Yong Xu, Wenwu Wang, Philip J. B. Jackson, and Mark D. Plumbley · 2019
Earlier work this paper cites.
End-to-end music source separation: Is it possible in the waveform domain?
Francesc Lluís, Jordi Pons, and Xavier Serra · 2019
Earlier work this paper cites.
Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation
Yi Luo and Nima Mesgarani · 2019
Earlier work this paper cites.
Cutting music source separation some Slakh: A dataset to study the impact of training data quality and quantity
Ethan Manilow, Gordon Wichern, Prem Seetharaman, and Jonathan Le Roux · 2019
Earlier work this paper cites.
Musdb18-hq - an uncompressed version of musdb18, August 2019
Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis, and Rachel Bittner · 2019
Earlier work this paper cites.
Sdr – half-baked or well done?
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R. Hershey · 2019
Earlier work this paper cites.
Protein design and variant prediction using autoregressive generative models
Jung-Eun Shin, Adam J Riesselman, Aaron W Kollasch, Conor McMahon, Elana Simon, Chris Sander, Aashish Manglik, Andrew C Kruse, and Debora S Marks · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon · 2019
Earlier work this paper cites.
Jukebox: A generative model for music
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Cited alongside, same era.
Source separation with deep generative priors
Vivek Jayaram and John Thickstun · 2020
Cited alongside, same era.
Unsupervised audio source separation using generative priors
Vivek Narayanaswamy, Jayaraman J. Thiagarajan, Rushil Anirudh, and Andreas Spanias · 2020
Cited alongside, same era.
Unsupervised sound separation using mixture invariant training
Scott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron Weiss, Kevin Wilson, and John Hershey · 2020
Cited alongside, same era.
Wavegrad: Estimating gradients for waveform generation
Nanxin Chen, Yu Zhang, Heiga Zen, Ron J Weiss, Mohammad Norouzi, and William Chan · 2021
Cited alongside, same era.
Lasaft: Latent source attentive frequency transformation for conditioned source separation
Woosung Choi, Minseok Kim, Jaehwa Chung, and Soonyoung Jung · 2021
A diffusion-inspired training strategy for singing voice extraction in the waveform domain
Genís Plaja-Roglans, Miron Marius, and Xavier Serra · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Universal speech enhancement with score-based diffusion
Joan Serrà, Santiago Pascual, Jordi Pons, R Oguz Araz, and Davide Scaini · 2022
Later among the works it cites.
Music source separation with generative flow
Ge Zhu, Jordan Darefsky, Fei Jiang, Anton Selitskiy, and Zhiyao Duan · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Hybrid spectrogram and waveform source separation
Alexandre Défossez · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Cited alongside, same era.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2021
Cited alongside, same era.
Parallel and flexible sampling from autoregressive models via langevin dynamics
Vivek Jayaram and John Thickstun · 2021
Cited alongside, same era.
Diffwave: A versatile diffusion model for audio synthesis
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro · 2021
Cited alongside, same era.
NU-Wave: A Diffusion Probabilistic Model for Neural Audio Upsampling
Junhyeok Lee and Seungu Han · 2021
Cited alongside, same era.
Musiclm: Generating music from text
Andrea Agostinelli, Timo I Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, et al · 2023
Closest in time.
Audiolm: a language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al · 2023
Closest in time.
Diffroll: Diffusion-based generative music transcription with unsupervised pretraining capability
Kin Wai Cheuk, Ryosuke Sawata, Toshimitsu Uesaka, Naoki Murata, Naoya Takahashi, Shusuke Takahashi, Dorien Herremans, and Yuki Mitsufuji · 2023
Closest in time.
Singsong: Generating musical accompaniments from singing
Chris Donahue, Antoine Caillon, Adam Roberts, Ethan Manilow, Philippe Esling, Andrea Agostinelli, Mauro Verzetti, Ian Simon, Olivier Pietquin, Neil Zeghidour, et al · 2023
Closest in time.
Audiogen: Textually guided audio generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi · 2023
Closest in time.
Audioldm: Text-to-audio generation with latent diffusion models
Haohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei, Xubo Liu, Danilo Mandic, Wenwu Wang, and Mark D Plumbley · 2023
Closest in time.
Separate and diffuse: Using a pretrained diffusion model for improving source separation
Shahar Lutati, Eliya Nachmani, and Lior Wolf · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Full-band general audio synthesis with score-based diffusion
Santiago Pascual, Gautam Bhattacharya, Chunghsin Yeh, Jordi Pons, and Joan Serrà · 2023
Closest in time.
Adversarial permutation invariant training for universal sound separation
Emilian Postolache, Jordi Pons, Santiago Pascual, and Joan Serrà · 2023
Closest in time.
Unsupervised vocal dereverberation with diffusion-based generative models
Koichi Saito, Naoki Murata, Toshimitsu Uesaka, Chieh-Hsin Lai, Yuhta Takida, Takao Fukui, and Yuki Mitsufuji · 2023
Closest in time.
Accelerating transformer inference for translation via parallel decoding
Andrea Santilli, Silvio Severino, Emilian Postolache, Valentino Maiorca, Michele Mancusi, Riccardo Marin, and Emanuele Rodola · 2023
Closest in time.
Diffiner: A Versatile Diffusion-based Generative Refiner for Speech Enhancement
Ryosuke Sawata, Naoki Murata, Yuhta Takida, Toshimitsu Uesaka, Takashi Shibuya, Shusuke Takahashi, and Yuki Mitsufuji · 2023
Closest in time.
Diffusion-based generative speech source separation
Robin Scheibler, Youna Ji, Soo-Whan Chung, Jaeuk Byun, Soyeon Choe, and Min-Seok Choi · 2023
Closest in time.
Moûsai: Text-to-music generation with long-context latent diffusion
Flavio Schneider, Zhijing Jin, and Bernhard Schölkopf · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Guided diffusion for inverse molecular design
Tomer Weiss, Eduardo Mayo Yanes, Sabyasachi Chakraborty, Luca Cosmo, Alex M. Bronstein, and Renana Gershoni-Poranne · 2023
Closest in time.
Diffsound: Discrete diffusion model for text-to-sound generation
Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu · 2023
Closest in time.
Conditioning and sampling in variational diffusion models for speech super-resolution
Chin-Yun Yu, Sung-Lin Yeh, György Fazekas, and Hao Tang · 2023
Closest in time.
Graph generation via spectral diffusion
Giorgia Minello, Alessandro Bicciato, Luca Rossi, Andrea Torsello, and Luca Cosmo · 2024
Closest in time.