Fetching the paper…
Reading the bibliography…
This paper presents CQT-Diff, a data-driven generative audio model that can, once trained, be used for solving various different audio inverse problems in a problem-agnostic setting.
“Constructing an invertible constant-Q transform with non-stationary Gabor frames,”
G. A. Velasco, N. Holighaus, M. Dörfler, and T. Grill, · 2011
Earlier work this paper cites.
“Audio declipping with social sparsity,”
K. Siedenburg, M. Kowalski, and M. Dörfler, · 2014
Earlier work this paper cites.
“Sparsity and cosparsity for audio declipping: A flexible non-convex approach,”
S. Kitić, N. Bertin, and R. Gribonval, · 2015
Earlier work this paper cites.
“Solving time-domain audio inverse problems using nonnegative tensor factorization,”
Ç. Bilen, A. Ozerov, and P. Pérez, · 2018
Earlier work this paper cites.
“FiLM: Visual reasoning with a general conditioning layer,”
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, · 2018
Earlier work this paper cites.
“TimbreTron: A WaveNet(CycleGAN(CQT(audio))) pipeline for musical timbre transfer,”
S. Huang, Q. Li, C. Anil, et al., · 2019
Earlier work this paper cites.
“Enabling factorized piano music modeling and generation with the MAESTRO dataset,”
C. Hawthorne, A. Stasyuk, A. Roberts, et al., · 2019
Earlier work this paper cites.
“Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms,”
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, · 2019
Earlier work this paper cites.
“GACELA: A generative adversarial context encoder for long audio inpainting of music,”
A. Marafioti, P. Majdak, N. Holighaus, and N. Perraudin, · 2020
Earlier work this paper cites.
“Denoising diffusion probabilistic models,”
J. Ho, A. Jain, and P. Abbeel, · 2020
Earlier work this paper cites.
“Data-driven harmonic filters for audio representation learning,”
M. Won, S. Chun, O. Nieto, and X. Serra, · 2020
Earlier work this paper cites.
“Analyzing and improving the image quality of StyleGAN,”
T. Karras, S. Laine, M. Aittala, et al., · 2020
Earlier work this paper cites.
“Fourier features let networks learn high frequency functions in low dimensional domains,”
M. Tancik, P. Srinivasan, B. Mildenhall, et al., · 2020
Cited alongside, same era.
“On filter generalization for music bandwidth extension using deep neural networks,”
S. Sulun and M. E. P. Davies, · 2020
Cited alongside, same era.
“Score-based generative modeling through stochastic differential equations,”
Y. Song, J. Sohl-Dickstein, D. P Kingma, et al., · 2021
Cited alongside, same era.
“DiffWave: A versatile diffusion model for audio synthesis,”
Z. Kong, W. Ping, J. Huang, et al., · 2021
Cited alongside, same era.
“CRASH: Raw audio score-based generative modeling for controllable high-resolution drum sound synthesis,”
S. Rouard and G. Hadjeres, · 2021
Cited alongside, same era.
“ILVR: Conditioning method for denoising diffusion probabilistic models,”
“Multi-instrument music synthesis with spectrogram diffusion,”
C. Hawthorne, I. Simon, A. Roberts, et al., · 2022
Closest in time.
“Denoising diffusion restoration models,”
B. Kawar, M. Elad, S. Ermon, and J. Song, · 2022
Closest in time.
“Improving diffusion models for inverse problems using manifold constraints,”
H. Chung, B. Sim, D. Ryu, and J. C. Ye, · 2022
Closest in time.
“It’s raw! Audio generation with state-space models,”
K. Goel, A. Gu, C. Donahue, and C. Re, · 2022
Closest in time.
“Realistic gramophone noise synthesis using a diffusion model,”
E. Moliner and V. Välimäki, · 2022
Closest in time.
“Speech enhancement and dereverberation with diffusion-based generative models,”
J. Richter, S. Welker, J. Lemercier, et al., · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Choi, S. Kim, Y. Jeong, et al., · 2021
Cited alongside, same era.
“Sequence-to-sequence piano transcription with transformers,”
C. Hawthorne, I. Simon, R. Swavely, et al., · 2021
Cited alongside, same era.
“Catch-a-waveform: Learning to generate audio from a single short example,”
G. Greshler, T. Shaham, and T. Michaeli, · 2021
Cited alongside, same era.
“BEHM-GAN: Bandwidth extension of historical music using generative adversarial networks,”
E. Moliner and V. Välimäki, · 2022
Cited alongside, same era.
“High-resolution image synthesis with latent diffusion models,”
R. Rombach, A. Blattmann, D. Lorenz, et al., · 2022
Cited alongside, same era.
“Video diffusion models,”
J. Ho, T. Salimans, A. Gritsenko, W. Chan, M. Norouzi, and D. J. Fleet, · 2022
Cited alongside, same era.
“Universal speech enhancement with score-based diffusion,”
J. Serrà, S. Pascual, J. Pons, R. O. Araz, and D. Scaini, · 2022
Cited alongside, same era.
Closest in time.
“Elucidating the design space of diffusion-based generative models,”
T. Karras, M. Aittala, T. Aila, and S. Laine, · 2022
Closest in time.
“NU-Wave 2: A general neural audio upsampling model for various sampling rates,”
S. Han and J. Lee, · 2022
Closest in time.
“Repaint: Inpainting using denoising diffusion probabilistic models,”
A. Lugmayr, M. Danelljan, A. Romero, et al., · 2022
Closest in time.
“Guided-TTS 2: A diffusion model for high-quality adaptive text-to-speech with untranscribed data,”
Sungwon Kim, Heeseung Kim, and Sungroh Yoon, · 2022
Closest in time.
“Diffusion posterior sampling for general noisy inverse problems,”
H. Chung, J. Kim, M. T. Mccann, et al., · 2023
Closest in time.