Fetching the paper…
Reading the bibliography…
Recently, diffusion models (DMs) have been increasingly used in audio processing tasks, including speech super-resolution (SR), which aims to restore high-frequency content given low-resolution speech utterances.
“Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,”
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, · 2001
Earlier work this paper cites.
Digital Audio Resampling Home Page
J. O. Smith, · 2002
Earlier work this paper cites.
“Deep unsupervised learning using nonequilibrium thermodynamics,”
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
D. P. Kingma and J. Ba, · 2015
Earlier work this paper cites.
“CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),” Available: https://doi.org/10.7488/ds/2645 , 2019
J. Yamagishi, C. Veaux, and K. MacDonald, · 2019
Earlier work this paper cites.
“WaveGlow: A flow-based generative network for speech synthesis,”
R. Prenger, R. Valle, and B. Catanzaro, · 2019
Earlier work this paper cites.
“Denoising diffusion probabilistic models,”
J. Ho, A. Jain, and P. Abbeel, · 2020
Earlier work this paper cites.
“NU-Wave: A diffusion probabilistic model for neural audio upsampling,”
J. Lee and S. Han, · 2021
Earlier work this paper cites.
“Score-based generative modeling through stochastic differential equations,”
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, · 2021
Cited alongside, same era.
“WaveGrad: Estimating gradients for waveform generation,”
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, · 2021
Cited alongside, same era.
“DiffWave: A versatile diffusion model for audio synthesis,”
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, · 2021
Cited alongside, same era.
“Variational diffusion models,”
D. P. Kingma, T. Salimans, B. Poole, and J. Ho, · 2021
Cited alongside, same era.
“A study on speech enhancement based on diffusion probabilistic model,”
Y.-J. Lu, Y. Tsao, and S. Watanabe, · 2021
Cited alongside, same era.
“WSRGlow: A glow-based waveform generative model for audio super-resolution,”
K. Zhang, Y. Ren, C. Xu, and Z. Zhao, · 2021
“NU-Wave 2: A general neural audio upsampling model for various sampling rates,”
S. Han and J. Lee, · 2022
Closest in time.
“Denoising diffusion restoration models,”
B. Kawar, M. Elad, S. Ermon, and J. Song, · 2022
Closest in time.
“Repaint: Inpainting using denoising diffusion probabilistic models,”
A. Lugmayr, M. Danelljan, A. Romero, F. Yu, R. Timofte, and L. Van Gool, · 2022
Closest in time.
“Improving diffusion models for inverse problems using manifold constraints,”
H. Chung, B. Sim, D. Ryu, and J. C. Ye, · 2022
Closest in time.
“Solving inverse problems in medical imaging with score-based generative models,”
Y. Song, L. Shen, L. Xing, and S. Ermon, · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Towards robust speech super-resolution,”
H. Wang and D. Wang, · 2021
Cited alongside, same era.
“Speech enhancement with score-based generative models in the complex stft domain,”
S. Welker, J. Richter, and T. Gerkmann, · 2022
Closest in time.
“Neural vocoder is all you need for speech super-resolution,”
H. Liu, W. Choi, X. Liu, Q. Kong, Q. Tian, and D. Wang, · 2022
Closest in time.