Fetching the paper…
Reading the bibliography…
Although recent mainstream waveform-domain end-to-end (E2E) neural audio codecs achieve impressive coded audio quality with a very low bitrate, the quality gap between the coded and natural audio is still significant.
“On the theory of the brownian motion,”
G. E. Uhlenbeck and L. S. Ornstein, · 1930
Earlier work this paper cites.
“Reverse-time diffusion equation models,”
B. D. Anderson, · 1982
Earlier work this paper cites.
Free Lossless Audio Codec
J. Coalson, · 2000
Earlier work this paper cites.
“Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,”
A. Rix, J. Beerends, M. Hollier, and A. Hekstra, · 2001
Earlier work this paper cites.
“The adaptive multirate wideband speech codec (AMR-WB),”
B. Bessette et al., · 2002
Earlier work this paper cites.
“MPEG-4 ALS: An emerging standard for lossless audio coding,”
Tilman Liebchen and Yuriy A Reznik, · 2004
Earlier work this paper cites.
“Estimation of non-normalized statistical models by score matching.,”
A. Hyvärinen and P. Dayan, · 2005
Earlier work this paper cites.
“A review of vector quantization techniques,”
A. Vasuki and P.T. Vanathi, · 2006
Earlier work this paper cites.
“An algorithm for intelligibility prediction of time–frequency weighted noisy speech,”
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, · 2011
Earlier work this paper cites.
“High-quality, low-delay music coding in the opus codec,”
J.-M. Valin, G. Maxwell, T. B. Terriberry, and K. Vos, · 2013
Earlier work this paper cites.
“Generative adversarial nets,”
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, · 2014
Earlier work this paper cites.
“Overview of the EVS codec architecture,”
M. Dietz et al., · 2015
Cited alongside, same era.
“CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,”
C. Veaux, J. Yamagishi, and K. MacDonald, · 2017
Cited alongside, same era.
“Noisy speech database for training speech enhancement algorithms and TTS models,”
C. Valentini-Botinhao, · 2017
Cited alongside, same era.
“End-to-end optimized speech coding with deep neural networks,”
S. Kankanahalli, · 2018
Cited alongside, same era.
“Low bit-rate speech coding with vq-vae and a wavenet decoder,”
C. Gârbacea and other, · 2019
Cited alongside, same era.
“Cascaded cross-module residual learning towards lightweight end-to-end speech coding,”
K. Zhen, J. Sung, M. S. Lee, S. Beack, and M. Kim, · 2019
Cited alongside, same era.
“SoundStream: An end-to-end neural audio codec,”
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, · 2021
Later among the works it cites.
“Score-based generative modeling through stochastic differential equations,”
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, · 2021
Later among the works it cites.
“DiffWave: A versatile diffusion model for audio synthesis,”
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, · 2021
Later among the works it cites.
“WaveGrad: Estimating gradients for waveform generation,”
N. Chen, Y. Zhang, H. Zen, R. J. Weiss, M. Norouzi, and W. Chan, · 2021
Later among the works it cites.
“High fidelity neural audio compression,”
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, · 2022
Later among the works it cites.
“Speech enhancement with score-based generative models in the complex STFT domain,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“MelGAN: generative adversarial networks for conditional waveform synthesis,”
K. Kumar et al., · 2019
Cited alongside, same era.
“Sdr–half-baked or well done?,”
J. L. Roux, S. Wisdom, H. Erdogan, and J. R. Hershey, · 2019
Cited alongside, same era.
“Denoising diffusion probabilistic models,”
J. Ho, A. Jain, and P. Abbeel, · 2020
Cited alongside, same era.
“HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,”
J. Kong, J. Kim, and J. Bae, · 2020
Cited alongside, same era.
“Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,”
R. Yamamoto, E. Song, and J.-M. Kim, · 2020
Cited alongside, same era.
S. Welker, J. Richter, and T. Gerkmann, · 2022
Later among the works it cites.
“Conditional diffusion probabilistic model for speech enhancement,”
Y.-J. Lu, Z.-Q. Wang, S. Watanabe, A. Richard, C. Yu, and Y. Tsao, · 2022
Later among the works it cites.
“AudioDec: An open-source streaming high-fidelity neural audio codec,”
Y.-C. Wu, I. D. Gebru, D. Marković, and A. Richard, · 2023
Later among the works it cites.
“Speech enhancement and dereverberation with diffusion-based generative models,”
J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, · 2023
Later among the works it cites.
“Neural speech phase prediction based on parallel estimation architecture and anti-wrapping losses,”
Y. Ai and Z.-H. Ling, · 2023
Later among the works it cites.