Fetching the paper…
Reading the bibliography…
The recent advancement of end-to-end neural audio codecs enables compressing audio at very low bitrates while reconstructing the output audio with high fidelity.
H. Dudley, “Remaking speech,” in J. Acoust. Soc. Amer
1939
Earlier work this paper cites.
B. Atal and M. R. Schroeder, “Adaptive predictive coding of speech signals,” in Conf. Comm. and Proc
1967
Earlier work this paper cites.
J. Makhoul, “Linear prediction: A tutorial review,” Proc. IEEE
1975
Earlier work this paper cites.
R. Gray, “Vector quantization,” IEEE Assp Mag
1984
Earlier work this paper cites.
S. Morishima, H. Harashima and Y. Katayama, “Speech coding based on a multi-layer neural network,” in Proc. IEEE Int. Conf. Commun., Including Supercomm Tech. Sessions
1990
Earlier work this paper cites.
K. Brandenburg, J. Herre, J. D. Johnston, Y. Mahieux, and E. Schroeder, “ASPEC: Adaptive spectral entropy coding of high quality music signals,” in Proc. 90th Conv. Aud. Eng. Soc
1991
Earlier work this paper cites.
T. Q. Nguyen, “Near-perfect-reconstruction pseudo-QMF banks,” in IEEE Trans. Signal Process
1994
Earlier work this paper cites.
K. Brandenburg and G. Stoll, “ISO/MPEG-1 audio: A genereic standard for coding of high quality digital audio,” J. Audio Eng. Soc
1994
Earlier work this paper cites.
M. Bosi et al
1997
Earlier work this paper cites.
W. Dobson, J. Yang, K. Smart and F. Guo, “High quality low complexity scalable wavelet audio coding,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller, “Efficient backprop,” in Neural Networks: Tricks of the Trade
1998
Earlier work this paper cites.
Method for objective measurement of perceived audio quality
1998
Earlier work this paper cites.
T. Painter and A. Spanias, “Perceptual coding of digital audio,” Proc. IEEE
2000
Earlier work this paper cites.
V. K. Goyal, “Theoretical foundations of transform coding,” IEEE Signal Process. Mag
2001
Earlier work this paper cites.
P. Kabal, “An examination and interpretation of ITU-R BS.1387: Perceptual evaluation of audio quality,” McGill University, Tech. Rep., 2002
2002
Earlier work this paper cites.
J. Makinen, B. Bessette, S. Bruhn, P. Ojala, R. Salami and A. Taleb, “AMR-WB+: A new audio coding standard for 3rd generation mobile audio services,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process
2005
Earlier work this paper cites.
A. Vasuki and P. Vanathi, “A review of vector quantization techniques,” IEEE Potentials
2006
Earlier work this paper cites.
A. Spanias, T. Painter, and V. Atti, Audio signal processing and coding
2006
Earlier work this paper cites.
Information technology — Generic coding of moving pictures and associated audio information — Part 7: Advanced Audio Coding (AAC)
2006
Earlier work this paper cites.
M. Neuendorf et al
2009
Earlier work this paper cites.
I. Goodfellow et al
2014
Earlier work this paper cites.
M. Dietz et al
2015
Earlier work this paper cites.
A. Rämö and H. Toukomaa, “Subjective quality evaluation of the 3GPP EVS codec,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proc. Int. Conf. Mach. Learn
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J, Sun, “Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification,” in Proc. IEEE Int. Conf. Comput. Vis
2015
Earlier work this paper cites.
ITU-R, “Recommendation BS.1534-3: Method for the subjective assessment of intermediate quality level of coding systems,” Int. Telecommun. Union
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” 2016, arXiv:1607.06450
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit
2016
Cited alongside, same era.
D.-A. Clevert, T. Unterthiner, and S. Hochreiter, “Fast and accurate deep network learning by exponential linear units (ELUs),” in Proc. Int. Conf. Learn. Representations
2016
Cited alongside, same era.
T. Salimans and D. P. Kingma, “Weight normalization: A simple reparameterization to accelerate training of deep neural networks,” in Proc. Adv. Neural Inf. Process. Syst
2016
Cited alongside, same era.
V. Nagarajan and J. Z. Kolter, “Gradient descent GAN optimization is locally stable,” in Proc. Adv. Neural Inf. Process. Syst
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Proc. Adv. Neural Inf. Process. Syst
2020
Later among the works it cites.
A. Biswas and D. Jia, “Audio codec enhancement with generative adversarial networks,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process
2020
Later among the works it cites.
F. Mentzer, G. D. Toderici, M. Tschannen, and E. Agustsson, “High-fidelity generative image compression,” in Proc. Adv. Neural Inf. Process. Syst
2020
Later among the works it cites.
S. De and S. L. Smith, “Batch normalization biases residual blocks towards the identity function in deep networks,” in Proc. Adv. Neural Inf. Process. Syst
2020
Later among the works it cites.
M. Chinen, F. S. C. Lim, J. Skoglund, N. Gureev, F. O’Gorman, and A. Hines, “ViSQOL v3: An open source production ready objective speech and audio metric,” in Proc. 25th Int. Conf. Qual. Multimedia Experience
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
2017
Cited alongside, same era.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in Proc. Adv. Neural Inf. Process. Syst
2017
Cited alongside, same era.
I. Loshchilov and F. Hutter, “SGDR: Stochastic gradient descent with restarts,” in Proc. Int. Conf. Learn. Representations
2017
Cited alongside, same era.
N. Kalchbrenner et al
2018
Cited alongside, same era.
A. van den Oord et al
2018
Cited alongside, same era.
W. B. Kleijn et al
2018
Cited alongside, same era.
S. Kankanahalli, “End-to-end optimized speech coding with deep neural networks,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process
2018
Cited alongside, same era.
2020
Later among the works it cites.
Y. Li, M. Tagliasacchi, O. Rybakov, V. Ungureanu, and D. Roblek, “Real-time speech frequency bandwidth extension,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process
2021
Later among the works it cites.
W. Jang, D. Lim, J. Yoon, B. Kim, and J. Kim, “UnivNet, A neural vocoder with multi-resolution spectrogram discriminators for high-fidelity waveform generation,” in Proc. Interspeech
2021
Later among the works it cites.
B. Heo et al
2021
Later among the works it cites.
A. Brock, S. De, and S. L. Smith, “Characterizing signal propagation to close the performance gap in unnormalized ResNets,” in Proc. Int. Conf. Learn. Representations
2021
Later among the works it cites.
ONNX Runtime developers, “ONNX Runtime,” 2021. [Online] Available: https://onnxruntime.ai
2021
Later among the works it cites.
T. Bachlechner, B. P. Majumder, H. Mao, G. Cottrell, and J. McAuley, “ReZero is all you need: Fast convergence at large depth,” in Proc. 37th Conf. Uncertainty Artif. Intell
2021
Later among the works it cites.
H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve and H. Jégou, “Going deeper with image transformers,” in IEEE/CVF Int. Conf. Comput. Vis
2021
Later among the works it cites.
S. Korse, N. Pia, K. Gupta, and G. Fuchs, “PostGAN: A GAN-based post-processor to enhance the quality of coded speech,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process
2022
Later among the works it cites.
J. Lin, K. Kalgaonkar, Q. He and X. Lei, “Speech enhancement for low bit rate speech codec,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process
2022
Later among the works it cites.
H. Y. Kim, J. W. Yoon, W. I. Cho, and N. S. Kim, “Neurally optimized decoder for low bitrate speech codec,” in IEEE Signal Process. Lett
2022
Later among the works it cites.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Trans. Audio, Speech, Lang. Process
2022
Later among the works it cites.
H. Dubey et al
2022
Later among the works it cites.
T. Bak, J. Lee, H. Bae, J. Yang, J.-S. Bae, and Y.-S. Joo, “Avocodo: Generative adversarial network for artifact-free vocoder,” in Proc. AAAI Conf. Artif. Intell
2023
Later among the works it cites.
S.-g. Lee, W. Ping, B. Ginsburg, B. Catanzaro, and S. Yoon, “BigVGAN: A universal neural vocoder with large-scale training,” in Proc. Int. Conf. Learn. Representations
2023
Later among the works it cites.
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” Trans. Mach. Learn. Res
2023
Later among the works it cites.
Y. -C. Wu, I. D. Gebru, D. Marković and A. Richard, “Audiodec: An open-source streaming high-fidelity neural audio codec,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process
2023
Later among the works it cites.
2023
Later among the works it cites.
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity audio compression with improved RVQGAN,” in Proc. Adv. Neural Inf. Process. Syst
2023
Later among the works it cites.
J. Yu et al
2023
Later among the works it cites.
W. Xiao et al
2023
Later among the works it cites.
Z. Du, S. Zhang, K. Hu and S. Zheng, “FunCodec: A Fundamental, Reproducible and Integrable Open-Source Toolkit for Neural Speech Codec,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process
2024
Closest in time.
L. Xu, J. Wang, J. Zhang and X. Xie, ”LightCodec: A high fidelity neural audio codec with low computation complexity,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process
2024
Closest in time.