Fetching the paper…
Reading the bibliography…
This paper explores the integration of model-based and data-driven approaches within the realm of neural speech and audio coding systems.
G. Fant, Acoustic theory of speech production: with calculations based on X-ray studies of Russian articulations . Walter de Gruyter, 1971, no. 2
1971
Earlier work this paper cites.
ISO/IEC 11172-3:1993, “Coding of moving pictures and associated audio for digital storage media at up to about 1.5 Mbit/s,” 1993
1993
Earlier work this paper cites.
T. Painter and A. Spanias, “Perceptual coding of digital audio,” Proceedings of the IEEE , vol. 88, no. 4, pp. 451–515, 2000
2000
Earlier work this paper cites.
B. Bessette et al. , “The adaptive multirate wideband speech codec (AMR-WB),” IEEE Transactions on Speech and Audio Processing , vol. 10, no. 8, pp. 620–636, 2002
2002
Earlier work this paper cites.
ISO/IEC 14496-3:2001/Amd 1:2003, “Information technology — coding of audio-visual objects — part 3: Audio — amendment 1: Bandwidth extension,” 2003
2003
Earlier work this paper cites.
ITU-T Recommendation G.722.2, “Wideband coding of speech at around 16 kbit/s using Adaptive Multi-Rate Wideband (AMR-WB),” 2003
2003
Earlier work this paper cites.
ISO (2006) ISO/IEC 13818-7:2006, “Information technology — Generic coding of moving pictures and associated audio information — Part 7: Advanced Audio Coding (AAC),” 2006
2006
Earlier work this paper cites.
ISO/IEC DIS 23003-3, “Information technology – MPEG audio technologies – part 3: Unified speech and audio coding,” 2011
2011
Earlier work this paper cites.
ISO/IEC 14496-3:2009/PDAM 3, “Transport of unified speech and audio coding (USAC),” 2011
2011
Earlier work this paper cites.
D. Rowe. (2011, http://www.tapr.org/pdf/DCC2011-Codec2-VK5DGR.pdf) Codec 2- open source speech coding at 2400 bits/s and below [online]
2011
Earlier work this paper cites.
J.-M. Valin, K. Vos, and T. B. Terriberry, “Definition of the Opus Audio Codec,” RFC 6716, Sep. 2012. [Online]. Available: https://www.rfc-editor.org/info/rfc6716
2012
Earlier work this paper cites.
Recommendation G.711.1, “Wideband embedded extension for ITU-T G.711 pulse code modulation,” 2012
2012
Earlier work this paper cites.
I. Goodfellow et al. , “Generative adversarial nets,” in Advances in neural information processing systems , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
K. Cho et al. , “Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation,” in Proc. of the Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2014
2014
Earlier work this paper cites.
ITU-R Recommendation BS 1534-3, “Method for the subjective assessment of intermediate quality levels of coding systems (MUSHRA),” 2015
2015
Cited alongside, same era.
A. van den Oord et al. , “WaveNet: A Generative Model for Raw Audio,” in Proc. 9th ISCA Workshop on Speech Synthesis Workshop (SSW 9) , 2016, p. 125
2016
Cited alongside, same era.
E. Agustsson et al. , “Soft-to-hard vector quantization for end-to-end learning compressible representations,” in Advances in Neural Information Processing Systems (NIPS) , 2017, pp. 1141–1151
2017
Cited alongside, same era.
S. Kankanahalli, “End-to-end optimized speech coding with deep neural networks,” in Proc. of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2018
2018
Cited alongside, same era.
W. B. Kleijn et al. , “WaveNet based low rate speech coding,” in Proc. of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2018, pp. 676–680
A. Biswas and D. Jia, “Audio codec enhancement with generative adversarial networks,” in Proc. of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2020, pp. 356–360
2020
Later among the works it cites.
K. Zhen, M. S. Lee, J. Sung, S. Beack, and M. Kim, “Psychoacoustic calibration of loss functions for efficient end-to-end neural audio coding,” IEEE Signal Processing Letters , vol. 27, pp. 2159–2163, 2020
2020
Later among the works it cites.
J. Byun, S. Shin, Y. Park, J. Sung, and S. Beack, “Development of a Psychoacoustic Loss Function for the Deep Neural Network (DNN)-Based Speech Coder,” in Proc. Interspeech , 2021, pp. 1694–1698
2021
Later among the works it cites.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Trans. Audio, Speech and Lang. Proc. , vol. 30, p. 495–507, jan 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
European Telecommunications Standards Institute (ETSI), “TR 103 590: Digital Enhanced Cordless Telecommunications (DECT); Study of Super Wideband Codec in DECT for narrowband, wideband and super-wideband audio communication including options of low delay audio connections,” 2018
2018
Cited alongside, same era.
N. Kalchbrenner et al. , “Efficient neural audio synthesis,” in Proc. of the International Conference on Machine Learning (ICML) , vol. 80, 2018, pp. 2410–2419
2018
Cited alongside, same era.
C. Garbacea, A. van den Oord, and Y. Li, “Low bit-rate speech coding with VQ-VAE and a WaveNet decoder,” in Proc. of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2019
2019
Cited alongside, same era.
J.-M. Valin and J. Skoglund, “A real-time wideband neural vocoder at 1.6 kb/s using LPCNet,” in Proc. Interspeech , 2019
2019
Cited alongside, same era.
Z. Zhao, H. Liu, and T. Fingscheidt, “Convolutional neural networks to enhance coded speech,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 4, pp. 663–678, 2019
2019
Cited alongside, same era.
K. Zhen, J. Sung, M. S. Lee, S. Beack, and M. Kim, “Cascaded cross-module residual learning towards lightweight end-to-end speech coding,” in Proc. Interspeech , 2019
2019
Cited alongside, same era.
3GPP TS 26.451, “Codec for Enhanced Voice Services (EVS); Voice Activity Detection (VAD),” 2020
2020
Cited alongside, same era.
S. Korse, N. Pia, K. Gupta, and G. Fuchs, “PostGAN: A gan-based post-processor to enhance the quality of coded speech,” in Proc. of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2022, pp. 831–835
2022
Later among the works it cites.
——, “Scalable and efficient neural speech coding: A hybrid design,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 12–25, 2022
2022
Later among the works it cites.
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” Transactions on Machine Learning Research , 2023
2023
Later among the works it cites.
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity audio compression with improved RVQGAN,” in Advances in Neural Information Processing Systems (NeurIPS) , 2023
2023
Later among the works it cites.
D. O’Shaughnessy, “Review of methods for coding of speech signals,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 8, 2023
2023
Later among the works it cites.
J.-M. Valin, J. Büthe, and A. Mustafa, “Low-bitrate redundancy coding of speech using a rate-distortion-optimized variational autoencoder,” in Proc. of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2023
2023
Later among the works it cites.
H. Yang, W. Lim, and M. Kim, “Neural feature predictor and discriminative residual coding for low-bitrate speech coding,” in Proc. of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2023
2023
Later among the works it cites.
X. Jiang, X. Peng, H. Xue, Y. Zhang, and Y. Lu, “Latent-domain predictive neural speech coding,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 2111–2123, 2023
2023
Later among the works it cites.
G. Davidson et al. , “High quality audio coding with MDCTNet,” in Proc. of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) , 2023
2023
Later among the works it cites.