Fetching the paper…
Reading the bibliography…
This paper introduces a novel neural audio codec targeting high waveform sampling rates and low bitrates named APCodec, which seamlessly integrates the strengths of parametric codecs and waveform codecs.
H. S. Black and J. Edson, “Pulse code modulation,” Transactions of the American Institute of Electrical Engineers , vol. 66, no. 1, pp. 895–899, 1947
1947
Earlier work this paper cites.
T. Tremain, “Linear predictive coding systems,” in Proc. ICASSP , vol. 1, 1976, pp. 474–478
1976
Earlier work this paper cites.
P. Kroon, E. Deprettere, and R. Sluyter, “Regular-pulse excitation–a novel approach to effective and efficient multipulse coding of speech,” IEEE transactions on acoustics, speech, and signal processing , vol. 34, no. 5, pp. 1054–1063, 1986
1986
Earlier work this paper cites.
D. O’Shaughnessy, “Linear predictive coding,” IEEE potentials , vol. 7, no. 1, pp. 29–32, 1988
1988
Earlier work this paper cites.
K. Brandenburg and G. Stoll, “ISO/MPEG-1 audio: A generic standard for coding of high-quality digital audio,” Journal of the Audio Engineering Society , vol. 42, no. 10, pp. 780–792, 1994
1994
Earlier work this paper cites.
R. Salami, C. Laflamme, J.-P. Adoul, and D. Massaloux, “A toll quality 8 kb/s speech codec for the personal communications system (pcs),” IEEE Transactions on Vehicular Technology , vol. 43, no. 3, pp. 808–816, 1994
1994
Earlier work this paper cites.
I. Recommendation, “Method for the subjective assessment of intermediate sound quality (MUSHRA),” ITU, BS , pp. 1543–1, 2001
2001
Earlier work this paper cites.
A. Vasuki and P. Vanathi, “A review of vector quantization techniques,” IEEE Potentials , vol. 25, no. 4, pp. 39–47, 2006
2006
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “A short-time objective intelligibility measure for time-frequency weighted noisy speech,” in Proc. ICASSP , 2010, pp. 4214–4217
2010
Earlier work this paper cites.
J.-M. Valin, G. Maxwell, T. B. Terriberry, and K. Vos, “High-quality, low-delay music coding in the opus codec,” in Audio Engineering Society Convention 135 , 2013
2013
Earlier work this paper cites.
A. L. Maas, A. Y. Hannun, and A. Y. Ng, “Rectifier nonlinearities improve neural network acoustic models,” in Proc. ICML , vol. 30, no. 1, 2013, p. 3
2013
Earlier work this paper cites.
M. Dietz, M. Multrus, V. Eksler, V. Malenovsky, E. Norvell, H. Pobloth, L. Miao, Z. Wang, L. Laaksonen, A. Vasilache et al. , “Overview of the EVS codec architecture,” in Proc. ICASSP , 2015, pp. 5698–5702
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals et al. , “Neural discrete representation learning,” in Proc. NIPS , vol. 30, 2017
2017
Earlier work this paper cites.
W. B. Kleijn, F. S. Lim, A. Luebs, J. Skoglund, F. Stimberg, Q. Wang, and T. C. Walters, “WaveNet based low rate speech coding,” in Proc. ICASSP , 2018, pp. 676–680
2018
Earlier work this paper cites.
S. Kankanahalli, “End-to-end optimized speech coding with deep neural networks,” in Proc. ICASSP , 2018, pp. 2521–2525
2018
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Proc. ICLR , 2018
2018
Cited alongside, same era.
J. Klejsa, P. Hedelin, C. Zhou, R. Fejgin, and L. Villemoes, “High-quality speech coding with sample RNN,” in Proc. ICASSP , 2019, pp. 7155–7159
2019
Cited alongside, same era.
J.-M. Valin and J. Skoglund, “A real-time wideband neural vocoder at 1.6 kb/s using LPCNet,” in Proc. Interspeech , 2019, pp. 3406–3410
2019
Cited alongside, same era.
C. Gârbacea, A. van den Oord, Y. Li, F. S. Lim, A. Luebs, O. Vinyals, and T. C. Walters, “Low bit-rate speech coding with VQ-VAE and a WaveNet decoder,” in Proc. ICASSP , 2019, pp. 735–739
2019
Cited alongside, same era.
K. Zhen, J. Sung, M. S. Lee, S. Beack, and M. Kim, “Cascaded cross-module residual learning towards lightweight end-to-end speech coding,” in Proc. Interspeech , 2019, pp. 3396–3400
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Zheng, L. Xiao, W. Tu, Y. Yang, and X. Xu, “CQNV: A combination of coarsely quantized bitstream and neural vocoder for low rate speech coding,” in Proc. Interspeech , 2023, pp. 171–175
2023
Later among the works it cites.
G. Davidson, M. Vinton, P. Ekstrand, C. Zhou, L. Villemoes, and L. Lu, “High quality audio coding with MDCTNet,” in Proc. ICASSP , 2023, pp. 1–5
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
K. Kumar, R. Kumar, T. de Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. de Brébisson, Y. Bengio, and A. C. Courville, “MelGAN: Generative adversarial networks for conditional waveform synthesis,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
J. Yamagishi, C. Veaux, K. MacDonald et al. , “CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),” University of Edinburgh. The Centre for Speech Technology Research (CSTR) , 2019
2019
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Proc. NIPS , vol. 33, 2020, pp. 17 022–17 033
2020
Cited alongside, same era.
M. Chinen, F. S. Lim, J. Skoglund, N. Gureev, F. O’Gorman, and A. Hines, “ViSQOL v3: An open source production ready objective speech and audio metric,” in Proc. QoMEX , 2020, pp. 1–6
2020
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in Proc. LREC , 2020, pp. 4218–4222
2020
Cited alongside, same era.
A. Mustafa, J. Büthe, S. Korse, K. Gupta, G. Fuchs, and N. Pia, “A streamwise GAN vocoder for wideband speech coding at very low bit rate,” in Proc. WASPAA , 2021, pp. 66–70
2021
Cited alongside, same era.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “SoundStream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Cited alongside, same era.
2023
Later among the works it cites.
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High fidelity neural audio compression,” Transactions on Machine Learning Research , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
L. Xu, J. Jiang, D. Zhang, X. Xia, L. Chen, Y. Xiao, P. Ding, S. Song, S. Yin, and F. Sohel, “An intra-BRNN and GB-RVQ based end-to-end neural audio codec,” in Proc. Interspeech , 2023, pp. 800–803
2023
Later among the works it cites.
T. Jenrungrot, M. Chinen, W. B. Kleijn, J. Skoglund, Z. Borsos, N. Zeghidour, and M. Tagliasacchi, “LMCodec: A low bitrate speech codec with causal transformer models,” in Proc. ICASSP , 2023, pp. 1–5
2023
Later among the works it cites.
W. Xiao, W. Liu, M. Wang, S. Yang, Y. Shi, Y. Kang, D. Su, S. Shang, and D. Yu, “Multi-mode neural speech coding based on deep generative networks,” in Proc. Interspeech , 2023, pp. 819–823
2023
Later among the works it cites.
Y.-C. Wu, I. D. Gebru, D. Marković, and A. Richard, “AudioDec: An open-source streaming high-fidelity neural audio codec,” in Proc. ICASSP , 2023, pp. 1–5
2023
Later among the works it cites.
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity audio compression with improved rvqgan,” Advances in Neural Information Processing Systems , vol. 36, 2023
2023
Later among the works it cites.
S. Woo, S. Debnath, R. Hu, X. Chen, Z. Liu, I. S. Kweon, and S. Xie, “ConvNeXt v2: Co-designing and scaling convnets with masked autoencoders,” in Proc. CVPR , 2023, pp. 16 133–16 142
2023
Later among the works it cites.
Y. Ai and Z.-H. Ling, “Neural speech phase prediction based on parallel estimation architecture and anti-wrapping losses,” in Proc. ICASSP , 2023, pp. 1–5
2023
Later among the works it cites.
Y. Ai and Z.-H. Ling, “APNet: An all-frame-level neural vocoder incorporating direct prediction of amplitude and phase spectra,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 2145–2157, 2023
2023
Later among the works it cites.
Y. Ren, T. Wang, J. Yi, L. Xu, J. Tao, C. Y. Zhang, and J. Zhou, “Fewer-token neural speech codec with time-invariant codes,” in Proc. ICASSP , 2024, pp. 12 737–12 741
2024
Closest in time.