Fetching the paper…
Reading the bibliography…
We present BigCodec, a low-bitrate neural speech codec.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” in Proc. NeurIPS , vol. 33, 2020, pp. 1877–1901
1901
Earlier work this paper cites.
P. B. Denes, “On the statistics of spoken english,” The Journal of the Acoustical Society of America , vol. 35, no. 6, pp. 892–904, 1963
1963
Earlier work this paper cites.
B. Atal, “Predictive coding of speech at low bit rates,” IEEE Transactions on Communications , vol. 30, no. 4, pp. 600–614, 1982
1982
Earlier work this paper cites.
D. V. Cicchetti, “Guidelines, criteria, and rules of thumb for evaluating normed and standardized assessment instruments in psychology.” Psychological assessment , vol. 6, no. 4, p. 284, 1994
1994
Earlier work this paper cites.
N. Tishby, F. C. Pereira, and W. Bialek, “The information bottleneck method,” arXiv preprint physics/0004057 , 2000
2000
Earlier work this paper cites.
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra, “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in Proc. ICASSP , vol. 2. IEEE, 2001, pp. 749–752
2001
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens, and J. Jensen, “An algorithm for intelligibility prediction of time–frequency weighted noisy speech,” IEEE Transactions on audio, speech, and language processing , vol. 19, no. 7, pp. 2125–2136, 2011
2011
Earlier work this paper cites.
J.-M. Valin, K. Vos, and T. B. Terriberry, “Definition of the Opus Audio Codec,” RFC 6716, Sep. 2012. [Online]. Available: https://www.rfc-editor.org/info/rfc6716
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
B. Series, “Method for the subjective assessment of intermediate quality level of audio systems,” International Telecommunication Union Radiocommunication Assembly , vol. 2, 2014
2014
Earlier work this paper cites.
M. Dietz, M. Multrus, V. Eksler, V. Malenovsky, E. Norvell, H. Pobloth, L. Miao, Z. Wang, L. Laaksonen, A. Vasilache et al. , “Overview of the evs codec architecture,” in Proc. ICASSP . IEEE, 2015, pp. 5698–5702
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in Proc. ICASSP . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
S. Van Kuyk, W. B. Kleijn, and R. C. Hendriks, “On the information rate of speech communication,” in Proc. ICASSP . IEEE, 2017, pp. 5625–5629
2017
Earlier work this paper cites.
A. van den Oord, O. Vinyals, and k. kavukcuoglu, “Neural discrete representation learning,” in Proc. NeurIPS , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds., vol. 30. Curran Associates, Inc., 2017
2017
Earlier work this paper cites.
X. Mao, Q. Li, H. Xie, R. Y. Lau, Z. Wang, and S. Paul Smolley, “Least squares generative adversarial networks,” in Proc. ICCV , 2017, pp. 2794–2802
2017
Earlier work this paper cites.
W. B. Kleijn, F. S. Lim, A. Luebs, J. Skoglund, F. Stimberg, Q. Wang, and T. C. Walters, “Wavenet based low rate speech coding,” in Proc. ICASSP . IEEE, 2018, pp. 676–680
2018
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in Proc. ICLR , 2018
2018
Cited alongside, same era.
M. Schoeffler, S. Bartoschek, F.-R. Stoter, M. Roess, S. Westphal, B. Edler, and J. Herre, “webmushra-a comprehensive framework for web-based listening tests,” Journal of Open Research Software , vol. 6, no. 7, 2018
2018
Cited alongside, same era.
K. Kumar, R. Kumar, T. De Boissiere, L. Gestin, W. Z. Teoh, J. Sotelo, A. De Brebisson, Y. Bengio, and A. C. Courville, “Melgan: Generative adversarial networks for conditional waveform synthesis,” in Proc. NeurIPS , vol. 32, 2019
2019
Cited alongside, same era.
C. K. Reddy, H. Dubey, V. Gopal, R. Cutler, S. Braun, H. Gamper, R. Aichner, and S. Srinivasan, “Icassp 2021 deep noise suppression challenge,” in Proc. ICASSP . IEEE, 2021, pp. 6623–6627
2021
Later among the works it cites.
T. Jayashankar, T. Koehler, K. Kalgaonkar, Z. Xiu, J. Wu, J. Lin, P. Agrawal, and Q. He, “Architecture for variable bitrate neural speech codec with configurable computation complexity,” in Proc. ICASSP . IEEE, 2022, pp. 861–865
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Communications of the ACM , vol. 63, no. 11, pp. 139–144, 2020
2020
Cited alongside, same era.
L. Ziyin, T. Hartwig, and M. Ueda, “Neural networks fail to learn periodic functions and how to fix it,” in Proc. NeurIPS , vol. 33, 2020, pp. 1583–1594
2020
Cited alongside, same era.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” in Proc. NeurIPS , vol. 33, 2020, pp. 17 022–17 033
2020
Cited alongside, same era.
R. Yamamoto, E. Song, and J.-M. Kim, “Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,” in Proc. ICASSP . IEEE, 2020, pp. 6199–6203
2020
Cited alongside, same era.
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A Large-Scale Multilingual Dataset for Speech Research,” in Proc. Interspeech , 2020, pp. 2757–2761
2020
Cited alongside, same era.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux, “Libri-light: A benchmark for asr with limited or no supervision,” in Proc. ICASSP , 2020, pp. 7669–7673, https://github.com/facebookresearch/libri-light
2020
Cited alongside, same era.
N. Zeghidour, A. Luebs, A. Omran, J. Skoglund, and M. Tagliasacchi, “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Later among the works it cites.
X. Jiang, X. Peng, H. Xue, Y. Zhang, and Y. Lu, “Latent-domain predictive neural speech coding,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 2111–2123, 2023
2023
Later among the works it cites.
T. Jenrungrot, M. Chinen, W. B. Kleijn, J. Skoglund, Z. Borsos, N. Zeghidour, and M. Tagliasacchi, “Lmcodec: A low bitrate speech codec with causal transformer models,” in Proc. ICASSP . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
Y.-C. Wu, I. D. Gebru, D. Marković, and A. Richard, “Audiodec: An open-source streaming high-fidelity neural audio codec,” in Proc. ICASSP . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity audio compression with improved rvqgan,” in Proc. NeurIPS , vol. 36, 2024
2024
Closest in time.
Y. Zheng, W. Tu, L. Xiao, and X. Xu, “Srcodec: Split-residual vector quantization for neural speech codec,” in Proc. ICASSP . IEEE, 2024, pp. 451–455
2024
Closest in time.
——, “Supercodec: A neural speech codec with selective back-projection network,” in Proc. ICASSP . IEEE, 2024, pp. 566–570
2024
Closest in time.
2024
Closest in time.