Fetching the paper…
Reading the bibliography…
Neural speech codecs have recently emerged as a focal point in the fields of speech compression and generation.
A. W. Rix, J. G. Beerends, M. P. Hollier , et al. , “Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,” in 2001 IEEE international conference on acoustics, speech, and signal processing. Proceedings (Cat. No. 01CH37221) , vol. 2. IEEE, 2001, pp. 749–752
2001
Earlier work this paper cites.
Z. Wang, A. C. Bovik, H. R. Sheikh , et al. , “Image quality assessment: from error visibility to structural similarity,” IEEE transactions on image processing , vol. 13, no. 4, pp. 600–612, 2004
2004
Earlier work this paper cites.
C. H. Taal, R. C. Hendriks, R. Heusdens , et al. , “A short-time objective intelligibility measure for time-frequency weighted noisy speech,” in 2010 IEEE international conference on acoustics, speech and signal processing . IEEE, 2010, pp. 4214–4217
2010
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Dietz, M. Multrus, V. Eksler , et al. , “Overview of the evs codec architecture,” in 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2015, pp. 5698–5702
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals , et al. , “Neural discrete representation learning,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar , et al. , “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
C. Gârbacea, A. van den Oord, Y. Li , et al. , “Low bit-rate speech coding with vq-vae and a wavenet decoder,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 735–739
2019
Earlier work this paper cites.
J. Klejsa, P. Hedelin, C. Zhou , et al. , “High-quality speech coding with sample rnn,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 7155–7159
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
M. Chinen, F. S. Lim, J. Skoglund , et al. , “Visqol v3: An open source production ready objective speech and audio metric,” in 2020 twelfth international conference on quality of multimedia experience (QoMEX) . IEEE, 2020, pp. 1–6
2020
Earlier work this paper cites.
J. Kong, J. Kim, and J. Bae, “Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,” Advances in neural information processing systems , vol. 33, pp. 17 022–17 033, 2020
2020
Earlier work this paper cites.
N. Zeghidour, A. Luebs, A. Omran , et al. , “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2021
2021
Earlier work this paper cites.
Y.-H. Chen, D.-Y. Wu, T.-H. Wu , et al. , “Again-vc: A one-shot voice conversion using activation guidance and adaptive instance normalization,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 5954–5958
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Cited alongside, same era.
R. Rombach, A. Blattmann, D. Lorenz , et al. , “High-resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 10 684–10 695
2022
Cited alongside, same era.
2022
Cited alongside, same era.
C. Ho Chan, K. Qian, Y. Zhang , et al. , “Speechsplit2.0: Unsupervised speech disentanglement for voice conversion without tuning autoencoder bottlenecks,” in ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 6332–6336
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Yang, Y. Pan, J. Yin , et al. , “Hybridformer: Improving squeezeformer with hybrid attention and nsr mechanism,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Later among the works it cites.
H. Xue, X. Peng, and Y. Lu, “Low-latency speech enhancement via speech token generation,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 661–665
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. A. Trinh and S. Braun, “Unsupervised speech enhancement with speech recognition embedding and disentanglement losses,” in ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 391–395
2022
Cited alongside, same era.
2023
Cited alongside, same era.
J. Betker, “Better speech synthesis through scaling,” arXiv preprint arXiv:2305.07243 , 2023
2023
Cited alongside, same era.
Z. Borsos, R. Marinier, D. Vincent , et al. , “Audiolm: a language modeling approach to audio generation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Z. Wang, Y. Chen, L. Xie , et al. , “Lm-vc: Zero-shot voice conversion via speech generation based on language models,” IEEE Signal Processing Letters , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
T. Jenrungrot, M. Chinen, W. B. Kleijn , et al. , “Lmcodec: A low bitrate speech codec with causal transformer models,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Cited alongside, same era.
2024
Closest in time.
J. Büthe, A. Mustafa, J.-M. Valin , et al. , “Nolace: Improving low-complexity speech codec enhancement through adaptive temporal shaping,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 476–480
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Du, S. Zhang, K. Hu , et al. , “Funcodec: A fundamental, reproducible and integrable open-source toolkit for neural speech codec,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 591–595
2024
Closest in time.
R. Kumar, P. Seetharaman, A. Luebs , et al. , “High-fidelity audio compression with improved rvqgan,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
H. Yang, I. Jang, and M. Kim, “Generative de-quantization for neural speech codec via latent diffusion,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 1251–1255
2024
Closest in time.
2024
Closest in time.
R. San Roman, Y. Adi, A. Deleforge , et al. , “From discrete tokens to high-fidelity audio using multi-band diffusion,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.