Fetching the paper…
Reading the bibliography…
Neural speech codecs have gained great attention for their outstanding reconstruction with discrete token representations.
L. M. Supplee, R. P. Cohn et al. , “Melp: the new federal standard at 2400 bps,” in 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing , vol. 2. IEEE, 1997, pp. 1591–1594
1997
Earlier work this paper cites.
R. ITU-R, “1534-1, method for the subjective assessment of intermediate quality levels of coding systems (mushra),”,” International Telecommunication Union , 2003
2003
Earlier work this paper cites.
C. H. Taal, Hendriks et al. , “A short-time objective intelligibility measure for time-frequency weighted noisy speech,” in 2010 IEEE international conference on acoustics, speech and signal processing . IEEE, 2010, pp. 4214–4217
2010
Earlier work this paper cites.
D. Rowe, “Codec 2-open source speech coding at 2400 bits/s and below,” in TAPR and ARRL 30th Digital Communications Conference , 2011, pp. 80–84
2011
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
J.-M. Valin, “Speex: A free codec for free speech,” arXiv preprint arXiv:1602.08668 , 2016
2016
Earlier work this paper cites.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , 2017, pp. 6309–6318
2017
Earlier work this paper cites.
J. Yamagishi, C. Veaux et al. , “Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),” 2019
2019
Earlier work this paper cites.
B. Desplanques, J. Thienpondt, and K. Demuynck, “Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,” in Interspeech 2020 , 2020, pp. 3830–3834
2020
Earlier work this paper cites.
A. Polyak, Y. Adi et al. , “Speech Resynthesis from Discrete Disentangled Self-Supervised Representations,” in Proc. Interspeech 2021 , 2021, pp. 3615–3619
2021
Earlier work this paper cites.
D. Wang, L. Deng et al. , “Vqmivc: Vector quantization and mutual information-based unsupervised speech representation disentanglement for one-shot voice conversion,” in Interspeech 2021 , 2021, pp. 1344–1348
2021
Earlier work this paper cites.
W. A. Jassim, J. Skoglund et al. , “Warp-q: Quality prediction for generative neural speech codecs,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 401–405
2021
Earlier work this paper cites.
N. Zeghidour, A. Luebs et al. , “Soundstream: An end-to-end neural audio codec,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 495–507, 2022
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
Y. Ren, M. Lei, Z. Huang et al. , “Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 7577–7581
2022
Cited alongside, same era.
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 11 976–11 986
2022
Cited alongside, same era.
S. Chen, C. Wang et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Cited alongside, same era.
Z. Ju, Y. Wang et al. , “Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,” in Forty-first International Conference on Machine Learning , 2024. [Online]. Available: https://openreview.net/forum?id=dVhrnjZJad
2024
Closest in time.
Y. Ren, T. Wang, J. Yi et al. , “Fewer-token neural speech codec with time-invariant codes,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 12 737–12 741
2024
Closest in time.
X. Zhang, D. Zhang, S. Li, Y. Zhou, and X. Qiu, “Speechtokenizer: Unified speech tokenizer for speech language models,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
H. Li, L. Xue et al. , “Single-codec: Single-codebook speech codec towards high-performance speech generation,” in Interspeech 2024 , 2024, pp. 3390–3394
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. Casanova, J. Weber et al. , “Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,” in International Conference on Machine Learning . PMLR, 2022, pp. 2709–2720
2022
Cited alongside, same era.
T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “Utmos: Utokyo-sarulab system for voicemos challenge 2022,” in Interspeech 2022 , 2022, pp. 4521–4525
2022
Cited alongside, same era.
2023
Cited alongside, same era.
Z. Borsos, R. Marinier, D. Vincent, E. Kharitonov et al. , “Audiolm: a language modeling approach to audio generation,” IEEE/ACM transactions on audio, speech, and language processing , vol. 31, pp. 2523–2533, 2023
2023
Cited alongside, same era.
X. Jiang, X. Peng, H. Xue, Y. Zhang, and Y. Lu, “Latent-domain predictive neural speech coding,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 2111–2123, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
J. Li, W. Tu, and L. Xiao, “Freevc: Towards high-quality text-free one-shot voice conversion,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5
2023
Cited alongside, same era.
2024
Cited alongside, same era.
H. Liu, X. Xu et al. , “Semanticodec: An ultra low bitrate semantic audio codec for general sound,” IEEE Journal of Selected Topics in Signal Processing , vol. 18, no. 8, pp. 1448–1461, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
R. Kumar, P. Seetharaman et al. , “High-fidelity audio compression with improved rvqgan,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Z. Jiang, J. Liu, Y. Ren et al. , “Mega-TTS 2: Boosting prompting mechanisms for zero-shot speech synthesis,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=mvMI3N4AvD
2024
Closest in time.
2024
Closest in time.
J. Lim and K. Kim, “Wav2vec-vc: Voice conversion via hidden representations of wav2vec 2.0,” in ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 10 326–10 330
2024
Closest in time.
S. Ji, Z. Jiang et al. , “Wavtokenizer: an efficient acoustic discrete codec tokenizer for audio language modeling,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=yBlVlS2Fd9
2025
Closest in time.