Fetching the paper…
Reading the bibliography…
Although discrete speech tokens have exhibited strong potential for language model-based speech generation, their high bitrates and redundant timbre information restrict the development of such models.
W. Verhelst and M. Roelands, “An overlap-add technique based on waveform similarity (WSOLA) for high quality time-scale modification of speech,” in Proc. IEEE ICASSP , vol. 2. IEEE, 1993, pp. 554–557
1993
Earlier work this paper cites.
A. Van Den Oord, O. Vinyals et al. , “Neural discrete representation learning,” Proc. NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust DNN embeddings for speaker recognition,” in Proc. IEEE ICASSP , 2018, pp. 5329–5333
2018
Earlier work this paper cites.
K. Qian, Y. Zhang, S. Chang, X. Yang, and M. Hasegawa-Johnson, “AutoVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss,” in Proc. ICML . PMLR, 2019, pp. 5210–5219
2019
Earlier work this paper cites.
H. Zen, V. Dang, R. Clark et al. , “LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,” in Proc. ISCA Interspeech , 2019, pp. 1526–1530
2019
Earlier work this paper cites.
A. Baevski, S. Schneider, and M. Auli, “vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations,” in Proc. ICLR , 2020
2020
Earlier work this paper cites.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,” Proc. NeurIPS , vol. 33, pp. 12 449–12 460, 2020
2020
Earlier work this paper cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Proc. ISCA Interspeech , 2020, pp. 5036–5040
2020
Earlier work this paper cites.
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” IEEE/ACM Trans. ASLP. , vol. 29, pp. 3451–3460, 2021
2021
Earlier work this paper cites.
S. wen Yang, P.-H. Chi, Y.-S. Chuang et al. , “SUPERB: Speech Processing Universal PERformance Benchmark,” in Proc. ISCA Interspeech , 2021, pp. 1194–1198
2021
Earlier work this paper cites.
C. Du, Y. Guo, X. Chen, and K. Yu, “VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature,” in Proc. ISCA Interspeech , 2022, pp. 1596–1600
2022
Earlier work this paper cites.
S. Chen, C. Wang, Z. Chen et al. , “WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Earlier work this paper cites.
K. Qian, Y. Zhang, H. Gao, J. Ni, C.-I. Lai, D. Cox, M. Hasegawa-Johnson, and S. Chang, “Contentvec: An improved self-supervised speech representation by disentangling speakers,” in Proc. ICML . PMLR, 2022, pp. 18 003–18 017
2022
Earlier work this paper cites.
C. H. Chan, K. Qian, Y. Zhang, and M. Hasegawa-Johnson, “SpeechSplit2.0: Unsupervised Speech Disentanglement for Voice Conversion without Tuning Autoencoder Bottlenecks,” in Proc. IEEE ICASSP , 2022, pp. 6332–6336
2022
Earlier work this paper cites.
A. Défossez, J. Copet, G. Synnaeve, and Y. Adi, “High Fidelity Neural Audio Compression,” Transactions on Machine Learning Research , 2023
2023
Cited alongside, same era.
X. Jiang, X. Peng, Y. Zhang, and Y. Lu, “Disentangled feature learning for real-time neural speech coding,” in Proc. IEEE ICASSP . IEEE, 2023, pp. 1–5
2023
Cited alongside, same era.
H.-J. Chang, A. H. Liu, and J. Glass, “Self-supervised fine-tuning for improved content representations by speaker-invariant clustering,” in Proc. ISCA Interspeech , 2023, pp. 2983–2987
2023
Cited alongside, same era.
D. Yang, J. Tian, X. Tan et al. , “UniAudio: Towards Universal Audio Generation with Large Language Models,” in Proc. ICML , 2024
2024
Cited alongside, same era.
Y. Yang, F. Shen, C. Du, Z. Ma, K. Yu, D. Povey, and X. Chen, “Towards universal speech discrete tokens: A case study for ASR and TTS,” in Proc. IEEE ICASSP , 2024, pp. 10 401–10 405
2024
Closest in time.
Z. Ju, Y. Wang, K. Shen et al. , “NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models,” in Proc. ICML , 2024
2024
Closest in time.
2024
Closest in time.
C. Du, Y. Guo, F. Shen, Z. Liu, Z. Liang, X. Chen, S. Wang, H. Zhang, and K. Yu, “UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding,” in Proc. AAAI , vol. 38, no. 16, 2024, pp. 17 924–17 932
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar, “High-fidelity audio compression with improved RVQGAN,” Proc. NeurIPS , vol. 36, 2024
2024
Cited alongside, same era.
P. Peng, P.-Y. Huang, S.-W. Li, A. Mohamed, and D. Harwath, “VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild,” in Proc. ACL , Aug. 2024, pp. 12 442–12 462
2024
Cited alongside, same era.
2024
Cited alongside, same era.
J. Shi, X. Ma, H. Inaguma, A. Sun, and S. Watanabe, “MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model,” in Proc. ISCA Interspeech , 2024, pp. 2569–2573
2024
Cited alongside, same era.
F. Shen, Y. Guo, C. Du, X. Chen, and K. Yu, “Acoustic BPE for speech generation with discrete tokens,” in Proc. IEEE ICASSP . IEEE, 2024, pp. 11 746–11 750
2024
Cited alongside, same era.
Y. Ren, T. Wang, J. Yi, L. Xu, J. Tao, C. Y. Zhang, and J. Zhou, “Fewer-token neural speech codec with time-invariant codes,” in Proc. IEEE ICASSP . IEEE, 2024, pp. 12 737–12 741
2024
Cited alongside, same era.
H. Li, L. Xue, H. Guo, X. Zhu, Y. Lv, L. Xie, Y. Chen, H. Yin, and Z. Li, “Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation,” in Proc. ISCA Interspeech , 2024, pp. 3390–3394
2024
Cited alongside, same era.
J. Li, Y. Guo, X. Chen, and K. Yu, “SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention,” in Proc. IEEE ICASSP , 2024, pp. 12 296–12 300
2024
Closest in time.
J. weon Jung, W. Zhang, J. Shi et al. , “ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models,” in Proc. ISCA Interspeech , 2024, pp. 4278–4282
2024
Closest in time.
H. Siuzdak, “Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis,” in Proc. ICLR , 2024
2024
Closest in time.
H. Liu, X. Xu, Y. Yuan, M. Wu, W. Wang, and M. D. Plumbley, “SemantiCodec: An ultra low bitrate semantic audio codec for general sound,” IEEE Journal of Selected Topics in Signal Processing , vol. 18, no. 8, pp. 1448–1461, 2024
2024
Closest in time.
D. Yang, H. Guo, Y. Wang, R. Huang, X. Li, X. Tan, X. Wu, and H. M. Meng, “UniAudio 1.5: Large language model-driven audio codec is a few-shot audio task learner,” in Proc. NeurIPS , 2024
2024
Closest in time.
S. Chen, C. Wang, Y. Wu et al. , “Neural codec language models are zero-shot text to speech synthesizers,” IEEE/ACM Trans. ASLP. , vol. 33, pp. 705–718, 2025
2025
Closest in time.
2025
Closest in time.
S. Ji, Z. Jiang, W. Wang et al. , “WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling,” in Proc. ICLR , 2025
2025
Closest in time.
J. D. Parker, A. Smirnov, J. Pons, C. Carr, Z. Zukowski, Z. Evans, and X. Liu, “Scaling transformers for low-bitrate high-quality speech coding,” in Proc. ICLR , 2025
2025
Closest in time.