Fetching the paper…
Reading the bibliography…
Recent advances in text-based large language models (LLMs), particularly in the GPT series and the o1 model, have demonstrated the effectiveness of scaling both training-time and inference-time compute.
Fastspeech 2: Fast and high-quality end-to-end text to speech, 2022
Ren, Y., Hu, C., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2006
Earlier work this paper cites.
Mls: A large-scale multilingual dataset for speech research
Pratap, V., Xu, Q., Sriram, A., Synnaeve, G., and Collobert, R · 2012
Earlier work this paper cites.
Mls: A large-scale multilingual dataset for speech research
Pratap, V., Xu, Q., Sriram, A., Synnaeve, G., and Collobert, R · 2012
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Kong, J., Kim, J., and Bae, J · 2020
Earlier work this paper cites.
Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio
Chen, G., Chai, S., Wang, G., Du, J., Zhang, W.-Q., Weng, C., Su, D., Povey, D., Trmal, J., Zhang, J., et al · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors
Reddy, C. K., Gopal, V., and Cutler, R · 2021
Earlier work this paper cites.
Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset
Zhou, K., Sisman, B., Liu, R., and Li, H · 2021
Earlier work this paper cites.
Speech Communication , 137:1–18, 2022
Emotional voice conversion: Theory, databases and esd · 2022
Earlier work this paper cites.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Chen, S., Wang, C., Chen, Z., Wu, Y., Liu, S., Chen, Z., Li, J., Kanda, N., Yoshioka, T., Xiao, X., et al · 2022
Earlier work this paper cites.
High fidelity neural audio compression
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2022
Earlier work this paper cites.
Utmos: Utokyo-sarulab system for voicemos challenge 2022
Saeki, T., Xin, D., Nakata, W., Koriyama, T., Takamichi, S., and Saruwatari, H · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Seamless: Multilingual expressive and streaming speech translation
Barrault, L., Chung, Y.-A., Meglioli, M. C., Dale, D., Dong, N., Duppenthaler, M., Duquenne, P.-A., Ellis, B., Elsahar, H., Haaheim, J., et al · 2023
Earlier work this paper cites.
Better speech synthesis through scaling
Betker, J · 2023
Cited alongside, same era.
Audiolm: a language modeling approach to audio generation
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Roblek, D., Teboul, O., Grangier, D., Tagliasacchi, M., et al · 2023
Cited alongside, same era.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
Kharitonov, E., Vincent, D., Borsos, Z., Marinier, R., Girgin, S., Pietquin, O., Sharifi, M., Tagliasacchi, M., and Zeghidour, N · 2023
Cited alongside, same era.
Diverse and expressive speech prosody prediction with denoising diffusion probabilistic model
Li, X., Liu, S., Lam, M. W., Wu, Z., Weng, C., and Meng, H · 2023
Cited alongside, same era.
emotion2vec: Self-supervised pre-training for speech emotion representation
High-fidelity audio compression with improved rvqgan
Kumar, R., Seetharaman, P., Luebs, A., Kumar, I., and Kumar, K · 2024
Later among the works it cites.
BASE TTS: lessons from building a billion-parameter text-to-speech model on 100k hours of data
Lajszczak, M., Cámbara, G., Li, Y., Beyhan, F., van Korlaar, A., Yang, F., Joly, A., Martín-Cortinas, Á., Abbas, A., Michalski, A., Moinet, A., Karlapati, S., Muszynska, E., Guo, H., Putrycz, B., Gambino, S. L., Yoo, K., Sokolova, E., and Drugman, T · 2024
Later among the works it cites.
Base tts: Lessons from building a billion-parameter text-to-speech model on 100k hours of data
Łajszczak, M., Cámbara, G., Li, Y., Beyhan, F., van Korlaar, A., Yang, F., Joly, A., Martín-Cortinas, Á., Abbas, A., Michalski, A., et al · 2024
Later among the works it cites.
Voicebox: Text-guided multilingual universal speech generation at scale
Le, M., Vyas, A., Shi, B., Karrer, B., Sari, L., Moritz, R., Williamson, M., Manohar, V., Adi, Y., Mahadeokar, J., et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ma, Z., Zheng, Z., Ye, J., Li, J., Gao, Z., Zhang, S., and Chen, X · 2023
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Neural codec language models are zero-shot text to speech synthesizers
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., et al · 2023
Cited alongside, same era.
Speak foreign languages with your own voice: Cross-lingual neural codec language modeling
Zhang, Z., Zhou, L., Wang, C., Chen, S., Wu, Y., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., He, L., Zhao, S., and Wei, F · 2023
Cited alongside, same era.
Seed-tts: A family of high-quality versatile speech generation models
Anastassiou, P., Chen, J., Chen, J., Chen, Y., Chen, Z., Chen, Z., Cong, J., Deng, L., Ding, C., Gao, L., et al · 2024
Cited alongside, same era.
Moshi: a speech-text foundation model for real-time dialogue
Défossez, A., Mazaré, L., Orsini, M., Royer, A., Pérez, P., Jégou, H., Grave, E., and Zeghidour, N · 2024
Cited alongside, same era.
E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts
Eskimez, S. E., Wang, X., Thakker, M., Li, C., Tsai, C.-H., Xiao, Z., Yang, H., Zhu, Z., Tang, M., Tan, X., et al · 2024
Cited alongside, same era.
Liu, H., Xu, X., Yuan, Y., Wu, M., Wang, W., and Plumbley, M. D · 2024
Later among the works it cites.
Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model benchmark
Ma, L., Guo, D., Song, K., Jiang, Y., Wang, S., Xue, L., Xu, W., Zhao, H., Zhang, B., and Xie, L · 2024
Later among the works it cites.
Finite scalar quantization: VQ-VAE made simple
Mentzer, F., Minnen, D., Agustsson, E., and Tschannen, M · 2024
Later among the works it cites.
Scaling transformers for low-bitrate high-quality speech coding
Parker, J. D., Smirnov, A., Pons, J., Carr, C., Zukowski, Z., Evans, Z., and Liu, X · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering
Song, Y., Chen, Z., Wang, X., Ma, Z., and Chen, X · 2024
Later among the works it cites.
Maskgct: Zero-shot text-to-speech with masked generative codec transformer
Wang, Y., Zhan, H., Liu, L., Zeng, R., Guo, H., Zheng, J., Zhang, Q., Zhang, X., Zhang, S., and Wu, Z · 2024
Later among the works it cites.
Bigcodec: Pushing the limits of low-bitrate neural speech codec
Xin, D., Tan, X., Takamichi, S., and Saruwatari, H · 2024
Later among the works it cites.
Codec does matter: Exploring the semantic shortcoming of codec for audio language model
Ye, Z., Sun, P., Lei, J., Lin, H., Tan, X., Dai, Z., Kong, Q., Chen, J., Pan, J., Liu, Q., et al · 2024
Later among the works it cites.
Speechtokenizer: Unified speech tokenizer for speech language models
Zhang, X., Zhang, D., Li, S., Zhou, Y., and Qiu, X · 2024
Later among the works it cites.
Scaling transformers for low-bitrate high-quality speech coding
Anonymous · 2025
Closest in time.
Test-time computing: from system-1 thinking to system-2 thinking
Ji, Y., Li, J., Ye, H., Wu, K., Xu, J., Mo, L., and Zhang, M · 2025
Closest in time.
Inference-time scaling for diffusion models beyond scaling denoising steps
Ma, N., Tong, S., Jia, H., Hu, H., Su, Y.-C., Zhang, M., Yang, X., Li, Y., Jaakkola, T., Jia, X., et al · 2025
Closest in time.