Fetching the paper…
Reading the bibliography…
Recent advances in text-to-speech synthesis have achieved notable success in generating high-quality short utterances for individual speakers.
Fisher english training speech part 1 transcripts ldc2004t19
Cieri, C · 2004
Earlier work this paper cites.
Paralinguistics in speech and language—state-of-the-art and the challenge
Schuller, B., Steidl, S., Batliner, A., Burkhardt, F., Devillers, L., MüLler, C., and Narayanan, S · 2013
Earlier work this paper cites.
wav2vec: Unsupervised pre-training for speech recognition
Schneider, S., Baevski, A., Collobert, R., and Auli, M · 2019
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Adiwardana, D., Luong, M.-T., So, D. R., Hall, J., Fiedel, N., Thoppilan, R., Yang, Z., Kulshreshtha, A., Nemade, G., Lu, Y., et al · 2020
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
DIALOGPT : Large-scale generative pre-training for conversational response generation
Zhang, Y., Sun, S., Galley, M., Chen, Y.-C., Brockett, C., Gao, X., Gao, J., Liu, J., and Dolan, B · 2020
Earlier work this paper cites.
W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Chung, Y.-A., Zhang, Y., Han, W., Chiu, C.-C., Qin, J., Pang, R., and Wu, Y · 2021
Earlier work this paper cites.
Beyond goldfish memory: Long-term open-domain conversation
Xu, J · 2021
Earlier work this paper cites.
AdaSpeech 3: Adaptive text to speech for spontaneous style
Yan, Y., Tan, X., Li, B., Zhang, G., Qin, T., Zhao, S., Shen, Y., Zhang, W.-Q., and Liu, T.-Y · 2021
Earlier work this paper cites.
Vector-quantized image modeling with improved vqgan
Yu, J., Li, X., Koh, J. Y., Zhang, H., Pang, R., Qin, J., Ku, A., Xu, Y., Baldridge, J., and Wu, Y · 2021
Earlier work this paper cites.
Audiolm: a language modeling approach to audio generation
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Teboul, O., Grangier, D., Tagliasacchi, M., and Zeghidour, N · 2022
Earlier work this paper cites.
Gao, Z., Zhang, S., McLoughlin, I., and Yan, Z · 2022
Earlier work this paper cites.
Bigvgan: A universal neural vocoder with large-scale training
Lee, S.-g., Ping, W., Ginsburg, B., Catanzaro, B., and Yoon, S · 2022
Cited alongside, same era.
High fidelity speech enhancement with band-split rnn
Yu, J., Luo, Y., Chen, H., Gu, R., and Weng, C · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Seamless: Multilingual expressive and streaming speech translation
Barrault, L., Chung, Y.-A., Meglioli, M. C., Dale, D., Dong, N., Duppenthaler, M., Duquenne, P.-A., Ellis, B., Elsahar, H., Haaheim, J., et al · 2023
Cited alongside, same era.
Vall-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers
Chen, S., Liu, S., Zhou, L., Liu, Y., Tan, X., Li, J., Zhao, S., Qian, Y., and Wei, F · 2024
Later among the works it cites.
Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation
He, H., Shang, Z., Wang, C., Li, X., Gu, Y., Hua, H., Liu, L., Yang, C., Li, J., Shi, P., et al · 2024
Later among the works it cites.
Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Ju, Z., Wang, Y., Shen, K., Tan, X., Xin, D., Yang, D., Liu, Y., Leng, Y., Song, K., Tang, S., et al · 2024
Later among the works it cites.
Base tts: Lessons from building a billion-parameter text-to-speech model on 100k hours of data
Łajszczak, M., Cámbara, G., Li, Y., Beyhan, F., van Korlaar, A., Yang, F., Joly, A., Martín-Cortinas, Á., Abbas, A., Michalski, A., et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Borsos, Z., Sharifi, M., Vincent, D., Kharitonov, E., Zeghidour, N., and Tagliasacchi, M · 2023
Cited alongside, same era.
High-fidelity audio compression with improved rvqgan
Kumar, R., Seetharaman, P., Luebs, A., Kumar, I., and Kumar, K · 2023
Cited alongside, same era.
Towards human-like spoken dialogue generation between ai agents from written dialogue
Mitsui, K., Hono, Y., and Sawada, K · 2023
Cited alongside, same era.
Generative spoken dialogue language modeling
Nguyen, T. A., Kharitonov, E., Copet, J., Adi, Y., Hsu, W.-N., Elkahky, A., Tomasello, P., Algayres, R., Sagot, B., Mohamed, A., et al · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
Peebles, W. and Xie, S · 2023
Cited alongside, same era.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers
Shen, K., Ju, Z., Tan, X., Liu, Y., Leng, Y., He, L., Qin, T., Zhao, S., and Bian, J · 2023
Cited alongside, same era.
Neural codec language models are zero-shot text to speech synthesizers
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., et al · 2023
Cited alongside, same era.
Seed-tts: A family of high-quality versatile speech generation models
Anastassiou, P., Chen, J., Chen, J., Chen, Y., Chen, Z., Chen, Z., Cong, J., Deng, L., Ding, C., Gao, L., et al · 2024
Cited alongside, same era.
Voicebox: Text-guided multilingual universal speech generation at scale
Le, M., Vyas, A., Shi, B., Karrer, B., Sari, L., Moritz, R., Williamson, M., Manohar, V., Adi, Y., Mahadeokar, J., et al · 2024
Later among the works it cites.
Spontts: modeling and transferring spontaneous style for tts
Li, H., Zhu, X., Xue, L., Song, Y., Chen, Y., and Xie, L · 2024
Later among the works it cites.
Autoregressive diffusion transformer for text-to-speech synthesis
Liu, Z., Wang, S., Inoue, S., Bai, Q., and Li, H · 2024
Later among the works it cites.
Maskgct: Zero-shot text-to-speech with masked generative codec transformer
Wang, Y., Zhan, H., Liu, L., Zeng, R., Guo, H., Zheng, J., Zhang, Q., Zhang, X., Zhang, S., and Wu, Z · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al · 2024
Later among the works it cites.
Autoprep: An automatic preprocessing framework for in-the-wild speech data
Yu, J., Chen, H., Bian, Y., Li, X., Luo, Y., Tian, J., Liu, M., Jiang, J., and Wang, S · 2024
Later among the works it cites.
Covomix: Advancing zero-shot speech generation for human-like multi-talker conversations
Zhang, L., Qian, Y., Zhou, L., Liu, S., Wang, D., Wang, X., Yousefi, M., Qian, Y., Li, J., He, L., et al · 2024
Later among the works it cites.
Slide: Integrating speech language model with llm for spontaneous spoken dialogue generation
Lu, H., Cheng, G., Luo, L., Zhang, L., Qian, Y., and Zhang, P · 2025
Closest in time.