Fetching the paper…
Reading the bibliography…
Benefiting from effective speech modeling, current Speech Large Language Models (SLLMs) have demonstrated exceptional capabilities in in-context speech generation and efficient generalization to unseen speakers.
Librispeech: An asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Neural ordinary differential equations, 2019
Chen, R. T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D · 2019
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus, 2020
Ardila, R., Branson, M., Davis, K., Henretty, M., Kohler, M., Meyer, J., Morais, R., Saunders, L., Tyers, F. M., and Weber, G · 2020
Earlier work this paper cites.
Conformer: Convolution-augmented transformer for speech recognition, 2020
Gulati, A., Qin, J., Chiu, C.-C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., and Pang, R · 2020
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Mls: A large-scale multilingual dataset for speech research
Pratap, V., Xu, Q., Sriram, A., Synnaeve, G., and Collobert, R · 2020
Earlier work this paper cites.
Unsupervised speech decomposition via triple information bottleneck
Qian, K., Zhang, Y., Chang, S., Hasegawa-Johnson, M., and Cox, D · 2020
Earlier work this paper cites.
Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio, 2021
Chen, G., Chai, S., Wang, G., Du, J., Zhang, W.-Q., Weng, C., Su, D., Povey, D., Trmal, J., Zhang, J., Jin, M., Khudanpur, S., Watanabe, S., Zhao, S., Zou, W., Li, X., Yao, X., Wang, Y., Wang, Y., You, Z., and Yan, Z · 2021
Earlier work this paper cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A · 2021
Earlier work this paper cites.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech, 2021
Kim, J., Kong, J., and Son, J · 2021
Earlier work this paper cites.
Speech resynthesis from discrete disentangled self-supervised representations, 2021
Polyak, A., Adi, Y., Copet, J., Kharitonov, E., Lakhotia, K., Hsu, W.-N., Mohamed, A., and Dupoux, E · 2021
Cited alongside, same era.
Audiolm: a language modeling approach to audio generation, 2022
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Teboul, O., Grangier, D., Tagliasacchi, M., and Zeghidour, N · 2022
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision, 2022
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2022
Cited alongside, same era.
Soundstorm: Efficient parallel audio generation, 2023
Borsos, Z., Sharifi, M., Vincent, D., Kharitonov, E., Zeghidour, N., and Tagliasacchi, M · 2023
Cited alongside, same era.
Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone, 2023
Casanova, E., Weber, J., Shulby, C., Junior, A. C., Gölge, E., and Ponti, M. A · 2023
Cited alongside, same era.
Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers, 2023
Shen, K., Ju, Z., Tan, X., Liu, Y., Leng, Y., He, L., Qin, T., Zhao, S., and Bian, J · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding, 2023
Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Audiobox: Unified audio generation with natural language prompts, 2023
Vyas, A., Shi, B., Le, M., Tjandra, A., Wu, Y.-C., Guo, B., Zhang, J., Zhang, X., Adkins, R., Ngan, W., Wang, J., Cruz, I., Akula, B., Akinyemi, A., Ellis, B., Moritz, R., Yungster, Y., Rakotoarison, A., Tan, L., Summers, C., Wood, C., Lane, J., Williamson, M., and Hsu, W.-N · 2023
Later among the works it cites.
Neural codec language models are zero-shot text to speech synthesizers, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Voicebox: Text-guided multilingual universal speech generation at scale, 2023
Le, M., Vyas, A., Shi, B., Karrer, B., Sari, L., Moritz, R., Williamson, M., Manohar, V., Adi, Y., Mahadeokar, J., and Hsu, W.-N · 2023
Cited alongside, same era.
Flow matching for generative modeling, 2023
Lipman, Y., Chen, R. T. Q., Ben-Hamu, H., Nickel, M., and Le, M · 2023
Cited alongside, same era.
Generative pre-training for speech with flow matching, 2023
Liu, A. H., Le, M., Vyas, A., Shi, B., Tjandra, A., and Hsu, W.-N · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Speechtokenizer: Unified speech tokenizer for speech large language models, 2023b
Zhang, X., Zhang, D., Li, S., Zhou, Y., and Qiu, X
Cited in the paper.
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., He, L., Zhao, S., and Wei, F · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2023
Later among the works it cites.
Uniaudio: An audio foundation model toward universal audio generation, 2023
Yang, D., Tian, J., Tan, X., Huang, R., Liu, S., Chang, X., Shi, J., Zhao, S., Bian, J., Wu, X., Zhao, Z., Watanabe, S., and Meng, H · 2023
Later among the works it cites.
SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities
Zhang, D., Li, S., Zhang, X., Zhan, J., Wang, P., Zhou, Y., and Qiu, X · 2023
Later among the works it cites.
Matcha-tts: A fast tts architecture with conditional flow matching, 2024
Mehta, S., Tu, R., Beskow, J., Éva Székely, and Henter, G. E · 2024
Closest in time.