Fetching the paper…
Reading the bibliography…
Current speech large language models build upon discrete speech representations, which can be categorized into semantic tokens and acoustic tokens.
Visqol: The virtual speech quality objective listener
Hines, A., Skoglund, J., Kokaram, A., and Harte, N · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
Method for the subjective assessment of intermediate quality level of audio systems
Series, B · 2014
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Layer normalization, 2016
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus), 2016
Clevert, D.-A., Unterthiner, T., and Hochreiter, S · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Earlier work this paper cites.
Autovc: Zero-shot voice style transfer with only autoencoder loss, 2019
Qian, K., Zhang, Y., Chang, S., Yang, X., and Hasegawa-Johnson, M · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M · 2020
Earlier work this paper cites.
Club: A contrastive log-ratio upper bound of mutual information, 2020
Cheng, P., Hao, W., Dai, S., Liu, J., Gan, Z., and Carin, L · 2020
Earlier work this paper cites.
Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Kong, J., Kim, J., and Bae, J · 2020
Earlier work this paper cites.
MLS: A large-scale multilingual dataset for speech research
Pratap, V., Xu, Q., Sriram, A., Synnaeve, G., and Collobert, R · 2020
Earlier work this paper cites.
Unsupervised speech decomposition via triple information bottleneck
Qian, K., Zhang, Y., Chang, S., Hasegawa-Johnson, M., and Cox, D · 2020
Earlier work this paper cites.
One-shot voice conversion by vector quantization
Wu, D.-Y. and Lee, H.-y · 2020
Cited alongside, same era.
W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training, 2021
Chung, Y.-A., Zhang, Y., Han, W., Chiu, C.-C., Qin, J., Pang, R., and Wu, Y · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A · 2021
Cited alongside, same era.
On generative spoken language modeling from raw audio
Lakhotia, K., Kharitonov, E., Hsu, W.-N., Adi, Y., Polyak, A., Bolte, B., Nguyen, T.-A., Copet, J., Baevski, A., Mohamed, A., et al · 2021
Cited alongside, same era.
Speech resynthesis from discrete disentangled self-supervised representations, 2021
Polyak, A., Adi, Y., Copet, J., Kharitonov, E., Lakhotia, K., Hsu, W.-N., Mohamed, A., and Dupoux, E · 2021
Cited alongside, same era.
Self-supervised speech representation learning: A review
Mohamed, A., yi Lee, H., Borgholt, L., Havtorn, J. D., Edin, J., Igel, C., Kirchhoff, K., Li, S.-W., Livescu, K., Maaloe, L., Sainath, T. N., and Watanabe, S · 2022
Later among the works it cites.
Polyvoice: Language models for speech to speech translation, 2023
Dong, Q., Huang, Z., Tian, Q., Xu, C., Ko, T., Zhao, Y., Feng, S., Li, T., Wang, K., Cheng, X., Yue, F., Bai, Y., Chen, X., Lu, L., Ma, Z., Wang, Y., Wang, M., and Wang, Y · 2023
Closest in time.
Textually pretrained speech language models, 2023
Hassid, M., Remez, T., Nguyen, T. A., Gat, I., Conneau, A., Kreuk, F., Copet, J., Defossez, A., Synnaeve, G., Dupoux, E., Schwartz, R., and Adi, Y · 2023
Closest in time.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision, 2023
Kharitonov, E., Vincent, D., Borsos, Z., Marinier, R., Girgin, S., Pietquin, O., Sharifi, M., Tagliasacchi, M., and Zeghidour, N · 2023
Closest in time.
Unifyspeech: A unified framework for zero-shot text-to-speech and voice conversion, 2023
Liu, H., Wang, T., Fu, R., Yi, J., Wen, Z., and Tao, J · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aishell-3: A multi-speaker mandarin tts corpus and the baselines, 2021
Shi, Y., Bu, H., Xu, X., Zhang, S., and Li, M · 2021
Cited alongside, same era.
Soundstream: An end-to-end neural audio codec, 2021
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2021
Cited alongside, same era.
Audiolm: a language modeling approach to audio generation, 2022
Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Teboul, O., Grangier, D., Tagliasacchi, M., and Zeghidour, N · 2022
Cited alongside, same era.
YourTTS: Towards zero-shot multi-speaker TTS and zero-shot voice conversion for everyone
Casanova, E., Weber, J., Shulby, C. D., Junior, A. C., Gölge, E., and Ponti, M. A · 2022
Cited alongside, same era.
Distilhubert: Speech representation learning by layer-wise distillation of hidden-unit bert
Chang, H.-J., Yang, S.-w., and Lee, H.-y · 2022
Cited alongside, same era.
WavLM: Large-scale self-supervised pre-training for full stack speech processing
Chen, S., Wang, C., Chen, Z., Wu, Y., Liu, S., Chen, Z., Li, J., Kanda, N., Yoshioka, T., Xiao, X., Wu, J., Zhou, L., Ren, S., Qian, Y., Qian, Y., Wu, J., Zeng, M., Yu, X., and Wei, F · 2022
Cited alongside, same era.
High fidelity neural audio compression, 2022
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2022
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Robust speech recognition via large-scale weak supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2023
Closest in time.
Audiopalm: A large language model that can speak and listen, 2023
Rubenstein, P. K., Asawaroengchai, C., Nguyen, D. D., Bapna, A., Borsos, Z., de Chaumont Quitry, F., Chen, P., Badawy, D. E., Han, W., Kharitonov, E., Muckenhirn, H., Padfield, D., Qin, J., Rozenberg, D., Sainath, T., Schalkwyk, J., Sharifi, M., Ramanovich, M. T., Tagliasacchi, M., Tudor, A., Velimirović, M., Vincent, D., Yu, J., Wang, Y., Zayats, V., Zeghidour, N., Zhang, Y., Zhang, Z., Zilka, L., and Frank, C · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.
Neural codec language models are zero-shot text to speech synthesizers, 2023
Wang, C., Chen, S., Wu, Y., Zhang, Z., Zhou, L., Liu, S., Chen, Z., Liu, Y., Wang, H., Li, J., He, L., Zhao, S., and Wei, F · 2023
Closest in time.
Hifi-codec: Group-residual vector quantization for high fidelity audio codec, 2023
Yang, D., Liu, S., Huang, R., Tian, J., Weng, C., and Zou, Y · 2023
Closest in time.
Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities, 2023
Zhang, D., Li, S., Zhang, X., Zhan, J., Wang, P., Zhou, Y., and Qiu, X · 2023
Closest in time.