Fetching the paper…
Reading the bibliography…
Recent advancements in audio language models have underscored the pivotal role of audio tokenization, which converts audio signals into discrete tokens, thereby facilitating the application of language model architectures to the audio domain.
The million song dataset
Bertin-Mahieux, T., Ellis, D. P., Whitman, B., and Lamere, P · 2011
Earlier work this paper cites.
Emovo corpus: an italian emotional speech database
Costantini, G., Iaderola, I., Paoloni, A., Todisco, M., et al · 2014
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Piczak, K. J · 2015
Earlier work this paper cites.
Deep convolutional networks on the pitch spiral for musical instrument recognition
Lostanlen, V. and Cella, C.-E · 2016
Earlier work this paper cites.
Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Veaux, C., Yamagishi, J., MacDonald, K., et al · 2017
Earlier work this paper cites.
Audiocaps: Generating captions for audios in the wild
Kim, C. D., Kim, B., Lee, H., and Kim, G · 2019
Earlier work this paper cites.
Libritts: A corpus derived from librispeech for text-to-speech
Zen, H., Dang, V., Clark, R., Zhang, Y., Weiss, R. J., Jia, Y., Chen, Z., and Wu, Y · 2019
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M · 2020
Earlier work this paper cites.
Clotho: An audio captioning dataset
Drossos, K., Lipping, S., and Virtanen, T · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Ho, J., Jain, A., and Abbeel, P · 2020
Earlier work this paper cites.
Mls: A large-scale multilingual dataset for speech research
Pratap, V., Xu, Q., Sriram, A., Synnaeve, G., and Collobert, R · 2020
Earlier work this paper cites.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Superb: Speech processing universal performance benchmark
Yang, S.-w., Chi, P.-H., Chuang, Y.-S., Lai, C.-I. J., Lakhotia, K., Lin, Y. Y., Liu, A. T., Shi, J., Chang, X., Lin, G.-T., et al · 2021
Earlier work this paper cites.
Soundstream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2021
Earlier work this paper cites.
wav2tok: Deep sequence tokenizer for audio retrieval
Banerjee, A. and Arora, V · 2022
Earlier work this paper cites.
High fidelity neural audio compression
Défossez, A., Copet, J., Synnaeve, G., and Adi, Y · 2022
Earlier work this paper cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Earlier work this paper cites.
Masked autoencoders that listen
Huang, P.-Y., Xu, H., Li, J., Baevski, A., Auli, M., Galuba, W., Metze, F., and Feichtenhofer, C · 2022
Cited alongside, same era.
Audiogen: Textually guided audio generation
Kreuk, F., Synnaeve, G., Polyak, A., Singer, U., Défossez, A., Copet, J., Parikh, D., Taigman, Y., and Adi, Y · 2022
Cited alongside, same era.
Flow matching for generative modeling
Lipman, Y., Chen, R. T., Ben-Hamu, H., Nickel, M., and Le, M · 2022
Cited alongside, same era.
Dnsmos p. 835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors
Reddy, C. K., Gopal, V., and Cutler, R · 2022
Cited alongside, same era.
Utmos: Utokyo-sarulab system for voicemos challenge 2022
Saeki, T., Xin, D., Nakata, W., Koriyama, T., Takamichi, S., and Saruwatari, H · 2022
Moshi: a speech-text foundation model for real-time dialogue
Défossez, A., Mazaré, L., Orsini, M., Royer, A., Pérez, P., Jégou, H., Grave, E., and Zeghidour, N · 2024
Later among the works it cites.
Du, Z., Chen, Q., Zhang, S., Hu, K., Lu, H., Yang, Y., Hu, H., Zheng, S., Gu, Y., Ma, Z., et al · 2024
Later among the works it cites.
Wavtokenizer: an efficient acoustic discrete codec tokenizer for audio language modeling
Ji, S., Jiang, Z., Wang, W., Chen, Y., Fang, M., Zuo, J., Yang, Q., Cheng, X., Wang, Z., Li, R., et al · 2024
Later among the works it cites.
Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Ju, Z., Wang, Y., Shen, K., Tan, X., Xin, D., Yang, D., Liu, Y., Leng, Y., Song, K., Tang, S., et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Musiclm: Generating music from text
Agostinelli, A., Denk, T. I., Borsos, Z., Engel, J., Verzetti, M., Caillon, A., Huang, Q., Jansen, A., Roberts, A., Tagliasacchi, M., et al · 2023
Cited alongside, same era.
Simple and controllable music generation
Copet, J., Kreuk, F., Gat, I., Remez, T., Kant, D., Synnaeve, G., Adi, Y., and Défossez, A · 2023
Cited alongside, same era.
Lp-musiccaps: Llm-based pseudo music captioning
Doh, S., Choi, K., Lee, J., and Nam, J · 2023
Cited alongside, same era.
Boosting large language model for speech synthesis: An empirical study
Hao, H., Zhou, L., Liu, S., Li, J., Hu, S., Wang, R., and Wei, F · 2023
Cited alongside, same era.
Speak, read and prompt: High-fidelity text-to-speech with minimal supervision
Kharitonov, E., Vincent, D., Borsos, Z., Marinier, R., Girgin, S., Pietquin, O., Sharifi, M., Tagliasacchi, M., and Zeghidour, N · 2023
Cited alongside, same era.
High-fidelity audio compression with improved RVQGAN
Kumar, R., Seetharaman, P., Luebs, A., Kumar, I., and Kumar, K · 2023
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Li, J., Li, D., Savarese, S., and Hoi, S · 2023
Cited alongside, same era.
Libriheavy: a 50,000 hours asr corpus with punctuation casing and context
Kang, W., Yang, X., Yao, Z., Kuang, F., Yang, Y., Guo, L., Lin, L., and Povey, D · 2024
Later among the works it cites.
Benchmarking representations for speech, music, and acoustic events
La Quatra, M., Koudounas, A., Vaiani, L., Baralis, E., Cagliero, L., Garza, P., and Siniscalchi, S. M · 2024
Later among the works it cites.
Single-codec: Single-codebook speech codec towards high-performance speech generation
Li, H., Xue, L., Guo, H., Zhu, X., Lv, Y., Xie, L., Chen, Y., Yin, H., and Li, Z · 2024
Later among the works it cites.
Semanticodec: An ultra low bitrate semantic audio codec for general sound
Liu, H., Xu, X., Yuan, Y., Wu, M., Wang, W., and Plumbley, M. D · 2024
Later among the works it cites.
Scaling transformers for low-bitrate high-quality speech coding
Parker, J. D., Smirnov, A., Pons, J., Carr, C., Zukowski, Z., Evans, Z., and Liu, X · 2024
Later among the works it cites.
SALMONN: Towards generic hearing abilities for large language models
Tang, C., Yu, W., Sun, G., Chen, X., Tan, T., Li, W., Lu, L., MA, Z., and Zhang, C · 2024
Later among the works it cites.
Spoken-term discovery using discrete speech units
van Niekerk, B., Zaïdi, J., Carbonneau, M.-A., and Kamper, H · 2024
Later among the works it cites.
Consistent and relevant: Rethink the query embedding in general sound separation
Wang, Y., Chen, H., Yang, D., Yu, J., Weng, C., Wu, Z., and Meng, H · 2024
Later among the works it cites.
Ts3-codec: Transformer-based simple streaming single codec
Wu, H., Kanda, N., Eskimez, S. E., and Li, J · 2024
Later among the works it cites.
An image is worth 32 tokens for reconstruction and generation
Yu, Q., Weber, M., Deng, X., Shen, X., Cremers, D., and Chen, L.-C · 2024
Later among the works it cites.
Glm-4-voice: Towards intelligent and human-like end-to-end spoken chatbot
Zeng, A., Du, Z., Liu, M., Wang, K., Jiang, S., Zhao, L., Dong, Y., and Tang, J · 2024
Later among the works it cites.
Scaling the codebook size of vqgan to 100,000 with a utilization rate of 99%
Zhu, L., Wei, F., Lu, Y., and Chen, D · 2024
Later among the works it cites.
Spirit-lm: Interleaved spoken and written language model
Nguyen, T. A., Muller, B., Yu, B., Costa-Jussa, M. R., Elbayad, M., Popuri, S., Ropers, C., Duquenne, P.-A., Algayres, R., Mavlyutov, R., et al · 2025
Closest in time.
Unisep: Universal target audio separation with language models at scale
Wang, Y., Chen, H., Yang, D., Li, W., Luo, D., Li, G., Yang, S., Wu, Z., Meng, H., and Wu, X · 2025
Closest in time.