Fetching the paper…
Reading the bibliography…
We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants.
LibriSpeech: An ASR corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Neural discrete representation learning
van den Oord, A., Vinyals, O., and Kavukcuoglu, K · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Texygen: A benchmarking platform for text generation models
Zhu, Y., Lu, S., Zheng, L., Guo, J., Zhang, W., Wang, J., and Yu, Y · 2018
Earlier work this paper cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Reimers, N. and Gurevych, I · 2019
Earlier work this paper cites.
Common Voice: A massively-multilingual speech corpus
Ardila, R., Branson, M., Davis, K., Kohler, M., Meyer, J., Henretty, M., Morais, R., Saunders, L., Tyers, F. M., and Weber, G · 2020
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M · 2020
Earlier work this paper cites.
Libri-Light: A benchmark for ASR with limited or no supervision
Kahn, J., Riviere, M., Zheng, W., Kharitonov, E., Xu, Q., Mazaré, P.-E., Karadayi, J., Liptchinsky, V., Collobert, R., Fuegen, C., et al · 2020
Earlier work this paper cites.
Nguyen, T. A., de Seyssel, M., Rozé, P., Rivière, M., Kharitonov, E., Baevski, A., Dunbar, E., and Dupoux, E · 2020
Earlier work this paper cites.
MLS: A large-scale multilingual dataset for speech research
Pratap, V., Xu, Q., Sriram, A., Synnaeve, G., and Collobert, R · 2020
Earlier work this paper cites.
Variable-rate discrete representation learning
Dieleman, S., Nash, C., Engel, J. H., and Simonyan, K · 2021
Earlier work this paper cites.
Combining recurrent, convolutional, and continuous-time models with linear state space layers
Gu, A., Johnson, I., Goel, K., Saab, K., Dao, T., Rudra, A., and Ré, C · 2021
Earlier work this paper cites.
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W., Bolte, B., Tsai, Y. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A · 2021
Earlier work this paper cites.
On generative spoken language modeling from raw audio
Lakhotia, K., Kharitonov, E., Hsu, W.-N., Adi, Y., Polyak, A., Bolte, B., Nguyen, T.-A., Copet, J., Baevski, A., Mohamed, A., et al · 2021
Earlier work this paper cites.
Long Range Arena: A benchmark for efficient Transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2021
Earlier work this paper cites.
VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation
Wang, C., Rivière, M., Lee, A., Wu, A., Talnikar, C., Haziza, D., Williamson, M., Pino, J. M., and Dupoux, E · 2021
Earlier work this paper cites.
It’s raw! Audio generation with state-space models
Goel, K., Gu, A., Donahue, C., and Ré, C · 2022
Earlier work this paper cites.
Efficiently modeling long sequences with structured state spaces
Gu, A., Goel, K., and Ré, C · 2022
Cited alongside, same era.
Text-free prosody-aware generative spoken language modeling
Kharitonov, E., Lee, A., Polyak, A., Adi, Y., Copet, J., Lakhotia, K., Nguyen, T. A., Rivière, M., Mohamed, A., Dupoux, E., and Hsu, W · 2022
Cited alongside, same era.
A survey of evaluation metrics used for NLG systems
Sai, A. B., Mohankumar, A. K., and Khapra, M. M · 2022
Cited alongside, same era.
SoundStream: An end-to-end neural audio codec
Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M · 2022
Cited alongside, same era.
A closer look into using large language models for automatic evaluation
Chiang, D. C. and Lee, H · 2023
Cited alongside, same era.
Textually pretrained speech language models
Hassid, M., Remez, T., Nguyen, T. A., Gat, I., Conneau, A., Kreuk, F., Copet, J., Défossez, A., Synnaeve, G., Dupoux, E., Schwartz, R., and Adi, Y · 2023
Gemma: Open models based on gemini research and technology
Gemma Team, Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al · 2024
Closest in time.
Zamba: A compact 7B SSM hybrid model
Glorioso, P., Anthony, Q., Tokpanov, Y., Whittington, J., Pilault, J., Ibrahim, A., and Millidge, B · 2024
Closest in time.
Gecko: Versatile text embeddings distilled from large language models
Lee, J., Dai, Z., Ren, X., Chen, B., Cer, D., Cole, J. R., Hui, K., Boratko, M., Kapadia, R., Ding, W., et al · 2024
Closest in time.
Audio Mamba: Pretrained audio state space model for audio tagging
Lin, J. and Hu, H · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The impact of positional encoding on length generalization in Transformers
Kazemnejad, A., Padhi, I., Ramamurthy, K. N., Das, P., and Reddy, S · 2023
Cited alongside, same era.
Resurrecting recurrent neural networks for long sequences
Orvieto, A., Smith, S. L., Gu, A., Fernando, A., Gülçehre, Ç., Pascanu, R., and De, S · 2023
Cited alongside, same era.
AudioPaLM: A large language model that can speak and listen
Rubenstein, P. K., Asawaroengchai, C., Nguyen, D. D., Bapna, A., Borsos, Z., de Chaumont Quitry, F., Chen, P., Badawy, D. E., Han, W., Kharitonov, E., Muckenhirn, H., Padfield, D., Qin, J., Rozenberg, D., Sainath, T. N., Schalkwyk, J., Sharifi, M., Ramanovich, M. T., Tagliasacchi, M., Tudor, A., Velimirovic, M., Vincent, D., Yu, J., Wang, Y., Zayats, V., Zeghidour, N., Zhang, Y., Zhang, Z., Zilka, L., and Frank, C. H · 2023
Cited alongside, same era.
The next chapter: A study of large language models in storytelling
Xie, Z., Cohn, T., and Lau, J. H · 2023
Cited alongside, same era.
Judging LLM-as-a-judge with MT-bench and Chatbot Arena
Zheng, L., Chiang, W., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Cited alongside, same era.
RecurrentGemma: Moving past Transformers for efficient open language models
Botev, A., De, S., Smith, S. L., Fernando, A., Muraru, G., Haroun, R., Berrada, L., Pascanu, R., Sessa, P. G., Dadashi, R., Hussenot, L., Ferret, J., Girgin, S., Bachem, O., Andreev, A., Kenealy, K., Mesnard, T., Hardin, C., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., Tafti, P., Joulin, A., Fiedel, N., Senter, E., Chen, Y., Srinivasan, S., Desjardins, G., Budden, D., Doucet, A., Vikram, S., Paszke, A., Gale, T., Borgeaud, S., Chen, C., Brock, A., Paterson, A., Brennan, J., Risdal, M., Gundluru, R., Devanathan, N., Mooney, P., Chauhan, N., Culliton, P., Martins, L. G., Bandy, E., Huntsperger, D., Cameron, G., Zucker, A., Warkentin, T., Peran, L., Giang, M., Ghahramani, Z., Farabet, C., Kavukcuoglu, K., Hassabis, D., Hadsell, R., Teh, Y. W., and de Frietas, N · 2024
Cited alongside, same era.
Maiti, S., Peng, Y., Choi, S., Jung, J., Chang, X., and Watanabe, S · 2024
Closest in time.
Exploring the capability of Mamba in speech applications
Miyazaki, K., Masuyama, Y., and Murata, M · 2024
Closest in time.
Spoken question answering and speech continuation using spectrogram-powered LLM
Nachmani, E., Levkovitch, A., Hirsch, R., Salazar, J., Asawaroengchai, C., Mariooryad, S., Rivlin, E., Skerry-Ryan, R. J., and Ramanovich, M. T · 2024
Closest in time.
Patro, B. N. and Agneeswaran, V. S · 2024
Closest in time.
SSAMBA: Self-supervised audio representation learning with Mamba state space model
Shams, S., Dindar, S. S., Jiang, X., and Mesgarani, N · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M. H. M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Closest in time.
STAB: Speech tokenizer assessment benchmark
Vashishth, S., Singh, H., Bharadwaj, S., Ganapathy, S., Asawaroengchai, C., Audhkhasi, K., Rosenberg, A., Bapna, A., and Ramabhadran, B · 2024
Closest in time.
SpeechAgents: Human-communication simulation with multi-modal multi-agent systems
Zhang, D., Li, Z., Wang, P., Zhang, X., Zhou, Y., and Qiu, X · 2024
Closest in time.
Gemini Team, Comanici, G., Bieber, E., Schaekermann, M., Pasupat, I., Sachdeva, N., Dhillon, I., Blistein, M., Ram, O., Zhang, D., Rosen, E., Marris, L., Petulla, S., Gaffney, C., Aharoni, A., Lintz, N., Pais, T. C., Jacobsson, H., Szpektor, I., Jiang, N.-J., Haridasan, K., Omran, A., Saunshi, N., Bahri, D., Mishra, G., et al · 2025
Closest in time.
Jamba: Hybrid Transformer-Mamba language models
Lenz, B., Lieber, O., Arazi, A., Bergman, A., Manevich, A., Peleg, B., Aviram, B., Almagor, C., Fridman, C., Padnos, D., Gissin, D., Jannai, D., Muhlgay, D., Zimberg, D., Gerber, E. M., Dolev, E., Krakovsky, E., Safahi, E., Schwartz, E., Cohen, G., and et al · 2025
Closest in time.
SpiRit-LM: Interleaved spoken and written language model
Nguyen, T. A., Muller, B., Yu, B., Costa-Jussa, M. R., Elbayad, M., Popuri, S., Ropers, C., Duquenne, P.-A., Algayres, R., Mavlyutov, R., et al · 2025
Closest in time.
HALL-E: Hierarchical neural codec language model for minute-long zero-shot text-to-speech synthesis
Nishimura, Y., Hirose, T., Ohi, M., Nakayama, H., and Inoue, N · 2025
Closest in time.