Fetching the paper…
Reading the bibliography…
Speech language models have recently demonstrated great potential as universal speech processing systems.
“The lj speech dataset,” https://keithito.com/LJ-Speech-Dataset/ , 2017
Keith Ito and Linda Johnson, · 2017
Earlier work this paper cites.
“Montreal forced aligner: Trainable text-speech alignment using kaldi.,”
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger, · 2017
Earlier work this paper cites.
“Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92),”
Junichi Yamagishi, Christophe Veaux, and Kirsten MacDonald, · 2019
Earlier work this paper cites.
“On generative spoken language modeling from raw audio,”
Kushal Lakhotia et al., · 2021
Earlier work this paper cites.
“The zero resource speech challenge 2021: Spoken language modelling,”
Ewan Dunbar et al., · 2021
Earlier work this paper cites.
“Text-free prosody-aware generative spoken language modeling,”
Eugene Kharitonov et al., · 2021
Earlier work this paper cites.
“Soundstream: An end-to-end neural audio codec,”
Neil Zeghidour, Alejandro Luebs, Ahmed Omran, Jan Skoglund, and Marco Tagliasacchi, · 2021
Earlier work this paper cites.
“Speech resynthesis from discrete disentangled self-supervised representations,”
Adam Polyak et al., · 2021
Earlier work this paper cites.
“Fsd50k: an open dataset of human-labeled sound events,”
Eduardo Fonseca et al., · 2021
Earlier work this paper cites.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu et al., · 2021
Earlier work this paper cites.
“It’s raw! audio generation with state-space models,”
Karan Goel, Albert Gu, Chris Donahue, and Christopher Ré, · 2022
Earlier work this paper cites.
“High fidelity neural audio compression,”
Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi, · 2022
Earlier work this paper cites.
“Salmonn: Towards generic hearing abilities for large language models,”
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang, · 2023
Earlier work this paper cites.
“Generative spoken dialogue language modeling,”
Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoit Sagot, Abdelrahman Mohamed, et al., · 2023
Earlier work this paper cites.
“Prosaudit, a prosodic benchmark for self-supervised speech models,”
Maureen de Seyssel et al., · 2023
Earlier work this paper cites.
“Audiolm: a language modeling approach to audio generation,”
Zalán Borsos et al., · 2023
Earlier work this paper cites.
“Generative spoken language model based on continuous word-sized audio tokens,”
Robin Algayres et al., · 2023
Cited alongside, same era.
“Spoken question answering and speech continuation using spectrogram-powered llm,”
Eliya Nachmani et al., · 2023
Cited alongside, same era.
“Speaking style conversion in the waveform domain using discrete self-supervised units,”
Gallil Maimon and Yossi Adi, · 2023
Cited alongside, same era.
“Analysing discrete self supervised speech representation for spoken language modeling,”
Amitay Sicherman and Yossi Adi, · 2023
Cited alongside, same era.
“Lauragpt: Listen, attend, understand, and regenerate audio with gpt,”
Qian Chen, Yunfei Chu, Zhifu Gao, Zerui Li, Kai Hu, Xiaohuan Zhou, Jin Xu, Ziyang Ma, Wen Wang, Siqi Zheng, et al., · 2023
“Speechalign: Aligning speech generation to human preferences,”
Dong Zhang, Zhaowei Li, Shimin Li, Xin Zhang, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu, · 2024
Closest in time.
“Audiochatllama: Towards general-purpose speech abilities for llms,”
Yassir Fathullah et al., · 2024
Closest in time.
“Beyond the turn-based game: Enabling real-time conversations with duplex models,”
Xinrong Zhang, Yingfa Chen, Shengding Hu, Xu Han, Zihang Xu, Yuanwei Xu, Weilin Zhao, Maosong Sun, and Zhiyuan Liu, · 2024
Closest in time.
“Modeling real-time interactive conversations as timed diarized transcripts,”
Garrett Tanzer, Gustaf Ahdritz, and Luke Melas-Kyriazi, · 2024
Closest in time.
“A full-duplex speech dialogue scheme based on large language models,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Audiopalm: A large language model that can speak and listen,”
Paul K Rubenstein et al., · 2023
Cited alongside, same era.
“Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models,”
Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou, · 2023
Cited alongside, same era.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford et al., · 2023
Cited alongside, same era.
“Expresso: A benchmark and analysis of discrete expressive speech resynthesis,”
Tu Anh Nguyen et al., · 2023
Cited alongside, same era.
“Llama 2: Open foundation and fine-tuned chat models,”
Hugo Touvron et al., · 2023
Cited alongside, same era.
“Speechprompt: Prompting speech language models for speech processing tasks,”
Kai-Wei Chang et al., · 2024
Cited alongside, same era.
“Textually pretrained speech language models,”
Michael Hassid et al., · 2024
Cited alongside, same era.
Peng Wang, Songshuo Lu, Yaohua Tang, Sijie Yan, Yuanjun Xiong, and Wei Xia, · 2024
Closest in time.
“Language model can listen while speaking,”
Ziyang Ma, Yakun Song, Chenpeng Du, Jian Cong, Zhuo Chen, Yuping Wang, Yuxuan Wang, and Xie Chen, · 2024
Closest in time.
“Sd-eval: A benchmark dataset for spoken dialogue understanding beyond words,”
Junyi Ao, Yuancheng Wang, Xiaohai Tian, Dekun Chen, Jun Zhang, Lu Lu, Yuxuan Wang, Haizhou Li, and Zhizheng Wu, · 2024
Closest in time.
“Nast: Noise aware speech tokenization for speech language models,”
Shoval Messica and Yossi Adi, · 2024
Closest in time.
“How should we extract discrete audio tokens from self-supervised models?,”
Pooneh Mousavi, Jarod Duret, Salah Zaiem, Luca Della Libera, Artem Ploujnikov, Cem Subakan, and Mirco Ravanelli, · 2024
Closest in time.
“Audiobench: A universal benchmark for audio large language models,”
Bin Wang et al., · 2024
Closest in time.
“Dynamic-superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech,”
Chien-yu Huang et al., · 2024
Closest in time.
“Air-bench: Benchmarking large audio-language models via generative comprehension,”
Qian Yang et al., · 2024
Closest in time.
“Echothief,” http://www.echothief.com ,
Christopher Warren, · 2024
Closest in time.
“Azure tts,” https://learn.microsoft.com/en-us/azure/ai-services/speech-service/text-to-speech ,
Microsoft, · 2024
Closest in time.
“Last: Language model aware speech tokenization,”
Arnon Turetzky and Yossi Adi, · 2024
Closest in time.
“Gpt-4o,” https://openai.com/index/gpt-4o-system-card/ ,
Open-AI, · 2024
Closest in time.