Fetching the paper…
Reading the bibliography…
Recent advances in Speech Large Language Models (Speech LLMs) have paved the way for unified architectures across diverse speech understanding tasks.
“Phone synchronous decoding with ctc lattice.,”
Zhehuai Chen, Wei Deng, Tao Xu, and Kai Yu, · 1927
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Phone synchronous speech recognition with ctc lattices,”
Zhehuai Chen, Yimeng Zhuang, Yanmin Qian, and Kai Yu, · 2017
Earlier work this paper cites.
“Label-synchronous decoding algorithm and its application in speech recognition,”
Zhehuai Chen, Wenlu Zheng, Yongbin You, Yanmin Qian, and Kai Yu, · 2019
Earlier work this paper cites.
“Modular end-to-end automatic speech recognition framework for acoustic-to-word model,”
Qi Liu, Zhehuai Chen, Hao Li, Mingkun Huang, Yizhou Lu, and Kai Yu, · 2020
Earlier work this paper cites.
“Optimizing alignment of speech and language latent spaces for end-to-end speech recognition and understanding,”
Wei Wang, Shuo Ren, Yao Qian, Shujie Liu, Yu Shi, Yanmin Qian, and Michael Zeng, · 2022
Earlier work this paper cites.
“Wenet 2.0: More productive end-to-end speech recognition toolkit,”
Binbin Zhang, Di Wu, Zhendong Peng, Xingchen Song, Zhuoyuan Yao, Hang Lv, Lei Xie, Chao Yang, Fuping Pan, and Jianwei Niu, · 2022
Earlier work this paper cites.
“Salmonn: Towards generic hearing abilities for large language models,”
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang, · 2023
Earlier work this paper cites.
“Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities,”
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu, · 2023
Cited alongside, same era.
“Slidespeech: A large-scale slide-enriched audio-visual corpus,” 2023
Haoxu Wang, Fan Yu, Xian Shi, Yuezhang Wang, Shiliang Zhang, and Ming Li, · 2023
Cited alongside, same era.
“Qwen2-audio technical report,”
Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al., · 2024
Cited alongside, same era.
“An embarrassingly simple approach for llm with strong asr capacity,”
Ziyang Ma, Guanrou Yang, Yifan Yang, Zhifu Gao, Jiaming Wang, Zhihao Du, Fan Yu, Qian Chen, Siqi Zheng, Shiliang Zhang, et al., · 2024
Cited alongside, same era.
“Seed-asr: Understanding diverse speech and contexts with llm-based speech recognition,”
“Osum: Advancing open speech understanding models with limited resources in academia,”
Xuelong Geng, Kun Wei, Qijie Shao, Shuiyun Liu, Zhennan Lin, Zhixian Zhao, Guojian Li, Wenjie Tian, Peikun Chen, Yangze Li, et al., · 2025
Closest in time.
“Low-resource domain adaptation for speech llms via text-only fine-tuning,”
Yangui Fang, Jing Peng, Xu Li, Yu Xi, Chengwei Zhang, Guohui Zhong, and Kai Yu, · 2025
Closest in time.
“Qwen2. 5-omni technical report,”
Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, et al., · 2025
Closest in time.
“Kimi-audio technical report,”
Ding Ding, Zeqian Ju, Yichong Leng, Songxiang Liu, Tong Liu, Zeyu Shang, Kai Shen, Wei Song, Xu Tan, Heyi Tang, et al., · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ye Bai, Jingping Chen, Jitong Chen, Wei Chen, Zhuo Chen, Chuang Ding, Linhao Dong, Qianqian Dong, Yujiao Du, Kepan Gao, et al., · 2024
Cited alongside, same era.
“On the landscape of spoken language models: A comprehensive survey,”
Siddhant Arora, Kai-Wei Chang, Chung-Ming Chien, Yifan Peng, Haibin Wu, Yossi Adi, Emmanuel Dupoux, Hung-Yi Lee, Karen Livescu, and Shinji Watanabe, · 2025
Cited alongside, same era.
“A survey on speech large language models for understanding,”
Jing Peng, Yucheng Wang, Bohan Li, Yiwei Guo, Hankun Wang, YanGui Fang, Yu Xi, Haoyu Li, Xu Li, Ke Zhang, Shuai Wang, and Kai Yu, · 2025
Cited alongside, same era.
Kai-Tuo Xu, Feng-Long Xie, Xu Tang, and Yao Hu, · 2025
Cited alongside, same era.
“Mmsu: A massive multi-task spoken language understanding and reasoning benchmark,” 2025
Dingdong Wang, Jincenzi Wu, Junan Li, Dongchao Yang, Xueyuan Chen, Tianhua Zhang, and Helen Meng, · 2025
Closest in time.
“Alignformer: Modality matching can achieve better zero-shot instruction-following speech-llm,”
Ruchao Fan, Bo Ren, Yuxuan Hu, Rui Zhao, Shujie Liu, and Jinyu Li, · 2025
Closest in time.
“Legoslm: Connecting llm with speech encoder using ctc posteriors,”
Rao Ma, Tongzhou Chen, Kartik Audhkhasi, and Bhuvana Ramabhadran, · 2025
Closest in time.