Fetching the paper…
Reading the bibliography…
We introduce Speech Information Retrieval (SIR), a new long-context task for Speech Large Language Models (Speech LLMs), and present SPIRAL, a 1,012-sample benchmark testing models' ability to extract critical details from approximately 90-second spoken inputs.
“Cortical oscillations and speech processing: emerging computational principles and operations”
Anne-Lise Giraud and David Poeppel · 2012
Earlier work this paper cites.
“Attention is All you Need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech”
Heiga Zen et al · 2019
Earlier work this paper cites.
“DREAM: A Challenge Data Set and Models for Dialogue-Based Reading Comprehension”
Kai Sun et al · 2019
Earlier work this paper cites.
“UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022”
Takaaki Saeki et al · 2022
Earlier work this paper cites.
“Transformer Quality in Linear Time”
Weizhe Hua, Zihang Dai, Hanxiao Liu and Quoc Le · 2022
Earlier work this paper cites.
“Learned token pruning for transformers”
Sehoon Kim et al · 2022
Earlier work this paper cites.
“SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities”
Dong Zhang et al · 2023
Earlier work this paper cites.
“On The Computational Complexity of Self-Attention”
Feyza Duman, Pruthuvi Wijewardena and Chinmay Hegde · 2023
Cited alongside, same era.
“Robust speech recognition via large-scale weak supervision”
Alec Radford et al · 2023
Cited alongside, same era.
“StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models”
Yinghao Li, Cong Han, Vinay Raghavan, Gavin Mischler and Nima Mesgarani · 2023
Cited alongside, same era.
“A Survey on Speech Large Language Models”
Jing Peng, Yucheng Wang, Yu Xi, Xv Li and Kai Yu · 2024
Cited alongside, same era.
“Qwen2-audio technical report”
Yunfei Chu et al · 2024
Cited alongside, same era.
“Llava-prumerge: Adaptive token reduction for efficient large multimodal models”
Yuzhang Shang, Mu Cai, Bingxin Xu, Yong Lee and Yan Yan · 2024
Closest in time.
“Listen, Think, and Understand”
Yuan Gong, Hongyin Luo, Alexander. Liu, Leonid Karlinsky and James. Glass · 2024
Closest in time.
“SALMONN: Towards Generic Hearing Abilities for Large Language Models”
Changli Tang et al · 2024
Closest in time.
“Speechverse: A large-scale generalizable audio language model”
Nilaksh Das et al · 2024
Closest in time.
“Llm inference unveiled: Survey and roofline model insights”
Zhihang Yuan et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
William Held et al · 2024
Cited alongside, same era.
“AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling”
Jun Zhan et al · 2024
Cited alongside, same era.
“Audiobench: A universal benchmark for audio large language models”
Bin Wang et al · 2024
Cited alongside, same era.
“WavLLM: Towards Robust and Adaptive Speech Large Language Model”
Shujie Hu et al · 2024
Closest in time.
“Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks”
Chien-yu Huang et al · 2025
Closest in time.
“MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark”
S Sakshi et al · 2025
Closest in time.