Fetching the paper…
Reading the bibliography…
Audio-language models have shown promising results in various sound understanding tasks, yet they remain limited in their ability to reason over the fine-grained semantics of sound.
“What in the World Do We Hear?: An Ecological Approach to Auditory Event Perception”
William. Gaver · 1993
Earlier work this paper cites.
“Similarity and Categorization of Environmental Sounds”
Brian Gygi, Gary Kidd and Charles Watson · 2007
Earlier work this paper cites.
“Freesound Technical Demo”
Frederic Font, Gerard Roma and Xavier Serra · 2013
Earlier work this paper cites.
“Learning Deep Features for Scene Recognition Using Places Database”
Bolei Zhou et al · 2014
Earlier work this paper cites.
“Audio Set: An Ontology and Human-Labeled Dataset for Audio Events”
Jort. Gemmeke et al · 2017
Earlier work this paper cites.
“Sound Categories: Category Formation and Evidence-Based Taxonomies”
Oliver Bones, Trevor Cox and William Davies · 2018
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“AudioCaps: Generating Captions for Audios in The Wild”
Chris Kim, Byeongchang Kim, Hyunmin Lee and Gunhee Kim · 2019
Earlier work this paper cites.
“Decoupled Weight Decay Regularization”
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
“Clotho: An Audio Captioning Dataset”
Konstantinos Drossos, Samuel Lipping and Tuomas Virtanen · 2020
Earlier work this paper cites.
“ZeRO: Memory Optimizations toward Training Trillion Parameter Models”
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase and Yuxiong He · 2020
Earlier work this paper cites.
“Training Verifiers to Solve Math Word Problems”
Karl Cobbe et al · 2021
Earlier work this paper cites.
“AST: Audio Spectrogram Transformer”
Yuan Gong, Yu-An Chung and James Glass · 2021
Earlier work this paper cites.
“Learning Transferable Visual Models From Natural Language Supervision”
Alec Radford et al · 2021
Earlier work this paper cites.
“What Do We Mean with Sound Semantics, Exactly? A Survey of Taxonomies and Ontologies of Everyday Sounds”
Bruno. Giordano et al · 2022
Earlier work this paper cites.
“LoRA: Low-rank Adaptation of Large Language Models”
Edward Hu et al · 2022
Earlier work this paper cites.
“BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation”
Junnan Li, Dongxu Li, Caiming Xiong and Steven Hoi · 2022
Earlier work this paper cites.
“Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering”
Samuel Lipping, Parthasaarathy Sudarsanam, Konstantinos Drossos and Tuomas Virtanen · 2022
Earlier work this paper cites.
“BEATs: Audio Pre-Training with Acoustic Tokenizers”
Sanyuan Chen et al · 2023
Earlier work this paper cites.
“Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models”
Yunfei Chu et al · 2023
Earlier work this paper cites.
“LP-MusicCaps: LLM-Based Pseudo Music Captioning”
Seungheon Doh, Keunwoo Choi, Jongpil Lee and Juhan Nam · 2023
Earlier work this paper cites.
“A Dataset for Audio-Visual Sound Event Detection in Movies”
Rajat Hebbar et al · 2023
Earlier work this paper cites.
“Teacher-Student Architecture for Knowledge Distillation: A Survey”
Chengming Hu et al · 2023
Earlier work this paper cites.
“Efficient Memory Management for Large Language Model Serving with PagedAttention”
Woosuk Kwon et al · 2023
Cited alongside, same era.
Junnan Li, Dongxu Li, Silvio Savarese and Steven.. Hoi · 2023
Cited alongside, same era.
“ACES: Evaluating Automated Audio Captioning Models on the Semantics of Sounds”
Gijs Wijngaard, Elia Formisano, Bruno. Giordano and Michel Dumontier · 2023
Cited alongside, same era.
“Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation”
Yusong Wu et al · 2023
Cited alongside, same era.
Marah Abdin et al · 2024
Cited alongside, same era.
“CogVLM: Visual Expert for Pretrained Language Models”
Weihan Wang et al · 2024
Later among the works it cites.
“MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models”
Benno Weck et al · 2024
Later among the works it cites.
“Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions”
Yi Yuan et al · 2024
Later among the works it cites.
“Video Instruction Tuning With Synthetic Data”
Yuanhan Zhang et al · 2024
Later among the works it cites.
“DETRs Beat YOLOs on Real-time Object Detection”
Yian Zhao et al · 2024
Later among the works it cites.
“L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“AudioSetCaps: Enriched Audio Captioning Dataset Generation Using Large Audio Language Models”
Jisheng Bai et al · 2024
Cited alongside, same era.
“Qwen2-Audio Technical Report”
Yunfei Chu et al · 2024
Cited alongside, same era.
“Gemini 1.5: Unlocking Multimodal Understanding across Millions of Tokens of Context”
Gemini Team et al · 2024
Cited alongside, same era.
“GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities”
Sreyan Ghosh et al · 2024
Cited alongside, same era.
“WavLLM: Towards Robust and Adaptive Speech Large Language Model”
Shujie Hu et al · 2024
Cited alongside, same era.
Albert. Jiang et al · 2024
Cited alongside, same era.
“Improving Text-To-Audio Models with Synthetic Captions”
Zhifeng Kong et al · 2024
Cited alongside, same era.
Pranjal Aggarwal and Sean Welleck · 2025
Closest in time.
“Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities”
Gheorghe Comanici et al · 2025
Closest in time.
“Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn’t”
Quy-Anh Dang and Chris Ngo · 2025
Closest in time.
“Mellow: A Small Audio Language Model for Reasoning”
Soham Deshmukh, Satvik Dixit, Rita Singh and Bhiksha Raj · 2025
Closest in time.
“XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models”
Yixin Dong et al · 2025
Closest in time.
Sreyan Ghosh et al · 2025
Closest in time.
“DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”
Daya Guo et al · 2025
Closest in time.
“MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models”
Yingxu He et al · 2025
Closest in time.
KimiTeam et al · 2025
Closest in time.
Gang Li et al · 2025
Closest in time.
“Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model”
Ziyang Ma et al · 2025
Closest in time.
“S1: Simple Test-Time Scaling”
Niklas Muennighoff et al · 2025
Closest in time.
“Expanding RL with Verifiable Rewards Across Diverse Domains”
Yi Su et al · 2025
Closest in time.
“AudioBench: A Universal Benchmark for Audio Large Language Models”
Bin Wang et al · 2025
Closest in time.
“Audio-Language Datasets of Scenes and Events: A Survey”
Gijs Wijngaard, Elia Formisano, Michele Esposito and Michel Dumontier · 2025
Closest in time.
“Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models”
Zhifei Xie et al · 2025
Closest in time.
“Qwen2.5-Omni Technical Report”
Jin Xu et al · 2025
Closest in time.
An Yang et al · 2025
Closest in time.