Fetching the paper…
Reading the bibliography…
Large Audio-Language Models (LALMs) have demonstrated remarkable performance in tasks involving audio perception and understanding, such as speech recognition and audio captioning.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” Proc. Neurips , 2022
2022
Earlier work this paper cites.
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” Proc. Neurips , 2022
2022
Earlier work this paper cites.
J. Liu, A. Liu, X. Lu, S. Welleck, P. West, R. L. Bras, Y. Choi, and H. Hajishirzi, “Generated knowledge prompting for commonsense reasoning,” Proc. ACL , 2022
2022
Earlier work this paper cites.
Y. Gong, A. H. Liu, H. Luo, L. Karlinsky, and J. Glass, “Joint audio and speech understanding,” Proc. ASRU , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Li, Y. Wu, J. Li, and S. Liu, “Prompting large language models for zero-shot domain adaptation in speech recognition,” Proc. ASRU , 2023
2023
Earlier work this paper cites.
J. Wu, Y. Gaur, Z. Chen, L. Zhou, Y. Zhu, T. Wang, J. Li, S. Liu, B. Ren, L. Liu et al. , “On decoder-only architecture for speech-to-text and large language model integration,” Proc. ASRU , 2023
2023
Earlier work this paper cites.
S.-L. Wu, X. Chang, G. Wichern, J.-w. Jung, F. Germain, J. Le Roux, and S. Watanabe, “BEATs-based audio captioning model with INSTRUCTOR embedding supervision and ChatGPT mix-up,” Proc. DCASE Challenge , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Z. Zhang, A. Zhang, M. Li, and A. Smola, “Automatic chain of thought prompting in large language models,” Proc. ICLR , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou, “Self-consistency improves chain of thought reasoning in language models,” Proc. ICLR , 2023
2023
Earlier work this paper cites.
C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “SALMONN: Towards generic hearing abilities for large language models,” Proc. ICLR , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
S. Ghosh, S. Kumar, A. Seth, C. K. R. Evuru, U. Tyagi, S. Sakshi, O. Nieto, R. Duraiswami, and D. Manocha, “GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities,” Proc. EMNLP , 2024
2024
Cited alongside, same era.
Z. Kong, A. Goel, R. Badlani, W. Ping, R. Valle, and B. Catanzaro, “Audio flamingo: A novel audio language model with few-shot learning and dialogue abilities,” Proc. ICML , 2024
2024
Cited alongside, same era.
Q. Yang, J. Xu, W. Liu, Y. Chu, Z. Jiang, X. Zhou, Y. Leng, Y. Lv, Z. Zhao, C. Zhou et al. , “Air-bench: Benchmarking large audio-language models via generative comprehension,” Proc. ACL , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
Y. Fathullah, C. Wu, E. Lakomkin, J. Jia, Y. Shangguan, K. Li, J. Guo, W. Xiong, J. Mahadeokar, O. Kalinli et al. , “Prompting large language models with speech recognition abilities,” Proc. ICASSP , 2024
2024
Cited alongside, same era.
W. Yu, C. Tang, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “Connecting speech encoder and large language model for ASR,” Proc. ICASSP , 2024
2024
Cited alongside, same era.
G. Yang, Z. Ma, F. Yu, Z. Gao, S. Zhang, and X. Chen, “MaLa-ASR: Multimedia-Assisted LLM-Based ASR,” Proc. Interspeech , 2024
2024
Cited alongside, same era.
G. Yang, Z. Ma, Z. Gao, S. Zhang, and X. Chen, “CTC-Assisted LLM-Based Contextual ASR,” Proc. SLT , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
J. Liu, G. Li, J. Zhang, C. Liu, H. Dinkel, Y. Wang, Z. Yan, Y. Wang, and B. Wang, “Leveraging ced encoder and large language models for automated audio captioning,” Proc. DCASE Challenge , 2024
2024
Cited alongside, same era.
2024
Later among the works it cites.
Z. Chu, J. Chen, Q. Chen, W. Yu, T. He, H. Wang, W. Peng, M. Liu, B. Qin, and T. Liu, “Navigate through enigmatic labyrinth a survey of chain of thought reasoning: Advances, frontiers and future,” Proc. ACL , 2024
2024
Later among the works it cites.
Z. Zhang, A. Zhang, M. Li, H. Zhao, G. Karypis, and A. Smola, “Multimodal chain-of-thought reasoning in language models,” Proc. TMLR , 2024
2024
Later among the works it cites.
Q. Chen, L. Qin, J. Zhang, Z. Chen, X. Xu, and W. Che, “M 3 CoT: A novel benchmark for multi-domain multi-step multi-modal chain-of-thought,” Proc. ACL , 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
W. Chen, Z. Ma, X. Li, X. Xu, Y. Liang, Z. Zheng, K. Yu, and X. Chen, “SLAM-AAC: Enhancing audio captioning with paraphrasing augmentation and CLAP-Refine through LLMs,” Proc. ICASSP , 2025
2025
Closest in time.
X. Li, W. Chen, Z. Ma, X. Xu, Y. Liang, Z. Zheng, Q. Kong, and X. Chen, “DRCap: Decoding CLAP latents with retrieval-augmented generation for zero-shot audio captioning,” Proc. ICASSP , 2025
2025
Closest in time.