Fetching the paper…
Reading the bibliography…
Audio language models process audio inputs using textual prompts for tasks like speech recognition and audio captioning.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “IEMOCAP: Interactive emotional dyadic motion capture database,” Language resources and evaluation , 2008
2008
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
J. F. Gemmeke, D. P. W. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter, “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2017, pp. 776–780
2017
Earlier work this paper cites.
C.-H. Lee, S.-L. Wu, C.-L. Liu, and H.-y. Lee, “Spoken SQuAD: A study of mitigating the impact of speech recognition errors on listening comprehension,” Proc. Interspeech 2018 , 2018
2018
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “Audiocaps: Generating captions for audios in the wild,” in NAACL-HLT , 2019
2019
Earlier work this paper cites.
H. Zen, V. Dang, R. Clark, Y. Zhang, R. J. Weiss, Y. Jia, Z. Chen, and Y. Wu, “Libritts: A corpus derived from librispeech for text-to-speech,” in Proc. Interspeech 2019 , 2019
2019
Earlier work this paper cites.
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common Voice: A Massively-Multilingual Speech Corpus,” in Proceedings of the 12th Conference on Language Resources and Evaluation (LREC 2020) , 2020
2020
Earlier work this paper cites.
K. Drossos, S. Lipping, and T. Virtanen, “Clotho: an audio captioning dataset,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 736–740
2020
Earlier work this paper cites.
Y. Gong, Y.-A. Chung, and J. Glass, “AST: Audio Spectrogram Transformer,” in Proc. Interspeech 2021 , 2021, pp. 571–575
2021
Earlier work this paper cites.
G. C. a Shuzhou Chai, G. Wang, J. Du, W.-Q. Zhang, C. Weng, D. Su, D. Povey, J. Trmal, J. Zhang, M. Jin, S. Khudanpur, S. Watanabe, S. Zhao, W. Zou, X. Li, X. Yao, Y. Wang, Y. Wang, Z. You, and Z. Yan, “GigaSpeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,” in Proc. Interspeech 2021 , 2021
2021
Earlier work this paper cites.
E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra, “Fsd50k: an open dataset of human-labeled sound events,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 829–852, 2021
2021
Earlier work this paper cites.
S. Hershey, D. P. W. Ellis, E. Fonseca, A. Jansen, C. Liu, R. Channing Moore, and M. Plakal, “The benefit of temporally-strong labels in audio event classification,” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 366–370
2021
Earlier work this paper cites.
C. Wang, A. Wu, J. Gu, and J. Pino, “Covost 2 and massively multilingual speech translation,” in Proc. Interspeech 2021 , 2021, pp. 2247–2251
2021
Earlier work this paper cites.
VISTEC, “Thai speech emotion dataset,” 2021. [Online]. Available: https://airesearch.in.th/releases/speech-emotion-dataset
2021
Cited alongside, same era.
E. J. Hu, yelong shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in ICLR , 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: bootstrapping language-image pre-training with frozen image encoders and large language models,” in Proceedings of the 40th International Conference on Machine Learning , ser. ICML’23. JMLR.org, 2023
2023
Later among the works it cites.
X. Li, S. Takamichi, T. Saeki, W. Chen, S. Shiota, and S. Watanabe, “Yodas: Youtube-oriented dataset for audio and speech,” in 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2023, pp. 1–8
2023
Later among the works it cites.
C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. MA, and C. Zhang, “SALMONN: Towards generic hearing abilities for large language models,” in ICLR , 2024
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Gong, A. H. Liu, H. Luo, L. Karlinsky, and J. Glass, “Joint audio and speech understanding,” in ASRU Workshop , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing GPT-4 with 90%* chatgpt quality,” March 2023
2023
Cited alongside, same era.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in ICML . PMLR, 2023, pp. 28 492–28 518
2023
Cited alongside, same era.
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “BEATs: audio pre-training with acoustic tokenizers,” in Proceedings of the 40th ICML , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
A. Zeng, X. Liu, Z. Du, Z. Wang, H. Lai, M. Ding, Z. Yang, Y. Xu, W. Zheng, X. Xia, W. L. Tam, Z. Ma, Y. Xue, J. Zhai, W. Chen, Z. Liu, P. Zhang, Y. Dong, and J. Tang, “GLM-130b: An open bilingual pre-trained model,” in ICLR , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. R. Glass, “Listen, think, and understand,” in ICLR , 2024
2024
Closest in time.
Y. Fathullah, C. Wu, E. Lakomkin, K. Li, J. Jia, Y. Shangguan, J. Mahadeokar, O. Kalinli, C. Fuegen, and M. Seltzer, “AudioChatLlama: Towards general-purpose speech abilities for LLMs,” in NAACL , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
A. Boribalburephan, Z. H. Aung, K. Pipatsrisawat, and T. Achakulvisut, “Thonburian Whisper: A fine-tuned whisper model for Thai automatic speech recognition,” 2024. [Online]. Available: https://huggingface.co/biodatlab/whisper-th-large-v3-combined
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.