Fetching the paper…
Reading the bibliography…
This work studies the capabilities of a large language model (LLM) to understand paralinguistic aspects of speech without fine-tuning its weights.
W. Kang and D. Roy, “Prompting Large Language Models with Audio for General-Purpose Speech Summarization,” in Proc. Interspeech , 2024, pp. 1955–1959
1959
Earlier work this paper cites.
D. Tannen, Spoken and Written Language: Exploring Orality and Literacy , ser. Advances in Discourse Processes. ABLEX Publishing Corporation, 1982
1982
Earlier work this paper cites.
C.-Y. Lin, “ROUGE: A Package for Automatic Evaluation of Summaries,” in Text Summarization Branches Out , 2004, pp. 74–81
2004
Earlier work this paper cites.
L. Van der Maaten and G. Hinton, “Visualizing Data using t-SNE,” Journal of Machine Learning Research , vol. 9, no. 11, 2008
2008
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” in International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al. , “Conformer: Convolution-augmented Transformer for Speech Recognition,” in Proc. Interspeech , 2020, pp. 5036–5040
2020
Earlier work this paper cites.
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, “BERTScore: Evaluating Text Generation with BERT,” in International Conference on Learning Representations , 2020
2020
Earlier work this paper cites.
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le, “Finetuned Language Models are Zero-Shot Learners,” in International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in Neural Information Processing Systems , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-Rank Adaptation of Large Language Models,” in International Conference on Learning Representations , 2022
2022
Earlier work this paper cites.
C.-C. Chiu, J. Qin, Y. Zhang, J. Yu, and Y. Wu, “Self-Supervised Learning with Random-Projection Quantizer for Speech Recognition,” in International Conference on Machine Learning . PMLR, 2022, pp. 3915–3924
2022
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Fathullah, C. Wu, E. Lakomkin, J. Jia, Y. Shangguan, K. Li, J. Guo, W. Xiong, J. Mahadeokar, O. Kalinli et al. , “Prompting Large Language Models with Speech Recognition Abilities,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 13 351–13 355
2024
Closest in time.
C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “SALMONN: Towards Generic Hearing Abilities for Large Language Models,” in International Conference on Learning Representations , 2024
2024
Closest in time.
2024
Closest in time.
Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. Glass, “Listen, Think, and Understand,” in International Conference on Learning Representations , 2024
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Wu, Y. Gaur, Z. Chen, L. Zhou, Y. Zhu, T. Wang, J. Li, S. Liu, B. Ren, L. Liu et al. , “On Decoder-Only Architecture For Speech-to-Text and Large Language Model Integration,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2023, pp. 1–8
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Gong, A. H. Liu, H. Luo, L. Karlinsky, and J. Glass, “Joint Audio and Speech Understanding,” in IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2023, pp. 1–8
2023
Cited alongside, same era.
2023
Cited alongside, same era.
T. A. Nguyen, W.-N. Hsu, A. D’Avirro, B. Shi, I. Gat, M. Fazel-Zarani, T. Remez, J. Copet, G. Synnaeve, M. Hassid, F. Kreuk, Y. Adi, and E. Dupoux, “Expresso: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis,” in Proc. Interspeech , 2023, pp. 4823–4827
2023
Cited alongside, same era.
Y. Gong, S. Khurana, L. Karlinsky, and J. Glass, “Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers,” in Proc. Interspeech , 2023, pp. 2798–2802
2023
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Closest in time.
Y. Fathullah, C. Wu, E. Lakomkin, K. Li, J. Jia, Y. Shangguan, J. Mahadeokar, O. Kalinli, C. Fuegen, and M. Seltzer, “AudioChatLlama: Towards General-Purpose Speech Abilities for LLMs,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , 2024, pp. 5522–5532
2024
Closest in time.
2024
Closest in time.
G.-T. Lin, P. G. Shivakumar, A. Gandhe, C.-H. H. Yang, Y. Gu, S. Ghosh, A. Stolcke, H.-y. Lee, and I. Bulyko, “Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue,” in IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 10 316–10 320
2024
Closest in time.
G.-T. Lin, C.-H. Chiang, and H.-y. Lee, “Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 6626–6642
2024
Closest in time.
2024
Closest in time.
Y. Xu, H. Chen, J. Yu, Q. Huang, Z. Wu, S.-X. Zhang, G. Li, Y. Luo, and R. Gu, “SECap: Speech Emotion Captioning with Large Language Model,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 17, 2024, pp. 19 323–19 331
2024
Closest in time.