Fetching the paper…
Reading the bibliography…
The impressive capability and versatility of large language models (LLMs) have aroused increasing attention in automatic speech recognition (ASR), with several pioneering studies attempting to build integrated ASR models by connecting a speech encoder with an LLM.
“Librispeech: An ASR corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, et al., · 2017
Earlier work this paper cites.
“Language models are few-shot learners,”
T. Brown, B. Mann, N. Ryder, et al., · 2020
Earlier work this paper cites.
“Common Voice: A massively-multilingual speech corpus,”
R. Ardila, M. Branson, K. Davis, et al., · 2020
Earlier work this paper cites.
W. Zeng, X. Ren, T. Su, et al., · 2021
Earlier work this paper cites.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
W. Hsu, B. Bolte, Y. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, · 2021
Earlier work this paper cites.
“W2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,”
Y. Chung, Y. Zhang, W. Han, C. Chiu, J. Qin, R. Pang, and Y. Wu, · 2021
Earlier work this paper cites.
“GigaSpeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,”
G. Chen, S. Chai, G. Wang, et al., · 2021
Earlier work this paper cites.
“Flamingo: A visual language model for few-shot learning,”
J.B. Alayrac, J. Donahue, P. Luc, et al., · 2022
Earlier work this paper cites.
OpenAI, · 2023
Earlier work this paper cites.
R. Anil, A.M. Dai, O. Firat, et al., · 2023
Earlier work this paper cites.
“LLaMA: Open and efficient foundation language models,”
H. Touvron, T. Lavril, et al., · 2023
Earlier work this paper cites.
“Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality,”
W. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, et al., · 2023
Earlier work this paper cites.
“AudioGPT: Understanding and generating speech, music, sound, and talking head,”
R. Huang, M. Li, D. Yang, et al., · 2023
Cited alongside, same era.
“HuggingGPT: Solving ai tasks with chatGPT and its friends in huggingface,”
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang, · 2023
Cited alongside, same era.
“SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,”
D. Zhang, S. Li, X. Zhang, J. Zhan, P. Wang, Y. Zhou, and X. Qiu, · 2023
Cited alongside, same era.
“AudioPaLM: A large language model that can speak and listen,”
P.K. Rubenstein, C. Asawaroengchai, D.D. Nguyen, et al., · 2023
Cited alongside, same era.
“Macaw-LLM: Multi-modal language modeling with image, audio, video, and text integration,”
C. Lyu, M. Wu, L. Wang, et al., · 2023
Closest in time.
“BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,”
J. Li, D. Li, S. Savarese, and S. Hoi, · 2023
Closest in time.
“Google USM: Scaling automatic speech recognition beyond 100 languages,”
Y. Zhang, W. Han, et al., · 2023
Closest in time.
“MiniGPT-4: Enhancing vision-language understanding with advanced large language models,”
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny, · 2023
Closest in time.
“InstructBLIP: Towards general-purpose vision-language models with instruction tuning,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Chen, M. Han, H. Zhao, Q. Zhang, J. Shi, S. Xu, and B. Xu, · 2023
Cited alongside, same era.
“LLaSM: Large language and speech model,”
Y. Shu, S. Dong, G. Chen, et al., · 2023
Cited alongside, same era.
“On decoder-only architecture for speech-to-text and large language model integration,”
J. Wu, Y. Gaur, Z. Chen, et al., · 2023
Cited alongside, same era.
“Prompting large language models with speech recognition abilities,”
Y. Fathullah, C. Wu, E. Lakomkin, J. Jia, Y. Shangguan, K. Li, J. Guo, W. Xiong, J. Mahadeokar, O. Kalinli, et al., · 2023
Cited alongside, same era.
“Prompting large language models for zero-shot domain adaptation in speech recognition,”
Y. Li, Y. Wu, J. Li, and S. Liu, · 2023
Cited alongside, same era.
“Can generative large language models perform asr error correction?,”
R. Ma, M. Qian, P. Manakul, M. Gales, and K. Knill, · 2023
Cited alongside, same era.
“Leveraging large language models for exploiting asr uncertainty,”
P. Dighe, Y. Su, S. Zheng, et al., · 2023
Cited alongside, same era.
“Robust speech recognition via large-scale weak supervision,”
A. Radford, J. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, · 2023
Cited alongside, same era.
W. Dai, J. Li, D. Li, A.M.H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, · 2023
Closest in time.
“Video-LLaMA: An instruction-tuned audio-visual language model for video understanding,”
H. Zhang, X. Li, and L. Bing, · 2023
Closest in time.
“PandaGPT: One model to instruction-follow them all,”
Y. Su, T. Lan, H. Li, J. Xu, Y. Wang, and D. Cai, · 2023
Closest in time.
“VideoLLM: Modeling video sequence with large language models,”
G. Chen, Y.D. Zheng, et al., · 2023
Closest in time.
“Video-ChatGPT: Towards detailed video understanding via large vision and language models,”
M. Maaz, H. Rasheed, S. Khan, and F.S. Khan, · 2023
Closest in time.
“Listen, think, and understand,”
Y. Gong, H. Luo, A.H. Liu, L. Karlinsky, and J. Glass, · 2023
Closest in time.
S. Liu, A. Hussain, C. Sun, and Y. Shan, · 2023
Closest in time.
“Random utterance concatenation based data augmentation for improving short-video speech recognition,”
Y. Lin, T. Han, H. Xu, et al., · 2023
Closest in time.