Fetching the paper…
Reading the bibliography…
Speech understanding as an element of the more generic video understanding using audio-visual large language models (av-LLMs) is a crucial yet understudied aspect.
The rich transcription 2004 spring meeting recognition evaluation
Garofolo, J. S., Fiscus, J. G., and Laprun, C. D · 2004
Earlier work this paper cites.
The rich transcription 2005 spring meeting recognition evaluation
Fiscus, J. G., Radde, N., Garofolo, J. S., Le, A., Ajot, J., and Laprun, C · 2005
Earlier work this paper cites.
The rich transcription 2006 spring meeting recognition evaluation
Fiscus, J. G., Ajot, J., Michel, M., and Garofolo, J. S · 2006
Earlier work this paper cites.
The rich transcription 2007 meeting recognition evaluation
Fiscus, J. G., Ajot, J., and Garofolo, J. S · 2007
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Bourdev, L., Girshick, R., Hays, J., Perona, P., Ramanan, D., Zitnick, C. L., and Dollár, P · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Young, P., Lai, A., Hodosh, M., and Hockenmaier, J · 2014
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Improved image captioning via policy gradient optimization of spider
Liu, S., Zhu, Z., Ye, N., Guadarrama, S., and Murphy, K · 2017
Earlier work this paper cites.
How2: A large-scale dataset for multimodal language understanding
Sanabria, R., Caglayan, O., Palaskar, S., Elliott, D., Barrault, L., Specia, L., and Metze, F · 2018
Earlier work this paper cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
AudioCaps: Generating captions for audios in the wild
Kim, C. D., Kim, B., Lee, H., and Kim, G · 2019
Earlier work this paper cites.
OCR-VQA: Visual question answering by reading text in images
Mishra, A., Shekhar, S., Singh, A. K., and Chakraborty, A · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., et al · 2020
Earlier work this paper cites.
VGGSound: A large-scale audio-visual dataset
Chen, H., Xie, W., Vedaldi, A., and Zisserman, A · 2020
Earlier work this paper cites.
Text-free image-to-speech synthesis using learned segmental units
Hsu, W.-N., Harwath, D., Song, C., and Glass, J · 2020
Earlier work this paper cites.
Textcaps: a dataset for image captioningwith reading comprehension
Sidorov, O., Hu, R., Rohrbach, M., and Singh, A · 2020
Cited alongside, same era.
NExT-QA: Next phase of question-answering to explaining temporal actions
Xiao, J., Shang, X., Yao, A., and Chua, T.-S · 2021
Cited alongside, same era.
Flamingo: A visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., et al · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., et al · 2022
Cited alongside, same era.
GLM: General language model pretraining with autoregressive blank infilling
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J · 2022
Cited alongside, same era.
Ego4D: Around the world in 3,000 hours of egocentric video
Macaw-LLM: Multi-modal language modeling with image, audio, video, and text integration
Lyu, C., Wu, M., Wang, L., Huang, X., Liu, B., Du, Z., Shi, S., and Tu, Z · 2023
Later among the works it cites.
Video-ChatGPT: Towards detailed video understanding via large vision and language models
Maaz, M., Rasheed, H., Khan, S., and Khan, F. S · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Peng, B., Li, C., He, P., Galley, M., and Gao, J · 2023
Later among the works it cites.
Mirasol3b: A multimodal autoregressive model for time-aligned and contextual modalities
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grauman, K., Westbury, A., Byrne, E., et al · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., et al · 2022
Cited alongside, same era.
Learning video representations from large language models
Zhao, Y., Misra, I., Krähenbühl, P., and Girdhar, R · 2022
Cited alongside, same era.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., et al · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing GPT-4 with 90%* ChatGPT quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
InstructBLIP: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Cited alongside, same era.
Piergiovanni, A., Noble, I., Kim, D., Ryoo, M. S., Gomes, V., and Angelova, A · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I · 2023
Later among the works it cites.
AudioPaLM: A large language model that can speak and listen
Rubenstein, P. K., Asawaroengchai, C., Nguyen, D. D., et al · 2023
Later among the works it cites.
PandaGPT: One model to instruction-follow them all
Su, Y., Lan, T., Li, H., Xu, J., Wang, Y., and Cai, D · 2023
Later among the works it cites.
Fine-grained audio-visual joint representations for multimodal large language models
Sun, G., Yu, W., Tang, C., Chen, X., Tan, T., Li, W., Lu, L., Ma, Z., and Zhang, C · 2023
Later among the works it cites.
SALMONN: Towards generic hearing abilities for large language models
Tang, C., Yu, W., Sun, G., Chen, X., Tan, T., Li, W., Lu, L., Ma, Z., and Zhang, C · 2023
Later among the works it cites.
LLaMA: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
What matters in training a gpt4-style language model with multimodal inputs?
Zeng, Y., Zhang, H., Zheng, J., Xia, J., Wei, G., Wei, Y., Zhang, Y., and Kong, T · 2023
Later among the works it cites.
BuboGPT: Enabling visual grounding in multi-modal LLMs
Zhao, Y., Lin, Z., Zhou, D., Huang, Z., Feng, J., and Kang, B · 2023
Later among the works it cites.
Extending large language models for speech and audio captioning
Tang, C., Yu, W., Sun, G., Chen, X., Tan, T., Li, W., Lu, L., Ma, Z., and Zhang, C · 2024
Closest in time.
Connecting speech encoder and large language model for ASR
Yu, W., Tang, C., Sun, G., Chen, X., Tan, T., Li, W., Lu, L., Ma, Z., and Zhang, C · 2024
Closest in time.