Fetching the paper…
Reading the bibliography…
Recently, integrating visual foundation models into large language models (LLMs) to form video understanding systems has attracted widespread attention.
Videograph: Recognizing minutes-long human activities in videos
Hussein, N.; Gavves, E.; and Smeulders, A. W. 2019b · 1905
Earlier work this paper cites.
Univl: A unified video and language pre-training model for multimodal understanding and generation
Luo, H.; Ji, L.; Shi, B.; Huang, H.; Duan, N.; Li, T.; Li, J.; Bharti, T.; and Zhou, M. 2020 · 2002
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Banerjee, S.; and Lavie, A. 2005 · 2005
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B. 2020 · 2005
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D.; and Dolan, W. B. 2011 · 2011
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
Kuehne, H.; Arslan, A.; and Serre, T. 2014 · 2014
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R.; Lawrence Zitnick, C.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J.; Mei, T.; Yao, T.; and Rui, Y. 2016 · 2016
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Xu, D.; Zhao, Z.; Xiao, J.; Wu, F.; Zhang, H.; He, X.; and Zhuang, Y. 2017 · 2017
Earlier work this paper cites.
Temporal segment networks for action recognition in videos
Wang, L.; Xiong, Y.; Wang, Z.; Qiao, Y.; Lin, D.; Tang, X.; and Van Gool, L. 2018 · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
Zhou, L.; Xu, C.; and Corso, J. 2018 · 2018
Earlier work this paper cites.
Scsampler: Sampling salient clips from video for efficient action recognition
Korbar, B.; Tran, D.; and Torresani, L. 2019 · 2019
Earlier work this paper cites.
Coin: A large-scale dataset for comprehensive instructional video analysis
Tang, Y.; Ding, D.; Rao, Y.; Zheng, Y.; Zhang, D.; Zhao, L.; Lu, J.; and Zhou, J. 2019 · 2019
Earlier work this paper cites.
Adaframe: Adaptive frame selection for fast video recognition
Wu, Z.; Xiong, C.; Ma, C.-Y.; Socher, R.; and Davis, L. S. 2019 · 2019
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Yu, Z.; Xu, D.; Yu, J.; Yu, T.; Zhao, Z.; Zhuang, Y.; and Tao, D. 2019 · 2019
Cited alongside, same era.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2021
Cited alongside, same era.
Graph-based high-order relation modeling for long-term action recognition
Zhou, J.; Lin, K.-Y.; Li, H.; and Zheng, W.-S. 2021 · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. 2022 · 2022
Cited alongside, same era.
Long movie clip classification with state-space video models
Islam, M. M.; and Bertasius, G. 2022 · 2022
Cited alongside, same era.
Llama-vid: An image is worth 2 tokens in large language models
Li, Y.; Wang, C.; and Jia, J. 2023 · 2023
Later among the works it cites.
Video-llava: Learning united visual representation by alignment before projection
Lin, B.; Zhu, B.; Ye, Y.; Ning, M.; Jin, P.; and Yuan, L. 2023 · 2023
Later among the works it cites.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Maaz, M.; Rasheed, H.; Khan, S.; and Khan, F. S. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, J.; Yang, Z.; Hu, X.; Li, L.; Lin, K.; Gan, Z.; Liu, Z.; Liu, C.; and Wang, L. 2022 · 2022
Cited alongside, same era.
VideoCoCa: Video-text modeling with zero-shot transfer from contrastive captioners
Yan, S.; Zhu, T.; Wang, Z.; Cao, Y.; Zhang, M.; Ghosh, S.; Wu, Y.; and Yu, J. 2022 · 2022
Cited alongside, same era.
Zero-shot video question answering via frozen bidirectional language models
Yang, A.; Miech, A.; Sivic, J.; Laptev, I.; and Schmid, C. 2022 · 2022
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; et al. 2023 · 2023
Cited alongside, same era.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. C. H. 2023 · 2023
Cited alongside, same era.
Eva: Exploring the limits of masked visual representation learning at scale
Fang, Y.; Wang, W.; Xie, B.; Sun, Q.; Wu, L.; Wang, X.; Huang, T.; Wang, X.; and Cao, Y. 2023 · 2023
Cited alongside, same era.
Wang, J.; Zhu, W.; Wang, P.; Yu, X.; Liu, L.; Omar, M.; and Hamid, R. 2023 · 2023
Later among the works it cites.
mplug-2: A modularized multi-modal foundation model across text, image and video
Xu, H.; Ye, Q.; Yan, M.; Shi, Y.; Ye, J.; Xu, Y.; Li, C.; Bi, B.; Qian, Q.; Wang, W.; et al. 2023 · 2023
Later among the works it cites.
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Zhang, H.; Li, X.; and Bing, L. 2023 · 2023
Later among the works it cites.
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Cheng, Z.; Leng, S.; Zhang, H.; Xin, Y.; Li, X.; Chen, G.; Zhu, Y.; Zhang, W.; Luo, Z.; Zhao, D.; et al. 2024 · 2024
Closest in time.
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
Guo, Y.; Liu, J.; Li, M.; Tang, X.; Chen, X.; and Zhao, B. 2024 · 2024
Closest in time.
Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
He, B.; Li, H.; Jang, Y. K.; Jia, M.; Cao, X.; Shah, A.; Shrivastava, A.; and Lim, S.-N. 2024 · 2024
Closest in time.
Visual instruction tuning
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2024 · 2024
Closest in time.
Pllava: Parameter-free llava extension from images to videos for video dense captioning
Xu, L.; Zhao, Y.; Zhou, D.; Lin, Z.; Ng, S. K.; and Feng, J. 2024 · 2024
Closest in time.
mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Ye, Q.; Xu, H.; Ye, J.; Yan, M.; Hu, A.; Liu, H.; Qian, Q.; Zhang, J.; and Huang, F. 2024 · 2024
Closest in time.