Fetching the paper…
Reading the bibliography…
Video-based multimodal large language models (Video-LLMs) possess significant potential for video understanding tasks.
Collecting highly parallel data for paraphrase evaluation
Chen, D.; and Dolan, W. B. 2011 · 2011
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Caba Heilbron, F.; Escorcia, V.; Ghanem, B.; and Carlos Niebles, J. 2015 · 2015
Earlier work this paper cites.
Spatiotemporal residual networks for video action recognition
Christoph, R.; and Pinz, F. A. 2016 · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Wang, L.; Xiong, Y.; Wang, Z.; Qiao, Y.; Lin, D.; Tang, X.; and Van Gool, L. 2016 · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J.; Mei, T.; Yao, T.; and Rui, Y. 2016 · 2016
Earlier work this paper cites.
Temporal modeling approaches for large-scale youtube-8m video understanding
Li, F.; Gan, C.; Liu, X.; Bian, Y.; Long, X.; Li, Y.; Li, Z.; Zhou, J.; and Wen, S. 2017 · 2017
Earlier work this paper cites.
Learnable pooling with context gating for video classification
Miech, A.; Laptev, I.; and Sivic, J. 2017 · 2017
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
Long-term feature banks for detailed video understanding
Wu, C.-Y.; Feichtenhofer, C.; Fan, H.; He, K.; Krahenbuhl, P.; and Girshick, R. 2019 · 2019
Earlier work this paper cites.
Multiscale vision transformers
Fan, H.; Xiong, B.; Mangalam, K.; Li, Y.; Yan, Z.; Malik, J.; and Feichtenhofer, C. 2021 · 2021
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J.; Li, D.; Xiong, C.; and Hoi, S. 2022 · 2022
Earlier work this paper cites.
Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition
Wu, C.-Y.; Li, Y.; Mangalam, K.; Fan, H.; Xiong, B.; Malik, J.; and Feichtenhofer, C. 2022 · 2022
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Sharegpt4v: Improving large multi-modal models with better captions
Chen, L.; Li, J.; Dong, X.; Zhang, P.; He, C.; Wang, J.; Zhao, F.; and Lin, D. 2023 · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; et al. 2023 · 2023
Cited alongside, same era.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023 · 2023
Cited alongside, same era.
Eva: Exploring the limits of masked visual representation learning at scale
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q.; Xu, H.; Xu, G.; Ye, J.; Yan, M.; Zhou, Y.; Wang, J.; Hu, A.; Shi, P.; Shi, Y.; et al. 2023 · 2023
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding
Zhang, H.; Li, X.; and Bing, L. 2023 · 2023
Later among the works it cites.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
Zhang, R.; Han, J.; Liu, C.; Gao, P.; Zhou, A.; Hu, X.; Yan, S.; Lu, P.; Li, H.; and Qiao, Y. 2023 · 2023
Later among the works it cites.
The claude 3 model family: Opus, sonnet, haiku
Anthropic, A. 2024 · 2024
Closest in time.
Sharegpt4video: Improving video understanding and generation with better captions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fang, Y.; Wang, W.; Xie, B.; Sun, Q.; Wu, L.; Wang, X.; Huang, T.; Wang, X.; and Cao, Y. 2023 · 2023
Cited alongside, same era.
LLaMA-VID: An image is worth 2 tokens in large language models
Li, Y.; Wang, C.; and Jia, J. 2023 · 2023
Cited alongside, same era.
Video-llava: Learning united visual representation by alignment before projection
Lin, B.; Zhu, B.; Ye, Y.; Ning, M.; Jin, P.; and Yuan, L. 2023 · 2023
Cited alongside, same era.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Maaz, M.; Rasheed, H.; Khan, S.; and Khan, F. S. 2023 · 2023
Cited alongside, same era.
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Ren, S.; Yao, L.; Li, S.; Sun, X.; and Hou, L. 2023 · 2023
Cited alongside, same era.
Moviechat: From dense token to sparse memory for long video understanding
Song, E.; Chai, W.; Wang, G.; Zhang, Y.; Zhou, H.; Wu, F.; Guo, X.; Ye, T.; Lu, Y.; Hwang, J.-N.; et al. 2023 · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Cited alongside, same era.
Chen, L.; Wei, X.; Li, J.; Dong, X.; Zhang, P.; Zang, Y.; Chen, Z.; Duan, H.; Lin, B.; Tang, Z.; et al. 2024 · 2024
Closest in time.
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024 · 2024
Closest in time.
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Fang, X.; Mao, K.; Duan, H.; Zhao, X.; Li, Y.; Lin, D.; and Chen, K. 2024 · 2024
Closest in time.
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Fu, C.; Dai, Y.; Luo, Y.; Li, L.; Ren, S.; Zhang, R.; Wang, Z.; Zhou, C.; Shen, Y.; Zhang, M.; et al. 2024 · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M.; Savinov, N.; Teplyashin, D.; Lepikhin, D.; Lillicrap, T.; Alayrac, J.-b.; Soricut, R.; Lazaridou, A.; Firat, O.; Schrittwieser, J.; et al. 2024 · 2024
Closest in time.
LVBench: An Extreme Long Video Understanding Benchmark
Wang, W.; He, Z.; Hong, W.; Cheng, Y.; Zhang, X.; Qi, J.; Huang, S.; Xu, B.; Dong, Y.; Ding, M.; et al. 2024 · 2024
Closest in time.
Pllava: Parameter-free llava extension from images to videos for video dense captioning
Xu, L.; Zhao, Y.; Zhou, D.; Lin, Z.; Ng, S. K.; and Feng, J. 2024 · 2024
Closest in time.