Fetching the paper…
Reading the bibliography…
Recent advancements in multimodal large language models (MLLMs) have opened new avenues for video understanding.
Collecting highly parallel data for paraphrase evaluation
David Chen and William Dolan · 2011
Earlier work this paper cites.
Tgif: A new dataset and benchmark on animated gif description
Yuncheng Li, Yale Song, Liangliang Cao, Joel Tetreault, Larry Goldberg, Alejandro Jaimes, and Jiebo Luo · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Tsm: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Earlier work this paper cites.
Efficient parameter-free clustering using first neighbor relations
Saquib Sarfraz, Vivek Sharma, and Rainer Stiefelhagen · 2019
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Zhou Yu, Dejing Xu, Jun Yu, Ting Yu, Zhou Zhao, Yueting Zhuang, and Dacheng Tao · 2019
Earlier work this paper cites.
Is space-time attention all you need for video understanding?
Gedas Bertasius, Heng Wang, and Lorenzo Torresani · 2021
Earlier work this paper cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Earlier work this paper cites.
Uniformer: Unified transformer for efficient spatiotemporal representation learning, 2022
Kunchang Li, Yali Wang, Peng Gao, Guanglu Song, Yu Liu, Hongsheng Li, and Yu Qiao · 2022
Earlier work this paper cites.
Video swin transformer
Ze Liu, Jia Ning, Yue Cao, Yixuan Wei, Zheng Zhang, Stephen Lin, and Han Hu · 2022
Earlier work this paper cites.
Less is more: Pay less attention in vision transformers
Zizheng Pan, Bohan Zhuang, Haoyu He, Jing Liu, and Jianfei Cai · 2022
Earlier work this paper cites.
Token merging: Your ViT but faster
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman · 2023
Cited alongside, same era.
Evaluating open-domain question answering in the era of large language models
Ehsan Kamalloo, Nouha Dziri, Charles LA Clarke, and Davood Rafiei · 2023
Cited alongside, same era.
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection, 2023
Bin Lin, Yang Ye, Bin Zhu, Jiaxi Cui, Munan Ning, Peng Jin, and Li Yuan · 2023
Cited alongside, same era.
Vista-LLaMA: Reliable Video Narrator via Equal Distance to Visual Tokens, 2023
Fan Ma, Xiaojie Jin, Heng Wang, Yuchen Xian, Jiashi Feng, and Yi Yang · 2023
Cited alongside, same era.
Egoschema: A diagnostic benchmark for very long-form video language understanding
Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik · 2023
Cited alongside, same era.
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding, 2024
Peng Jin, Ryuichi Takanobu, Wancai Zhang, Xiaochun Cao, and Li Yuan · 2024
Closest in time.
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM, 2024
Wonkyun Kim, Changin Choi, Wonseok Lee, and Wonjong Rhee · 2024
Closest in time.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Muhammad Maaz, Hanoona Rasheed, Salman Khan, and Fahad Shahbaz Khan · 2024
Closest in time.
Foundation models for video understanding: A survey
Neelu Madan, Andreas Møgelmose, Rajat Modi, Yogesh S Rawat, and Thomas B Moeslund · 2024
Closest in time.
Mm1: Methods, analysis & insights from multimodal llm pre-training
Brandon McKinzie, Zhe Gan, Jean-Philippe Fauconnier, Sam Dodge, Bowen Zhang, Philipp Dufter, Dhruti Shah, Xianzhi Du, Futang Peng, Floris Weers, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yunlong Tang, Jing Bi, Siting Xu, Luchuan Song, Susan Liang, Teng Wang, Daoan Zhang, Jie An, Jingyang Lin, Rongyi Zhu, et al · 2023
Cited alongside, same era.
Self-chained image-language model for video localization and question answering
Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal · 2023
Cited alongside, same era.
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding, 2023
Hang Zhang, Xin Li, and Lidong Bing · 2023
Cited alongside, same era.
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs, 2024
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and Lidong Bing · 2024
Cited alongside, same era.
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis, 2024
Chaoyou Fu, Yuhan Dai, Yondong Luo, Lei Li, Shuhuai Ren, Renrui Zhang, Zihan Wang, Chenyu Zhou, Yunhang Shen, Mengdan Zhang, Peixian Chen, Yanwei Li, Shaohui Lin, Sirui Zhao, Ke Li, Tong Xu, Xiawu Zheng, Enhong Chen, Rongrong Ji, and Xing Sun · 2024
Cited alongside, same era.
Vlm2vec: Training vision-language models for massive multimodal embedding tasks
Ziyan Jiang, Rui Meng, Xinyi Yang, Semih Yavuz, Yingbo Zhou, and Wenhu Chen · 2024
Cited alongside, same era.
Shimin Chen, Yitian Yuan, Shaoxiang Chen, Zequn Jie, and Lin Ma
Cited in the paper.
Closest in time.
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding, 2024
Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, Yan Lu, Jenq-Neng Hwang, and Gaoang Wang · 2024
Closest in time.
Internvideo2: Scaling foundation models for multimodal video understanding
Yi Wang, Kunchang Li, Xinhao Li, Jiashuo Yu, Yinan He, Guo Chen, Baoqi Pei, Rongkun Zheng, Zun Wang, Yansong Shi, et al · 2024
Closest in time.
Freeva: Offline mllm as training-free video assistant
Wenhao Wu · 2024
Closest in time.
Self-chained image-language model for video localization and question answering
Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal · 2024
Closest in time.
Multimodal guidance network for missing-modality inference in content moderation
Zhuokai Zhao, Harish Palani, Tianyi Liu, Lena Evans, and Ruth Toner · 2024
Closest in time.
A survey on generative ai and llm for video generation, understanding, and streaming
Pengyuan Zhou, Lin Wang, Zhi Liu, Yanbin Hao, Pan Hui, Sasu Tarkoma, and Jussi Kangasharju · 2024
Closest in time.