Fetching the paper…
Reading the bibliography…
Long-form video understanding is complicated by the high redundancy of video data and the abundance of query-irrelevant information.
Some methods for classification and analysis of multivariate observations
James MacQueen et al · 1967
Earlier work this paper cites.
Cross-modal and hierarchical modeling of video and text
Bowen Zhang, Hexiang Hu, and Fei Sha · 2018
Earlier work this paper cites.
Tsm: Temporal shift module for efficient video understanding
Ji Lin, Chuang Gan, and Song Han · 2019
Earlier work this paper cites.
Liteeval: A coarse-to-fine framework for resource efficient video recognition, 2019
Zuxuan Wu, Caiming Xiong, Yu-Gang Jiang, and Larry S. Davis · 2019
Earlier work this paper cites.
HERO: Hierarchical encoder for video+language omni-representation pre-training, 2020
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu · 2020
Earlier work this paper cites.
Openclip, 2021
Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt · 2021
Earlier work this paper cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Junbin Xiao, Xindi Shang, Angela Yao, and Tat-Seng Chua · 2021
Earlier work this paper cites.
Unsupervised temporal video grounding with deep semantic clustering, 2022
Daizong Liu, Xiaoye Qu, Yinzhen Wang, Xing Di, Kai Zou, Yu Cheng, Zichuan Xu, and Pan Zhou · 2022
Earlier work this paper cites.
Lgdn: Language-guided denoising network for video-language modeling, 2022
Haoyu Lu, Mingyu Ding, Nanyi Fei, Yuqi Huo, and Zhiwu Lu · 2022
Earlier work this paper cites.
Learning from untrimmed videos: Self-supervised video representation learning with hierarchical consistency, 2022
Zhiwu Qing, Shiwei Zhang, Ziyuan Huang, Yi Xu, Xiang Wang, Mingqian Tang, Changxin Gao, Rong Jin, and Nong Sang · 2022
Earlier work this paper cites.
MeMViT: Memory-augmented multiscale vision transformer for efficient long-term video recognition, 2022
Chao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer · 2022
Earlier work this paper cites.
Hiervl: Learning hierarchical video-language embeddings
Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, and Kristen Grauman · 2023
Earlier work this paper cites.
Zero-shot video question answering with procedural programs
Rohan Choudhury, Koichiro Niinuma, Kris M. Kitani, and Laszlo A. Jeni · 2023
Earlier work this paper cites.
Long Story Short: a summarize-then-search method for long video question answering, 2023
Jiwan Chung and Youngjae Yu · 2023
Earlier work this paper cites.
CogAgent: A visual language model for GUI agents, 2023
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxuan Zhang, Juanzi Li, Bin Xu, Yuxiao Dong, Ming Ding, and Jie Tang · 2023
Earlier work this paper cites.
Mistral 7b, 2023
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Earlier work this paper cites.
Large language models are temporal and causal reasoners for video question answering, 2023
Dohwan Ko, Ji Soo Lee, Wooyoung Kang, Byungseok Roh, and Hyunwoo J. Kim · 2023
Earlier work this paper cites.
Intentqa: Context-aware video intent reasoning
Jiapeng Li, Ping Wei, Wenjuan Han, and Lifeng Fan · 2023
Earlier work this paper cites.
Vista-LLaMA: Reliable video narrator via equal distance to visual tokens, 2023
Fan Ma, Xiaojie Jin, Heng Wang, Yuchen Xian, Jiashi Feng, and Yi Yang · 2023
Earlier work this paper cites.
Pg-video-llava: Pixel grounding large video-language models
Shehan Munasinghe, Rusiru Thushara, Muhammad Maaz, Hanoona Abdul Rasheed, Salman Khan, Mubarak Shah, and Fahad Khan · 2023
Cited alongside, same era.
Vipergpt: Visual inference via python execution for reasoning
Dídac Surís, Sachit Menon, and Carl Vondrick · 2023
Cited alongside, same era.
Retrieval-based video language model for efficient long video question answering
Jiaqi Xu, Cuiling Lan, Wenxuan Xie, Xuejin Chen, and Yan Lu · 2023
Cited alongside, same era.
Hierarchical video-moment retrieval and step-captioning
Abhay Zala, Jaemin Cho, Satwik Kottur, Xilun Chen, Barlas Oguz, Yashar Mehdad, and Mohit Bansal · 2023
Cited alongside, same era.
Learning video representations from large language models
Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar · 2023
Cited alongside, same era.
MoReVQA: Exploring modular reasoning models for video question answering
Juhong Min, Shyamal Buch, Arsha Nagrani, Minsu Cho, and Cordelia Schmid · 2024
Closest in time.
GPT-4o blog, 2024
OpenAI · 2024
Closest in time.
Dinov2: Learning robust visual features without supervision, 2024
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Hervé Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski · 2024
Closest in time.
Too many frames, not all useful:efficient strategies for long-form video qa, 2024
Jongwoo Park, Kanchana Ranasinghe, Kumara Kahatapitiya, Wonjeong Ryoo, Donghyun Kim, and Michael S. Ryoo · 2024
Closest in time.
Understanding long videos in one multimodal language model pass, 2024
Kanchana Ranasinghe, Xiang Li, Kumara Kahatapitiya, and Michael S. Ryoo · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al · 2024
Cited alongside, same era.
Videollama 2: Advancing spatial-temporal modeling and audio understanding in video-llms
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, and Lidong Bing · 2024
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid et al · 2024
Cited alongside, same era.
Videoagent: A memory-augmented multimodal agent for video understanding
Yue Fan, Xiaojian Ma, Rujie Wu, Yuntao Du, Jiaqi Li, Zhi Gao, and Qing Li · 2024
Cited alongside, same era.
Ma-lmm: Memory-augmented large multimodal model for long-term video understanding, 2024
Bo He, Hengduo Li, Young Kyun Jang, Menglin Jia, Xuefei Cao, Ashish Shah, Abhinav Shrivastava, and Ser-Nam Lim · 2024
Cited alongside, same era.
Video ReCap: Recursive captioning of hour-long videos
Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan, Lorenzo Torresani, and Gedas Bertasius · 2024
Cited alongside, same era.
Mixtral of experts, 2024
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Cited alongside, same era.
TimeChat: A time-sensitive multimodal large language model for long video understanding, 2024
Shuhuai Ren, Linli Yao, Shicheng Li, Xu Sun, and Lu Hou · 2024
Closest in time.
TV-TREES: Multimodal entailment trees for neuro-symbolic video reasoning, 2024
Kate Sanders, Nathaniel Weir, and Benjamin Van Durme · 2024
Closest in time.
Interpolating video-llms: Toward longer-sequence lmms in a training-free manner, 2024
Yuzhang Shang, Bingxin Xu, Weitai Kang, Mu Cai, Yuheng Li, Zehao Wen, Zhen Dong, Kurt Keutzer, Yong Jae Lee, and Yan Yan · 2024
Closest in time.
Longvu: Spatiotemporal adaptive compression for long video-language understanding, 2024
Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu, Jun Chen, Chenchen Zhu, Zechun Liu, Fanyi Xiao, Balakrishnan Varadarajan, Florian Bordes, Zhuang Liu, Hu Xu, Hyunwoo J. Kim, Bilge Soran, Raghuraman Krishnamoorthi, Mohamed Elhoseiny, and Vikas Chandra · 2024
Closest in time.
Moviechat: From dense token to sparse memory for long video understanding, 2024
Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, Yan Lu, Jenq-Neng Hwang, and Gaoang Wang · 2024
Closest in time.
EVA-CLIP-18B: Scaling clip to 18 billion parameters
Quan Sun, Jinsheng Wang, Qiying Yu, Yufeng Cui, Fan Zhang, Xiaosong Zhang, and Xinlong Wang · 2024
Closest in time.
Koala: Key frame-conditioned long video-LLM, 2024
Reuben Tan, Ximeng Sun, Ping Hu, Jui hsien Wang, Hanieh Deilamsalehy, Bryan A. Plummer, Bryan Russell, and Kate Saenko · 2024
Closest in time.
Longvlm: Efficient long video understanding via large language models, 2024
Yuetian Weng, Mingfei Han, Haoyu He, Xiaojun Chang, and Bohan Zhuang · 2024
Closest in time.
Freeva: Offline mllm as training-free video assistant, 2024
Wenhao Wu · 2024
Closest in time.
Slowfast-llava: A strong training-free baseline for video large language models, 2024
Mingze Xu, Mingfei Gao, Zhe Gan, Hong-You Chen, Zhengfeng Lai, Haiming Gang, Kai Kang, and Afshin Dehghan · 2024
Closest in time.
DoraemonGPT: Toward understanding dynamic scenes with large language models (exemplified as a video agent), 2024
Zongxin Yang, Guikun Chen, Xiaodi Li, Wenguan Wang, and Yi Yang · 2024
Closest in time.
Self-chained image-language model for video localization and question answering
Shoubin Yu, Jaemin Cho, Prateek Yadav, and Mohit Bansal · 2024
Closest in time.
Longvitu: Instruction tuning for long-form video understanding, 2025
Rujie Wu, Xiaojian Ma, Hai Ci, Yue Fan, Yuxuan Wang, Haozhe Zhao, Qing Li, and Yizhou Wang · 2025
Closest in time.