Fetching the paper…
Reading the bibliography…
Video large language models (Video-LLMs) have made significant progress in understanding videos.
The coincidence approach to stochastic point processes
Odile Macchi · 1975
Earlier work this paper cites.
An analysis of approximations for maximizing submodular set functions—i
George L Nemhauser, Laurence A Wolsey, and Marshall L Fisher · 1978
Earlier work this paper cites.
Submodular functions and convexity
László Lovász · 1983
Earlier work this paper cites.
Key frame selection by motion analysis
Wayne Wolf · 1996
Earlier work this paper cites.
Submodular functions and optimization
Satoru Fujishige · 2005
Earlier work this paper cites.
Support vector machines and kernel algorithms
Bernhard Schölkopf and Alex Smola · 2005
Earlier work this paper cites.
Determinantal point processes for machine learning
Alex Kulesza, Ben Taskar, et al · 2012
Earlier work this paper cites.
Key frame extraction using edge change ratio for shot segmentation
Azra Nasreen and G Dr Shobha · 2013
Earlier work this paper cites.
Diverse sequential subset selection for supervised video summarization
Boqing Gong, Wei-Lun Chao, Kristen Grauman, and Fei Sha · 2014
Earlier work this paper cites.
Submodular function maximization
Andreas Krause and Daniel Golovin · 2014
Earlier work this paper cites.
Learning transferable features with deep adaptation networks
Mingsheng Long, Yue Cao, Jianmin Wang, and Michael Jordan · 2015
Earlier work this paper cites.
Key-frame extraction by analysis of histograms of video frames using statistical methods
C Va Sheena and NK Narayanan · 2015
Earlier work this paper cites.
Fair and diverse DPP-based data summarization
Elisa Celis, Vijay Keswani, Damian Straszak, Amit Deshpande, Tarun Kathuria, and Nisheeth Vishnoi · 2018
Earlier work this paper cites.
Fast greedy map inference for determinantal point process to improve recommendation diversity
Laming Chen, Guoxin Zhang, and Eric Zhou · 2018
Earlier work this paper cites.
An efficient evolutionary algorithm for subset selection with general cost constraints
Chao Bian, Chao Feng, Chao Qian, and Yang Yu · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Earlier work this paper cites.
Katna: Tool for automating video keyframe extraction, video compression, image autocrop and smart image resize tasks
KeplerLab · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
k-SDPP: fixed-size video summarization via sequential determinantal point processes
Jiping Zheng and Ganfeng Lu · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Cited alongside, same era.
BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Cited alongside, same era.
Qwen-VL: A frontier large vision-language model with versatile abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou · 2023
Cited alongside, same era.
VideoChat: Chat-centric video understanding
KunChang Li, Yinan He, Yi Wang, Yizhuo Li, Wenhai Wang, Ping Luo, Yali Wang, Limin Wang, and Yu Qiao · 2023
Cited alongside, same era.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Video-LLaVA: Learning united visual representation by alignment before projection
Bin Lin, Bin Zhu, Yang Ye, Munan Ning, Peng Jin, and Li Yuan · 2024
Later among the works it cites.
LLaVA-NeXT: Improved reasoning, OCR, and world knowledge
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee · 2024
Later among the works it cites.
Ovis: Structural embedding alignment for multimodal large language model
Shiyin Lu, Yang Li, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, and Han-Jia Ye · 2024
Later among the works it cites.
Video-XL: Extra-long vision language model for hour-scale video understanding
Yan Shu, Peitian Zhang, Zheng Liu, Minghao Qin, Junjie Zhou, Tiejun Huang, and Bo Zhao · 2024
Later among the works it cites.
MovieChat: From dense token to sparse memory for long video understanding
Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Towards efficient generative large language model serving: A survey from algorithms to systems
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Hongyi Jin, Tianqi Chen, and Zhihao Jia · 2023
Cited alongside, same era.
Inducing point allocation for sparse gaussian processes in high-throughput bayesian optimisation
Henry B Moss, Sebastian W Ober, and Victor Picheny · 2023
Cited alongside, same era.
ChatGPT: Optimizing language models for dialogue
OpenAI · 2023
Cited alongside, same era.
Enhancing unsupervised domain adaptation by exploiting the conceptual consistency of multiple self-supervised tasks
Hui Sun and Ming Li · 2023
Cited alongside, same era.
Compositional exemplars for in-context learning
Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu, and Lingpeng Kong · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer · 2023
Cited alongside, same era.
Introducing the next generation of Claude
AI Anthropic · 2024
Cited alongside, same era.
Later among the works it cites.
Koala: Key frame-conditioned long Video-LLM
Reuben Tan, Ximeng Sun, Ping Hu, Jui hsien Wang, Hanieh Deilamsalehy, Bryan A. Plummer, Bryan Russell, and Kate Saenko · 2024
Later among the works it cites.
Efficient large language models: A survey
Zhongwei Wan, Xin Wang, Che Liu, Samiul Alam, Yu Zheng, Jiachen Liu, Zhongnan Qu, Shen Yan, Yi Zhu, Quanlu Zhang, Mosharaf Chowdhury, and Mi Zhang · 2024
Later among the works it cites.
LongVideoBench: A benchmark for long-context interleaved video-language understanding
Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li · 2024
Later among the works it cites.
Can I trust your answer? visually grounded video question answering
Junbin Xiao, Angela Yao, Yicong Li, and Tat-Seng Chua · 2024
Later among the works it cites.
Effective long-context scaling of foundation models
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, et al · 2024
Later among the works it cites.
LongVILA: Scaling long-context visual language models for long videos
Fuzhao Xue, Yukang Chen, Dacheng Li, Qinghao Hu, Ligeng Zhu, Xiuyu Li, Yunhao Fang, Haotian Tang, Shang Yang, Zhijian Liu, et al · 2024
Later among the works it cites.
MiniCPM-V: A GPT-4V level MLLM on your phone
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, et al · 2024
Later among the works it cites.
Frame-Voyager: Learning to query frames for video large language models
Sicheng Yu, Chengkai Jin, Huanyu Wang, Zhenghao Chen, Sheng Jin, Zhongrong Zuo, Xioalei Xu, Zhenbang Sun, Bingni Zhang, Jiawei Wu, et al · 2024
Later among the works it cites.
MLVU: A comprehensive benchmark for multi-task long video understanding
Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Shitao Xiao, Xi Yang, Yongping Xiong, Bo Zhang, Tiejun Huang, and Zheng Liu · 2024
Later among the works it cites.
M-LLM based video frame selection for efficient video understanding
Kai Hu, Feng Gao, Xiaohan Nie, Peng Zhou, Son Tran, Tal Neiman, Lingyun Wang, Mubarak Shah, Raffay Hamid, Bing Yin, et al · 2025
Closest in time.
TimeCraft: Navigate weakly-supervised temporal grounded video question answering via bi-directional reasoning
Huabin Liu, Xiao Ma, Cheng Zhong, Yang Zhang, and Weiyao Lin · 2025
Closest in time.
Adaptive keyframe sampling for long video understanding
Xi Tang, Jihao Qiu, Lingxi Xie, Yunjie Tian, Jianbin Jiao, and Qixiang Ye · 2025
Closest in time.
ViLA: Efficient video-language alignment for video question answering
Xijun Wang, Junbang Liang, Chun-Kai Wang, Kenan Deng, Yu Lou, Ming C. Lin, and Shan Yang · 2025
Closest in time.